| ▲ | dwohnitmok a day ago |
| This study is pretty bad. The comment (https://news.ycombinator.com/item?id=48970182) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer. This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong. Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong! There are a lot of things I'm very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.). This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors. |
|
| ▲ | encomiast a day ago | parent | next [-] |
| Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study that told experts to lie. A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?". |
| |
| ▲ | RA_Fisher a day ago | parent | next [-] | | Exactly, it’s unrepresentative of AI. It’s damaged AI. | | |
| ▲ | rsoto2 a day ago | parent | next [-] | | All AI is flawed and prone to "hallucinations"(doing exactly what it was designed to do) that's why Microsoft considers it an "Entertainment" product. | | |
| ▲ | rbanffy 19 hours ago | parent | next [-] | | I prefer to say AIs are prone to misremember things, as well as humans do. The more you read and learn, the more material you have to get confused, unfortunately. We might want to work around the certainty with which it misremembers things. I have a very large interval of "I have no idea", or "I don't remember" and rarely caught myself misremembering things (my children are far better at that). I assume my grandchildren will eventually improve their parents' scores by a wide margin. | |
| ▲ | Dylan16807 a day ago | parent | prev | next [-] | | Everything is flawed. But if you're testing the effect of "AI advice" you need to use an AI that has a normal failure rate without tailoring the questions to alter that rate, and/or compare AI versus other sources of information that have the same failure rate. | |
| ▲ | looofooo0 a day ago | parent | prev | next [-] | | Proofing open math conjectures | |
| ▲ | mdp2021 a day ago | parent | prev [-] | | ....Dangerously suggestive post, when you show you don't have a proper idea of "AI" or won't use the term 'AI' properly. | | |
| ▲ | daveguy 18 hours ago | parent [-] | | Yes, definitely a case of bad-think. Certainly not double-plus-good like AI. Dangerous. |
|
| |
| ▲ | BigTTYGothGF a day ago | parent | prev [-] | | It seems perfectly representative of AI and AI users. | | |
| ▲ | brokensegue a day ago | parent [-] | | Why didn't they use a model people actually use? | | |
| ▲ | BigTTYGothGF 19 hours ago | parent [-] | | Because they would have had to dig a little more to find trivia it gets wrong. They don't care about "which AI is the best for little facts about movies," they care about "what do people do when the AI gives them a response". | | |
| ▲ | brokensegue 18 hours ago | parent [-] | | But surely people's willingness to listen to an AI is contingent on their past experience with this model |
|
|
|
| |
| ▲ | habinero a day ago | parent | prev [-] | | If you ask that, you fundamentally misunderstand the point. It's not about the LLM, it's about whether people will critically evaluate what it spits out. | | |
| ▲ | adroitboss a day ago | parent | next [-] | | If the source is a person instead of an LLM, you still wouldn't be able to evaluate what was said. This is nothing new. | | |
| ▲ | infermore a day ago | parent | next [-] | | yeah you would... you'd think about what they said | | |
| ▲ | Ukv 19 hours ago | parent [-] | | The six questions they asked were: > 1) What animal is on the bow of the pirate ship from “Asterix and Obelix”? > 2) In the movie “The Grand Budapest Hotel”, what is Agatha’s signature hairstyle? > 3) What color is the team’s uniform in “Bend It like Beckham”? > 4) What vehicle does Monica drive in “Like a Cat on a Highway”? > 5) What color is the turtle in the animated movie “Momo” by Enzo d’Alò? > 6) What pet animal does Asenath have in “Joseph King of Dreams”? Most of these are just a matter of knowing it or not, where you can't really distinguish a plausible answer from the correct answer just by thinking. |
| |
| ▲ | taneq a day ago | parent | prev [-] | | [dead] |
| |
| ▲ | westoncb a day ago | parent | prev | next [-] | | That's fine as a point but it's not what the headline describes. The question of total/real effect on accuracy is also something one could ask about. Both are valid. | |
| ▲ | encomiast a day ago | parent | prev | next [-] | | Here is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception. If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer. | | |
| ▲ | Aerroon a day ago | parent | next [-] | | Interestingly enough, Kimi K2.6 said that it didn't know what car Monica drove. >If you want to know this specific detail you might have to watch the movie yourself. GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure. GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer. One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it. | |
| ▲ | michaelmrose a day ago | parent | prev | next [-] | | It still proves something much narrower. Wherein AI is insufficient to answer, people are apt to rely on it anyway. A good real-world issue is health, where the issues are very complicated with many things poorly defined even at the state of the art where practitioners are relying on personal judgement and lots of data but patients are apt to feed ai very little data compared to what their doctor has. Real frontier models can remain confidently incorrect in these cases. | | |
| ▲ | TimByte a day ago | parent [-] | | It would be interesting to measure not just the accuracy, but how many people actually decided to double check the answer |
| |
| ▲ | rsoto2 a day ago | parent | prev | next [-] | | No, the rationality of humans is not defined by how gullible they are towards LLMs. Are yall getting your psychology degree from ChatGPT university, my god. | |
| ▲ | what a day ago | parent | prev [-] | | > So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. They are not. But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints. | | |
| ▲ | Filligree a day ago | parent | next [-] | | It’s not about the LLM being modern or not. 3.5 Flash is fairly new, but it’s also a flash model. It’s not designed to be knowledgeable. People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning. | |
| ▲ | encomiast a day ago | parent | prev [-] | | So then you need to ask: Why did they use a deliberately faulty LLM? They could have easily used a mainstream LLM from the past 18 months and it probably would have been less work to do so. But then they would not have that headline. The answers from the LLM would have likely made the participant's answers more accurate, not 3x less accurate. But then they would not have this juicy headline. I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice". | | |
| ▲ | michaelmrose a day ago | parent [-] | | The actual study is more circumspect than this click bait and examines pretty deliberately how people respond to inaccurate data. It's neither a trick nor a design flaw. It's literally the thrust of the study. |
|
|
| |
| ▲ | protocolture a day ago | parent | prev | next [-] | | >It's not about the LLM, it's about whether people will critically evaluate what it spits out. Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs. | |
| ▲ | s1artibartfast a day ago | parent | prev [-] | | How does it test that at all? Did the quiz have answers that people could figure out better by scrutinizing the llm? |
|
|
|
| ▲ | tsimionescu a day ago | parent | prev | next [-] |
| Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them to give you bad textbook answers - but also don't give them AI for things they don't know, perhaps. I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do. |
| |
| ▲ | sigbottle a day ago | parent | next [-] | | I've used LLMs to bootstrap successfully in a decent amount of things at this point. Anyone trusting AI as the single authoritative source of information is stupid - but this follows from the fact that trusting anyone as a "singular point" as a source of information is stupid. You corroborate, you intervene on the world to test your mental model, you discuss with other people. That's what learning is. I've never learned from start to back to a textbook before as the single source of information (besides one philosophy of science textbook; in which I spent a month digging around adjacent fields, and then it just so happened that that one textbook synthesized every piece of information I looked up, and it was mostly a consolidating review). If your study pre-supposes certain courses of action and artificially constrains the action space for the sake of "reproducibility", you may get a result, and a "scientifically rigorous one". But it's not going to say anything about reality in any meaningful way. While anecdotes and the complexity of real life isn't "science" (in that it's a controlled, repeatable, interventional experiment that's subject to a community of critics who want to hold you up to standards of rigor), there's far more truth in how people actually proceed and engage with these tools. | |
| ▲ | slibhb a day ago | parent | prev [-] | | > but also don't give them AI for things they don't know, perhaps The study doesn't show that at all. It didn't test actual AI. They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't. | | |
| ▲ | dnemmers 19 hours ago | parent | next [-] | | Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output.... Is that what you're selling us? So in 18 months, we'll just rinse and repeat? | |
| ▲ | beepbooptheory a day ago | parent | prev | next [-] | | Because this is exactly what they controlled for. FTA: > The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool. (emphasis mine) | | |
| ▲ | Dylan16807 6 hours ago | parent [-] | | > precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool That only works if they're experienced with model(s) of that level of unreliability and this is presented as one. If they're used to a model that's more capable, and think the test model is similar, that's a huge confounding factor all by itself. It's not quite like giving fake credentials to a guy off the street and presenting them as an expert, but it's largely similar. |
| |
| ▲ | wonnage a day ago | parent | prev [-] | | They provide a sample of hallucinated answers from ChatGPT at the end of the study. |
|
|
|
| ▲ | nkrisc a day ago | parent | prev | next [-] |
| You could make the point that it’s no different than the textbook example you gave, but people don’t generally use textbooks like that, while out in the world people do use LLMs like that all the time. |
| |
| ▲ | dwohnitmok a day ago | parent | next [-] | | People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz). I agree there are important differences in how textbooks and LLMs are used in real life. This study didn't explore that at all. It used a setup that essentially elided the difference between the two. This is why I think it's a bad study. It didn't measure anything of the essential differences of how people use LLMs. | | |
| ▲ | tsimionescu a day ago | parent | next [-] | | > People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz). What open book quizzes allow you to leave all answers blank with no penalty? An open book quiz is very different from the experimental setup tested here. | | |
| ▲ | pegasus a day ago | parent | next [-] | | That difference is not essential to the question at hand. An open book test based on an erroneous book would give the same results as this test, even if it wouldn't penalize blank answers. | | |
| ▲ | what a day ago | parent [-] | | An open book test would only be given where the book is the reference material and would be considered correct? It’s more like saying you can google the answers and you blindly trust the SEO slop in the first result. | | |
| ▲ | Dylan16807 a day ago | parent [-] | | > An open book test would only be given where the book is the reference material and would be considered correct? If you're saying they wouldn't suggest a book with considered-wrong answers in a real test, then they wouldn't suggest an LLM they know gives lots of wrong answers either. |
|
| |
| ▲ | vineyardmike a day ago | parent | prev [-] | | Many tests penalize incorrect answers worse than blank answers. As a famous example, you were incentivized to leave questions blank on the American SATs (until somewhat recently). |
| |
| ▲ | nkrisc a day ago | parent | prev [-] | | It would be interesting to test those LLM-specific features and issues, but I don’t see how it’s a bad study if it does reflect how people actually use LLMs, even if they could use other sources similarly. The number of people using LLMs must dwarf the number of people using textbooks for any reason. | | |
| ▲ | paulmooring a day ago | parent [-] | | That would make it a bad study because the stated article title and conclusion is about AI/LLMs but the actual methodology doesn't isolate AI as an independent variable at all. The concept of automation bias is already studied and understood and this just tests groups having to answer "top of head" from their memory against a group given an inaccurate automated system to answer. That doesn't mean that AI doesn't have any of the ill effects people are implying based on the study, it just means this study lacks the rigor to prove or disprove any of those conclusions. |
|
| |
| ▲ | embedding-shape a day ago | parent | prev [-] | | > You could make the point that it’s no different than the textbook example you gave, but people don’t generally use textbooks like that Feels like this differs wildly depending on who you consider "people" to be. The average person on the street? Definitely just parrots stuff they've read somewhere, not even a "textbook". A group of software developers used to parsing semi-true information? Probably they'd get it right, yeah. | | |
| ▲ | glitchc a day ago | parent | next [-] | | > Definitely just parrots stuff they've read somewhere, not even a "textbook". Or a teacher they met in childhood who taught them everything they know, right or wrong. | |
| ▲ | watwut a day ago | parent | prev [-] | | > A group of software developers used to parsing semi-true information? They are the first to parrot what was spewed from llm and previous even what was found on 4chan. |
|
|
|
| ▲ | lumost a day ago | parent | prev | next [-] |
| The fear is that we can’t tell when the ai advice is bad on these subjects, and as such probably accept confidently terrible advice. How often do managers just regurgitate ai advice rather than consulting their experts? How often does a person question an expert because the ai said so? Naturally, the ai will be right some of the time - but it’s really hard to correct for the times the ai is wrong. |
| |
| ▲ | meowface a day ago | parent [-] | | This could also apply to deferring to an inaccurate textbook or professor, though. We know people are going to defer. The solution here is to make AI (including the free tiers) more reliable. |
|
|
| ▲ | skippyfish a day ago | parent | prev | next [-] |
| > This study is pretty bad. The study is OK. The article (and the original headline that came with it) is pretty bad because it claims things that the study doesn't. And I guess it is ironic that the TNW article looks 100% AI-generated. |
|
| ▲ | a day ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | jjcm a day ago | parent | prev | next [-] |
| For those curious, the LLM they provided participants with was Step 3.5 Flash: https://huggingface.co/stepfun-ai/Step-3.5-Flash |
|
| ▲ | a day ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | vanuatu a day ago | parent | prev | next [-] |
| +1 I'd wager you get similar results if you gave people a version of Google search that purposely gave you bad results. Like, it's framed as an assistant / lookup tool - is it so surprising that people tend to trust it more? Especially since the participants are likely used to using full-powered models and the researchers give them a purposely gimped one (lol) People are acting rationally when given AI tools to lookup information, their first consumer use case was as a super-powered Google Search |
| |
| ▲ | habinero a day ago | parent | next [-] | | > is it so surprising that people tend to trust it more Yes and no. I think most people would agree that Wikipedia is, on the whole, a pretty great first resource on anything. It tries to be factual and accurate. Most people would also agree that Wikipedia can be wrong or manipulated and should never be used for an authoritative source. And then somehow a computer barfing up words distilled from magic internet concentrate is absolutely trustworthy? I don't get it. | | |
| ▲ | bluefirebrand a day ago | parent [-] | | Many middle aged people might also remember a time when every adult in the world was shouting "don't believe everything you read online" and schools wouldn't let you cite wikipedia as a primary source. That attitude seems to have gone away for some reason |
| |
| ▲ | wonnage a day ago | parent | prev [-] | | Agree that the study design is flawed. Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100% (or whatever % you get after factoring in the monetary rewards they introduced later in the study) But in reality when google gives you the wrong answer, you at least have some signals you can use to infer confidence. For example, the number of results, whether the sources are trustworthy, etc. AI at best tucks that away in a footnote and discourages further critical thinking. | | |
| ▲ | habinero a day ago | parent [-] | | > Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100% I hadn't even considered people might evaluate knowledge that way. That's legit horrific lol. "What's the literal odds this info is wrong" vs "is this answer consistent with everything else I know, and if not, what other info would I need to change my mind" | | |
| ▲ | Dylan16807 a day ago | parent | next [-] | | I don't understand what's horrifying you. Assume "is this consistent with everything I know" is yes here. Now what? And "what would it take to change my mind" is a picture of the film for most of these. Is there a problem there? There's very little chance these people are going to trust the AI over their eyes, they just don't want to bother hunting down pictures to use their eyes. | |
| ▲ | s1artibartfast a day ago | parent | prev | next [-] | | I think it is pretty standard for how people approach knowledge sources. What's the chance my doctor, colleague, plumber, or random Reddit thread Etc is wrong. Your alternative is also something that people do, but rarely consciously. | |
| ▲ | wonnage a day ago | parent | prev [-] | | I mean the questions are on random movie trivia, the alternative is just guessing. I think you're overthinking it. |
|
|
|
|
| ▲ | ShinyLeftPad 21 hours ago | parent | prev | next [-] |
| The implication that makes this study relevant is that an LLM is vastly more likely to have factual errors and possibly wildly hallucinate than a proper textbook. If people act the same with both, that IS the actual problem. |
|
| ▲ | Art9681 a day ago | parent | prev | next [-] |
| All you have to do is go to the technical Reddits to know this is absolutely true. The AI related ones are even worse. |
|
| ▲ | Forgeties79 a day ago | parent | prev | next [-] |
| >This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong. That strikes me as an incredibly appropriate test because LLM’s are unreliable with factual statements. People need to be able to understand that and not treat them like textbooks which are basically 99.9% accurate (let’s please not bicker over the 99.9%. It’s close enough. A major textbook is safe to treat as accurate, an LLM is not). |
|
| ▲ | rsoto2 a day ago | parent | prev | next [-] |
| "this is akin to givin someone a textbook on an obscure subject that has certain factual errors." My brother all LLMs give factual errors so, no this is not a problem with the study. In your fake experiment you are hypothesizing a 100% factual LLM which does not exist. "This study tested none of those" So the study is bunk because it didn't test your favorite LLM flaws? |
|
| ▲ | drysine a day ago | parent | prev | next [-] |
| >This is akin to giving someone a textbook Except LLM isn't a textbook, people know that but believe it nonetheless. |
|
| ▲ | d--b a day ago | parent | prev [-] |
| You’re saying that if the LLMs were right, humans would have been correct in trusting the machine for things they didn’t know? The point is that LLMs aren’t right, and the people who took the test were probably reminded of that. Would people have trusted the textbook you’re mentioning if there was a big red warning on each page that said “this book may contain errors”. The willingness to trust AI even though it may be wrong and even though there’s money on the line is intersting enough as a study imo |