Remix.run Logo
encomiast a day ago

Here is what the study says:

"The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."

So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.

If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.

Aerroon a day ago | parent | next [-]

Interestingly enough, Kimi K2.6 said that it didn't know what car Monica drove.

>If you want to know this specific detail you might have to watch the movie yourself.

GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure.

GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer.

One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it.

michaelmrose a day ago | parent | prev | next [-]

It still proves something much narrower. Wherein AI is insufficient to answer, people are apt to rely on it anyway.

A good real-world issue is health, where the issues are very complicated with many things poorly defined even at the state of the art where practitioners are relying on personal judgement and lots of data but patients are apt to feed ai very little data compared to what their doctor has.

Real frontier models can remain confidently incorrect in these cases.

TimByte a day ago | parent [-]

It would be interesting to measure not just the accuracy, but how many people actually decided to double check the answer

rsoto2 a day ago | parent | prev | next [-]

No, the rationality of humans is not defined by how gullible they are towards LLMs. Are yall getting your psychology degree from ChatGPT university, my god.

what a day ago | parent | prev [-]

> So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

They are not.

But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.

Filligree a day ago | parent | next [-]

It’s not about the LLM being modern or not. 3.5 Flash is fairly new, but it’s also a flash model. It’s not designed to be knowledgeable.

People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.

encomiast a day ago | parent | prev [-]

So then you need to ask: Why did they use a deliberately faulty LLM? They could have easily used a mainstream LLM from the past 18 months and it probably would have been less work to do so. But then they would not have that headline. The answers from the LLM would have likely made the participant's answers more accurate, not 3x less accurate. But then they would not have this juicy headline.

I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice".

michaelmrose a day ago | parent [-]

The actual study is more circumspect than this click bait and examines pretty deliberately how people respond to inaccurate data. It's neither a trick nor a design flaw. It's literally the thrust of the study.