Remix.run Logo
tsimionescu a day ago

Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them to give you bad textbook answers - but also don't give them AI for things they don't know, perhaps.

I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do.

sigbottle a day ago | parent | next [-]

I've used LLMs to bootstrap successfully in a decent amount of things at this point.

Anyone trusting AI as the single authoritative source of information is stupid - but this follows from the fact that trusting anyone as a "singular point" as a source of information is stupid. You corroborate, you intervene on the world to test your mental model, you discuss with other people. That's what learning is. I've never learned from start to back to a textbook before as the single source of information (besides one philosophy of science textbook; in which I spent a month digging around adjacent fields, and then it just so happened that that one textbook synthesized every piece of information I looked up, and it was mostly a consolidating review).

If your study pre-supposes certain courses of action and artificially constrains the action space for the sake of "reproducibility", you may get a result, and a "scientifically rigorous one". But it's not going to say anything about reality in any meaningful way. While anecdotes and the complexity of real life isn't "science" (in that it's a controlled, repeatable, interventional experiment that's subject to a community of critics who want to hold you up to standards of rigor), there's far more truth in how people actually proceed and engage with these tools.

slibhb a day ago | parent | prev [-]

> but also don't give them AI for things they don't know, perhaps

The study doesn't show that at all. It didn't test actual AI.

They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.

dnemmers 19 hours ago | parent | next [-]

Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

Is that what you're selling us?

So in 18 months, we'll just rinse and repeat?

beepbooptheory a day ago | parent | prev | next [-]

Because this is exactly what they controlled for. FTA:

> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

(emphasis mine)

Dylan16807 6 hours ago | parent [-]

> precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool

That only works if they're experienced with model(s) of that level of unreliability and this is presented as one.

If they're used to a model that's more capable, and think the test model is similar, that's a huge confounding factor all by itself. It's not quite like giving fake credentials to a guy off the street and presenting them as an expert, but it's largely similar.

wonnage a day ago | parent | prev [-]

They provide a sample of hallucinated answers from ChatGPT at the end of the study.