Remix.run Logo
slibhb a day ago

> but also don't give them AI for things they don't know, perhaps

The study doesn't show that at all. It didn't test actual AI.

They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.

dnemmers 19 hours ago | parent | next [-]

Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

Is that what you're selling us?

So in 18 months, we'll just rinse and repeat?

beepbooptheory a day ago | parent | prev | next [-]

Because this is exactly what they controlled for. FTA:

> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

(emphasis mine)

Dylan16807 6 hours ago | parent [-]

> precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool

That only works if they're experienced with model(s) of that level of unreliability and this is presented as one.

If they're used to a model that's more capable, and think the test model is similar, that's a huge confounding factor all by itself. It's not quite like giving fake credentials to a guy off the street and presenting them as an expert, but it's largely similar.

wonnage a day ago | parent | prev [-]

They provide a sample of hallucinated answers from ChatGPT at the end of the study.