Remix.run Logo
fxwin a day ago

You can read the study here: https://osf.io/preprints/psyarxiv/5y6m4_v1

TLDR: They studied both cases (Access to a LLM/Chat interface which gives a wrong answer when asked, and access to pregenerated (wrong) answers). Both experiments yielded similar results.

tsimionescu a day ago | parent [-]

That is not a good summary. In all their versions of the experiment, they presented the tool as AI, and used actual LLM answers. In the first example, participants were directly interacting with a real, though small, LLM, but they had technical issues because of that - 10% of the time the LLM setup failed to present an answer at all. So they redid the experiment with pre-generated answers - the participants saw the same UI, but when they asked the question of the AI, they instead got one of 3 pre-generated answers from that AI, to avoid the technical issues.

fxwin a day ago | parent [-]

im not sure which part of my summary you take issue with then? the reason for study 1b is not relevant to answer the question asked in the comment i replied to

tsimionescu a day ago | parent [-]

It is relevant, because of (1) framing (people thought they were getting answers from AI, not a Google search, and any associated reputation works in that way); and (2) sycophancy and other similar characteristics of the answer's text, which were present in the answers presented in 1b and wouldn't be in a Google search.

Basically, the difference between 1a and 1b is not at all relevant to the question of whether the observed behavior is caused by AI or simply by faulty tools. The difference between 1a and 1b was designed specifically to be transparent to the actual test takers, and only to eliminate some confounding variable (technical issues in 1a).