Remix.run Logo
tsimionescu a day ago

It is relevant, because of (1) framing (people thought they were getting answers from AI, not a Google search, and any associated reputation works in that way); and (2) sycophancy and other similar characteristics of the answer's text, which were present in the answers presented in 1b and wouldn't be in a Google search.

Basically, the difference between 1a and 1b is not at all relevant to the question of whether the observed behavior is caused by AI or simply by faulty tools. The difference between 1a and 1b was designed specifically to be transparent to the actual test takers, and only to eliminate some confounding variable (technical issues in 1a).