| ▲ | dwohnitmok a day ago |
| People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz). I agree there are important differences in how textbooks and LLMs are used in real life. This study didn't explore that at all. It used a setup that essentially elided the difference between the two. This is why I think it's a bad study. It didn't measure anything of the essential differences of how people use LLMs. |
|
| ▲ | tsimionescu a day ago | parent | next [-] |
| > People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz). What open book quizzes allow you to leave all answers blank with no penalty? An open book quiz is very different from the experimental setup tested here. |
| |
| ▲ | pegasus a day ago | parent | next [-] | | That difference is not essential to the question at hand. An open book test based on an erroneous book would give the same results as this test, even if it wouldn't penalize blank answers. | | |
| ▲ | what a day ago | parent [-] | | An open book test would only be given where the book is the reference material and would be considered correct? It’s more like saying you can google the answers and you blindly trust the SEO slop in the first result. | | |
| ▲ | Dylan16807 a day ago | parent [-] | | > An open book test would only be given where the book is the reference material and would be considered correct? If you're saying they wouldn't suggest a book with considered-wrong answers in a real test, then they wouldn't suggest an LLM they know gives lots of wrong answers either. |
|
| |
| ▲ | vineyardmike a day ago | parent | prev [-] | | Many tests penalize incorrect answers worse than blank answers. As a famous example, you were incentivized to leave questions blank on the American SATs (until somewhat recently). |
|
|
| ▲ | nkrisc a day ago | parent | prev [-] |
| It would be interesting to test those LLM-specific features and issues, but I don’t see how it’s a bad study if it does reflect how people actually use LLMs, even if they could use other sources similarly. The number of people using LLMs must dwarf the number of people using textbooks for any reason. |
| |
| ▲ | paulmooring a day ago | parent [-] | | That would make it a bad study because the stated article title and conclusion is about AI/LLMs but the actual methodology doesn't isolate AI as an independent variable at all. The concept of automation bias is already studied and understood and this just tests groups having to answer "top of head" from their memory against a group given an inaccurate automated system to answer. That doesn't mean that AI doesn't have any of the ill effects people are implying based on the study, it just means this study lacks the rigor to prove or disprove any of those conclusions. |
|