Remix.run Logo
▲ ashleyn an hour ago

>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.

Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.

What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.

Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.

▲michaelchisari an hour ago | parent [-]

| few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output

That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.

▲JBits an hour ago | parent | next [-]

I would question the narrative that humans lack the capacity to verify the output and would instead argue the people lack the incentive to verify the output.

The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.

For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.

▲skydhash an hour ago | parent | prev [-]

Humans are not immortal and cannot spend all their time into review (especially unpaid). Even today, there’s so much knowledge around that you have to be specialist of a narrow domain to get to the frontier. Even in computing which is just approaching a century of existence.