Remix.run Logo
▲ dofm 5 hours ago

> One is that AI will continue hallucinating in a manner that is not easy to verify

It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.

Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

▲user43928 4 hours ago | parent | next [-]

And why not?

For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.

Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.

Why would this not also be the case for mathematics?

▲catlifeonmars an hour ago | parent [-]

I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.

▲bonoboTP 30 minutes ago | parent [-]

That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.

▲user43928 23 minutes ago | parent [-]

To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.

"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.

But is running a malformed command that does not achieve the expected effect itself a hallucination?

▲antonvs 5 hours ago | parent | prev [-]

> It is an old saw at this point

An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.