Remix.run Logo
▲ jayd16 5 hours ago

> The LLMS are very good at logic, by the way.

It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.

▲chimprich 4 hours ago | parent [-]

I'm rather surprised to hear this. This feels like a post from about 18 months ago. I can't remember the last time I encountered a genuine code hallucination from a frontier model. They have other issues, but rarely this.

What kind of domain are you working in?

▲smrtinsert 10 minutes ago | parent | next [-]

Fintech, easy to trigger some sort of failure mode or obvious gap once specs get detailed enough and the prompt intents become specific enough. Doesn't require exhausting available context. Opus 5 did feel like a regression, will hold out judgement on 5.5 which seems much more promising.

▲weakfish 4 hours ago | parent | prev | next [-]

I see subtle ones at least daily, misunderstanding a component or hallucination of a spec for something. I’m in blockchain.

▲daveguy 41 minutes ago | parent [-]

But you should see how great it is at crud social media apps and 2D scrolling games!

▲jayd16 4 hours ago | parent | prev [-]

C++ and Unreal Engine but it has full source access. If you're actually trying to deep dive on bugs, it's still confidently wrong a lot of the time.

It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though.

I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic?

If it was so good at not making mistakes, why even have tests? It's nonsensical on its face.