Remix.run Logo
▲ lifeisloving 2 hours ago

I use models all day everyday, have unlimited access to all models. The curve is not going "verticle". I have all the workflows and meta agentic tooling, im not holding it wrong. Its bad, not everything is a 20th percentile problem.

There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.

There is however a exponential curve of slop, and an ever increasing number of people who's minds are completely captured by these things.

▲tkz1312 an hour ago | parent [-]

As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.

▲lifeisloving 39 minutes ago | parent | next [-]

Then you'd know well that 2hr of LLM code generation can easily be about 4-8hrs of review, and that review can be brutal.

I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.

The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.

▲tkz1312 28 minutes ago | parent [-]

You don't need to review proof code.

In the last 9 months or so llms have gone from just another useful proof tactic (like grind or sledgehammer or sat solvers) to being so good at writing proofs that I don't even bother to try myself anymore.

▲gr_norm an hour ago | parent | prev [-]

I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.

▲tkz1312 26 minutes ago | parent [-]

I'm pretty convinced most serious software will have some kind of proof system inside within the next few years.