Remix.run Logo
dominotw 5 hours ago

Very few problems really have "all the tooling to verify its hypotheses" though. even if you want to construct such an harness.

Also let me ask you why we need better and better and models if what we have already can produce good output with 'all the tooling to verify its hypotheses'

dragonwriter 4 hours ago | parent | next [-]

> Also let me ask you why we need better and better and models if what we have already can produce good output with 'all the tooling to verify its hypotheses'

“Good” isn’t “perfect” and even if it was, the ability to produce perfect output with all the tooling to verify its hypotheses could still be improved, in time and token efficiency, by better models producing fewer spurious hypotheses, rejecting those it does generate faster, and taking fewer unnecessary steps in confirming its good hypotheses.

an hour ago | parent [-]
[deleted]
fn-mote 5 hours ago | parent | prev | next [-]

> Very few problems really have "all the tooling to verify its hypotheses"

This is such a blanket dismissal that I can’t agree or disagree.

Maybe very few of YOUR problems are this way. At least mention some problem domains.

Recent experiences: compiler-related (helpful), UI-related (agree it isn’t testable but the design iteration is quick, easy, and correct), debugging technical configuration problems (useless; I basically have to solve each problem myself before the LLM recognizes it).

dominotw an hour ago | parent [-]

> This is such a blanket dismissal that I can’t agree or disagree.

> Maybe very few of YOUR problems are this way. At least mention some problem domains.

i did in second part of my comment. why do you think billions are being poured into ai if ai can already do verifable tasks.

simonw 5 hours ago | parent | prev [-]

We don't. You could freeze model development today and we would be able to use our current models to effectively optimize existing code for years to come.