| ▲ | dominotw 5 hours ago | |||||||
Very few problems really have "all the tooling to verify its hypotheses" though. even if you want to construct such an harness. Also let me ask you why we need better and better and models if what we have already can produce good output with 'all the tooling to verify its hypotheses' | ||||||||
| ▲ | dragonwriter 4 hours ago | parent | next [-] | |||||||
> Also let me ask you why we need better and better and models if what we have already can produce good output with 'all the tooling to verify its hypotheses' “Good” isn’t “perfect” and even if it was, the ability to produce perfect output with all the tooling to verify its hypotheses could still be improved, in time and token efficiency, by better models producing fewer spurious hypotheses, rejecting those it does generate faster, and taking fewer unnecessary steps in confirming its good hypotheses. | ||||||||
| ||||||||
| ▲ | fn-mote 5 hours ago | parent | prev | next [-] | |||||||
> Very few problems really have "all the tooling to verify its hypotheses" This is such a blanket dismissal that I can’t agree or disagree. Maybe very few of YOUR problems are this way. At least mention some problem domains. Recent experiences: compiler-related (helpful), UI-related (agree it isn’t testable but the design iteration is quick, easy, and correct), debugging technical configuration problems (useless; I basically have to solve each problem myself before the LLM recognizes it). | ||||||||
| ||||||||
| ▲ | simonw 5 hours ago | parent | prev [-] | |||||||
We don't. You could freeze model development today and we would be able to use our current models to effectively optimize existing code for years to come. | ||||||||