| ▲ | ianjbutler 2 hours ago | |
> the model will perform worse Depends on whether and how you want to rank stability in terms of better/worse. Models are diverging on this, which seems increasingly clear.. i.e. Fable isn't stable, but Opus isn't clever, and they hit different kinds of walls. So both the theory (diluting the correctness reward) and the practice (hard split on plan/implement/review work) seems to be pointing towards a strongly multi-model and highly agentic / harness-driven / complex-system kind of future instead of singleton monolithic super-smart models. The do-everything model with solid reasoning AND solid results, and the honest/introspective helpful agent that doesn't actively resist governance may be at odds. Stable reasoning doesn't matter for pen-testing, and correct-answer with broken processes and fragile abstractions won't matter for math/science/coding. | ||