| ▲ | jstummbillig 2 hours ago | |||||||||||||
I am trying to estimate if my reaction to seeing GPT-5.6 Sol last on that list is reasonable or or mostly emotional and find that I have no way of telling. | ||||||||||||||
| ▲ | WD-42 an hour ago | parent | next [-] | |||||||||||||
Why would you get emotional over a model? They got you that good? | ||||||||||||||
| ▲ | CompoundEyes an hour ago | parent | prev | next [-] | |||||||||||||
I do think it’s the wizard not the wand at this point given a decent model. These benchmarks don’t have the wizard. Otherwise I wouldn’t see others in the exact same codebase struggle and underutilize agents while others thrive using the exact same ones. | ||||||||||||||
| ||||||||||||||
| ▲ | didgeoridoo an hour ago | parent | prev [-] | |||||||||||||
Sol failing mostly on “unverified assumptions” and rarely hitting “integration errors” seems about right to me. I think Sol is second only to Astra (and miles ahead of even Fable) in architecting & engineering the right implementation — but only if you are extremely specific and provide tight guidelines and guardrails. If you give it a one-liner… you’re going to have a bad (SHA-256-hash-verified) time. | ||||||||||||||
| ||||||||||||||