| ▲ | andai 5 hours ago | ||||||||||||||||
Trustworthy vibe coding. Much better than the other kind! Not sure I really understand the comparisons though. They emphasize the cost savings relative to Haiku, but Haiku kinda sucks at this task, and Leanstral is worse? If you're optimizing for correctness, why would "yeah it sucks but it's 10 times cheaper" be relevant? Or am I misunderstanding something? On the promising side, Opus doesn't look great at this benchmark either — maybe we can get better than Opus results by scaling this up. I guess that's the takeaway here. | |||||||||||||||||
| ▲ | flowerbreeze 4 hours ago | parent | next [-] | ||||||||||||||||
They haven't made the chart very clear, but it seems it has configurable passes and at 2 passes it's better than Haiku and Sonnet and at 16 passes starts closing in on Opus although it's not quite there, while consistently being less expensive than Sonnet. | |||||||||||||||||
| |||||||||||||||||
| ▲ | DrewADesign 4 hours ago | parent | prev [-] | ||||||||||||||||
It’s really not hard — just explicitly ask for trustworthy outputs only in your prompt, and Bob’s your uncle. | |||||||||||||||||
| |||||||||||||||||