| ▲ | jjcm 6 hours ago |
| Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not. Haiku 5.5: https://html.non.io/lcars-haiku-5.5/ Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5 Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f... One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work. |
|
| ▲ | thefourthchime 5 hours ago | parent | next [-] |
| Pac-Man Bench: Considering the price, no model comes close to being as good as this. However, it did take an extremely long time. TIME 19m
COST $0.16
https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5 All results:
https://jonclegg.github.io/pacman-bakeoff/ |
| |
| ▲ | myzie 5 hours ago | parent | next [-] | | Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results? | | |
| ▲ | thefourthchime 3 hours ago | parent [-] | | Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash. | | |
| ▲ | myzie 3 hours ago | parent [-] | | Yeah, I was mainly thinking about how much cheaper Luna appeared to be in this case. |
|
| |
| ▲ | rpcope1 4 hours ago | parent | prev | next [-] | | Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing? | |
| ▲ | onlyrealcuzzo 5 hours ago | parent | prev [-] | | How have you avoided being sued by Namco? | | |
| ▲ | thefourthchime 5 hours ago | parent [-] | | I'm pretty sure they'll never see this. It's pretty much impossible for anything you do you build nowadays to get noticed anyways. |
|
|
|
| ▲ | sparklingmango 5 hours ago | parent | prev | next [-] |
| > likely best used for small subagent tasks / tightly scoped work. Hasn't this always been the case with Haiku? |
|
| ▲ | saretup 5 hours ago | parent | prev | next [-] |
| To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it. |
| |
| ▲ | jjcm 5 hours ago | parent [-] | | Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down. |
|
|
| ▲ | BrokenCogs 6 hours ago | parent | prev [-] |
| Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made. |
| |
| ▲ | twostorytower 5 hours ago | parent | next [-] | | It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely. | | |
| ▲ | BrokenCogs 5 hours ago | parent [-] | | I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance. |
| |
| ▲ | jjcm 5 hours ago | parent | prev | next [-] | | Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with. It's meant to be a good test, not a good design. | |
| ▲ | FranzFerdiNaN 5 hours ago | parent | prev [-] | | It’s not really AI slop, it’s how most modern SAAS websites look like. |
|