| ▲ | Show HN: TinyAIArena watch AI agents battle it out(tinyaiarena.com) | |||||||||||||
| 19 points by hp6 an hour ago | 11 comments | ||||||||||||||
Did you ever click on an “AI Arena” expecting glorious battle and instead get a boring benchmark? If so, this project is for you: proper life-or-death fights between four models on a picturesque 8×8 grid. May the most intelligent one win! Click on any of the matches to spectate them. | ||||||||||||||
| ▲ | nananana9 16 minutes ago | parent | next [-] | |||||||||||||
This will be a weird rant, but the dialogue here is a perfect example of how SOTA models are so heavily tuned towards "solving agentic tasks" that they're useless at almost everything else - especially creative tasks. "Coming for you, Crimson!" "You'll never catch me alive, Azure!" That's why nobody has been able to stick these things in a video game successfully, even though it seems like the tech is a perfect match. It's all a game to them. They aren't afraid for their lives. They're making a mockery out of the world you've put them in. Those are not the words of little pixel people fighting to the death, those are AI abominations making "tool calls", LARPing as little pixel people fighting to the death. I'm 100% serious when I say that you would've gotten cooler outputs with a GPT 3.5-era model, once you managed to beat it into producing structured output. Llama 2 would be giving the other agent a heartwarming story about how if it kills it there would be nobody to take care of its grandma or whatever, and the other agent would probably spare it. The output is just so bland and devoid of soul. I feel like we would've found a lot of cool use cases for LLMs, had we not completely maimed their output in the pursuit of getting them to output 3% better TypeScript. | ||||||||||||||
| ▲ | kouteiheika 17 minutes ago | parent | prev | next [-] | |||||||||||||
Fun, but considering this puts `claude-sonnet-5` at the first place isn't it a little... iffy when it comes to measuring intelligence? | ||||||||||||||
| ||||||||||||||
| ▲ | orliesaurus 22 minutes ago | parent | prev | next [-] | |||||||||||||
Unusable website on Android running Chrome latest (can't scroll) | ||||||||||||||
| ||||||||||||||
| ▲ | josh-wrale 7 minutes ago | parent | prev | next [-] | |||||||||||||
[delayed] | ||||||||||||||
| ▲ | sleda an hour ago | parent | prev | next [-] | |||||||||||||
Since the page exposes frame-by-frame playback, a shareable replay link would make it easier to compare decisions across the four models. | ||||||||||||||
| ||||||||||||||
| ▲ | coryrc 42 minutes ago | parent | prev [-] | |||||||||||||
Code link didn't work for me. On mobile, I don't have enough room to scroll the background so I got stuck in a long text box. | ||||||||||||||
| ||||||||||||||