| ▲ | rao-v 5 hours ago | |
It’s really interesting how much little choices make the result better or worse. Astra and one of the GLMs added bright lights, and thus looked so much better to my eye. Genuinely happy with some of the Qwen 3.8 results (especially since I can run that model at Q8). Interesting to see how much better (at this task) Pi (OMP) is over Opencode as a harness. I’d love to see a few more with outcomes that are as easy to judge but less subjective. I’ve got a toy project going to make a fun to watch battle simulator where an LLM (or two if playing vs) has to write programs that control multiple bots (each with their own line of sight and limited battle context) that have to coordinate and fight alongside each other. Goal is to have the LLM update the code based on current situations maybe 5-10 times in a 5 min simulated battle. Exploring even allow the bots to request new programming and score based on number of reprogram steps. | ||
| ▲ | alvins82 5 hours ago | parent | next [-] | |
Interesting also that Astra and Sol didn't check visually using the browser. I suspect they would have done even better if they had. | ||
| ▲ | riversflow 3 hours ago | parent | prev [-] | |
huh. goes to show how much browser selection impacts things. the OMP Qwen result is completely unresponsive in free cam on my iphone, whereas the opencode version is responsive, has multitouch zoom and multitouch pan. the glm models do really well with this as well, but all of the anthropic models have pretty gimped camera controls, even astra, which while pretty wont let me pan and only allows me to rotate about 120 degrees! I’m really impressed with the Qwen 27b OpenCode result. I suppose the only thing that the prompt asks for is the cinematic view, and honestly they all kinda fail on the “subtle volumetric-style fog planes”, none of them have more fog when you get farther from a light source. | ||