| ▲ | achrono 3 hours ago | |
Sounds obvious but just try using the models for anything outside the evals. Take something arcane from Greek history, use it to create a masked linguistic puzzle, which you then ask the model to solve mathematically, all wrapped as an ask to generate ASCII art. Yes, all these elements exist in some form in the evals but the key is in how utterly unconventional the elements are that you pick and in how you combine them. I have consistently noticed Opus 4.8 and GPT-5.6 far outshine the Chinese models. Gemini is sort of middle of the road, Grok is better than Gemini but not really close to Opus/GPT. OAI & Anthropic still remain unbeaten by a wide margin in my eyes. | ||
| ▲ | skohan 3 hours ago | parent | next [-] | |
At that point aren't you just edge-case testing? Surely most of your use-cases are not novel tasks that combine obscure domains. It seems to me the real way to evaluate the value of a model is how it performs in your real-life workflows. | ||
| ▲ | CamperBob2 12 minutes ago | parent | prev [-] | |
That problem sounds reminiscent of one I like to use as a benchmark, which is to request that the model create an .SVG of a logarithmic spiral of 50 numbered stones. Qwen 3.8 27B absolutely knocked that one out of the park, where a lot of larger models have failed outright or otherwise performed suboptimally. Can you share an example of the Greek-history puzzle prompts you're talking about? | ||