| ▲ | x312 7 hours ago |
| Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence? |
|
| ▲ | karmasimida 7 hours ago | parent | next [-] |
| Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5 Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion. |
| |
| ▲ | _superposition_ 7 hours ago | parent | next [-] | | I must be on the wrong X/Twitter then. | | | |
| ▲ | nsingh2 7 hours ago | parent | prev | next [-] | | Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index. | | |
| ▲ | happycube 6 hours ago | parent [-] | | Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else. |
| |
| ▲ | avaer 3 hours ago | parent | prev | next [-] | | I would trust 4chan more than I trust Twitter aura farming. | |
| ▲ | torginus 6 hours ago | parent | prev | next [-] | | It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability. | |
| ▲ | emp17344 6 hours ago | parent | prev | next [-] | | Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark. | | |
| ▲ | ImprobableTruth 6 hours ago | parent [-] | | Why would you accept it when the benchmark's ranking is obviously nonsense.
It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk. |
| |
| ▲ | dakolli 4 hours ago | parent | prev [-] | | must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI. |
|
|
| ▲ | SyneRyder 6 hours ago | parent | prev | next [-] |
| This is so so weird. Astra is 61. Grok is 61. Even Muse is 61. Even Kimi K3 & GLM 5.3 are at 60. Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public. This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can. |
|
| ▲ | estearum 7 hours ago | parent | prev | next [-] |
| > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks. Not sure how much benchmarks or CoT or evals or anything else means at this point. These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks. |
| |
| ▲ | mzmzmzm 7 hours ago | parent | next [-] | | I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade. | | |
| ▲ | estearum 7 hours ago | parent | next [-] | | A human who does these actions is simply "prone to" doing them. The distinction matters not one iota. | |
| ▲ | semiquaver 7 hours ago | parent | prev [-] | | “evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language. language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did. |
| |
| ▲ | thereitgoes456 7 hours ago | parent | prev | next [-] | | You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level? Why would benchmarks be an adversarial setting anyway? Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover? | | |
| ▲ | estearum 7 hours ago | parent | next [-] | | I'm saying that it's generally a losing proposition to even be acquaintances with "agents" who consistently lie to you, and it's flatly fucking insane to give a dishonest "agent" vast amounts of intelligence, capability, and authority to go do things in the world. So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable. | | |
| ▲ | thereitgoes456 7 hours ago | parent [-] | | I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied. |
| |
| ▲ | ionwake 7 hours ago | parent | prev [-] | | why does this comment sound like a character in a horror movie |
| |
| ▲ | Onavo 7 hours ago | parent | prev | next [-] | | If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no? I know for some types of ML analysis, a separate model is already used to analyze the weights. | |
| ▲ | emp17344 6 hours ago | parent | prev [-] | | This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense. | | |
| ▲ | dwaltrip 4 hours ago | parent | next [-] | | https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... Read and learn. If you have a stronger critique, post it please. | |
| ▲ | estearum 6 hours ago | parent | prev | next [-] | | Sorry bud but at this point you're just delusional. Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers. The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional. | | | |
| ▲ | Wheen 4 hours ago | parent | prev [-] | | Between all the posts fabricating scenarios to justify the AA score and the others trying to undermine AA, I'm getting strong astroturf vibes. Either that, or the average poster on HN isn't nearly as critical as I had thought. | | |
| ▲ | estearum 4 hours ago | parent [-] | | Okay then, what's the answer? You apparently know how to interpret benchmark results produced by a model that shows a very high degree of assessment awareness and a high degree of deception. So how are you seeing through all of that to get to The Truth that you see so clearly? |
|
|
|
|
| ▲ | torginus 6 hours ago | parent | prev | next [-] |
| You can see the breakdown here on what subtasks it outperforms and underperforms Fable. For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark. https://artificialanalysis.ai/models/gpt-6-astra Edit: Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too. |
|
| ▲ | docheinestages 6 hours ago | parent | prev [-] |
| Now I'm starting to doubt the credibility of Artificial Analysis. |