| ▲ | Grok 4.6(x.ai) |
| 108 points by iLuddite an hour ago | 105 comments |
| |
|
| ▲ | nater5000 21 minutes ago | parent | next [-] |
| It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with. Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it wouldn't slow him down a bit. He's too rich to be held accountable, and that makes it impossible to trust his businesses. It's a funny dynamic that I don't think is appreciated enough, but I know that if Google or Amazon or OpenAI or Anthropic (etc.) got caught doing something like that, the backlash would be astounding and the reputation hit they'd take would be brutal. Here, Musk would just awkwardly come out attacking people for not letting him behave unethically even more than he already is, and that'd be it. Beyond that, the obvious astroturfing that occurs on this site (along with reddit, etc.) when it comes to Grok isn't helping. All I hear about Claude, GPT, Gemini, etc., are how terrible they are, yet any discussion of Grok seems to always revolve around sensible, but confident, assertions that it's actually a great product and every new release is the point where Grok finally catches up. |
| |
| ▲ | KerrAvon 5 minutes ago | parent | next [-] | | Tesla has lost both house battery and car sales in my family -- we're talking hundreds of thousands of dollars -- simply because we don't trust him not to remotely shut off our power/cars for petty political reasons. Also, if you want true privacy you should run AI models on local hardware. (Guess which country's models dominate SOTA/near SOTA open weights? Yes, it's China, and it's not even close. You can run full-fat DeepSeek locally for (just) under $10K USD.) | |
| ▲ | itsdesmond 18 minutes ago | parent | prev | next [-] | | Someone in another comment thread whataboutism’d a Chinese LLM. This isn’t a good gotcha. Musk has amplified the concept of “remigration” which is the forced deportation of non-whites. He would have me violently removed. I do not need to contextualize my decision within possible ethical quandaries. | |
| ▲ | re-thc 11 minutes ago | parent | prev | next [-] | | > It's crazy that I'd literally trust a Chinese AI company with my data It's crazy how much Chinese = bad the media or US companies have washed into you. Why lump it together? Like any place and any company there are good and bad 1s. It's not the Wild West over there... | | |
| ▲ | bellowsgulch 4 minutes ago | parent | next [-] | | [delayed] | |
| ▲ | crimsoneer 8 minutes ago | parent | prev [-] | | I mean, the Chinese government doesn't really believe in checks and balances, or corporations as autonomous to the state. That's not a conspiracy, that's just how the CCP sees it (ask Jack Ma). You could argue the US has the Cloud Act, and obviously their respect for rules based law and order as a concept has heavily deteriorated, for but it's a very different kettle of fish to a regime who just doesn't even believe in the concept. | | |
| ▲ | KerrAvon 2 minutes ago | parent | next [-] | | So have you looked at what's happened in the US over the past 10 years? The US has much further to fall, but it's falling very, very quickly and if there's ever another Democratic president they're going to have to rebuild a lot of the government from scratch. | |
| ▲ | toasty228 7 minutes ago | parent | prev [-] | | Meanwhile Trump is building a surveillance state with all his tech executives friends who all massively benefit from government sponsored schemes, it's TOTALLY different! | | |
| ▲ | crimsoneer a few seconds ago | parent [-] | | At the risk of stating the obvious, Trump has had his tariff policy killed off in the courts (although it'll obviously come back in some form) and in a few months is going to have (probably not great) midterm elections. And there are pretty open efforts to commit genocide in Xinjiang to preserve a nationalist myth of ethnic purity. So, you know, yes. |
|
|
| |
| ▲ | oulipo 16 minutes ago | parent | prev | next [-] | | Nobody wants a nazi AI | | |
| ▲ | aturek 10 minutes ago | parent | next [-] | | A number of HN commenters want the nazi AI! Which certainly makes me distrust their judgement in other domains. | |
| ▲ | neonstatic 5 minutes ago | parent | prev [-] | | And rightfully so. Unfortunately, they are perfectly fine with a marxist-leninist AI, and that's troubling. |
| |
| ▲ | agustechbro 6 minutes ago | parent | prev [-] | | Is so dumb your attitude to mix political positions with technology and tools, it will hold you back, even worst, it is dangerous for you because that means you are aboslutely sure about your ideas. What a blindly and wasteful way to live a life. |
|
|
| ▲ | causal 29 minutes ago | parent | prev | next [-] |
| Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity. Other reasons? Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually. |
| |
| ▲ | logancbrown 26 minutes ago | parent | next [-] | | Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers.
So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model. | | |
| ▲ | sm0ss117 24 minutes ago | parent | next [-] | | Yea, this reads as LLMs are a pretty obvious technology to develop(for the highly intelligent researchers who are there). Also there's probably a lot of actual divergence in model capabilities and skills that concealed by the fairly narrow set of tests we run them against nowadays. Like wasn't Grok 4.20 super targeted at non-coding tasks. | |
| ▲ | causal 18 minutes ago | parent | prev [-] | | GPUs might explain the remarkably concurrent timing. Data access doesn't really explain it unless all labs simultaneously got access to some treasure trove of data. |
| |
| ▲ | glimshe 27 minutes ago | parent | prev | next [-] | | 4) There's nothing terribly special about Anthropic. No moat. | | |
| ▲ | causal 25 minutes ago | parent [-] | | Agreed, but my suspicion is tied to the timing. Catching up eventually is to be expected. Having similar jumps in capability ready at the same time is odd. | | |
| ▲ | dash2 16 minutes ago | parent [-] | | Maybe "readiness" is quite a flexible category? You're mid-training for your next model; a rival releases something; you clear the boards and release the model without completing the training run? | | |
| ▲ | causal 12 minutes ago | parent [-] | | Touche, aborted training runs probably do happen often. Closed model providers have zero incentive to announce a new model with less-than-best benchmarks. |
|
|
| |
| ▲ | inerte 6 minutes ago | parent | prev | next [-] | | No, it has happened to almost every other "sota" model before. There used to be a meme with a circular arrow going through Anthropic, OpenAI, Google as a hype circle. Now we can drop Google and add a couple of Chinese companies. It's not an explanation of why it happens, I am just pointing Fable is not an exception, it has happened with almost every other model release by all these companies over the last 2-3 years. | |
| ▲ | extr 26 minutes ago | parent | prev | next [-] | | It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now. | | |
| ▲ | causal 22 minutes ago | parent [-] | | Does not explain timing | | |
| ▲ | extr 13 minutes ago | parent [-] | | keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later. | | |
| ▲ | causal 11 minutes ago | parent [-] | | Yeah that would make more sense, it's probably a tight community and word gets around when something starts working. |
|
|
| |
| ▲ | ayewo 17 minutes ago | parent | prev | next [-] | | > 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse. Mythos (Preview) was publicly announced in April 2026 [1] which means other labs have had 4 months to catch up, not 2 months. Assuming everyone had access to Mythos from the start, your expression, similar to other folks would have been "Mythos-level intelligence" and not "Fable-level intelligence". 1: https://news.ycombinator.com/item?id=47679258 | | |
| ▲ | causal 14 minutes ago | parent [-] | | Fair point. Still a very quick turnaround considering the other labs would have to figure out both HOW to train a Mythos-level model and then do the work (and Grok is the last to catch up), but certainly more plausible than a 2 month window. |
| |
| ▲ | moomin 22 minutes ago | parent | prev | next [-] | | Yeah, I’m not convinced that there are any models as smart as Fable. Opus 5 definitely isn’t for all it has great benchmark scores. Fable displays judgement in a way I haven’t seen from any other model. | | |
| ▲ | causal 20 minutes ago | parent [-] | | Yeah as models get better, valid benchmarks become more "trust me bro". |
| |
| ▲ | jerf 25 minutes ago | parent | prev | next [-] | | Possibility: They're all hitting the same plateau of what LLMs can do with their current architectures. I'm not stating this as a fact, but it's a hypothesis I'm keeping in my mix. | | |
| ▲ | moduspol 16 minutes ago | parent [-] | | It's possible, though I was thinking the same when GPT 5 released and it was kind of a nothing burger. Then I threw out that hypothesis with Opus 4.5. |
| |
| ▲ | Jcampuzano2 24 minutes ago | parent | prev | next [-] | | I'm pretty sure both Anthropic and OpenAI haven't necessarily been secretive that they have internal models that are much more capable than commercially available ones. It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right. | | |
| ▲ | causal 22 minutes ago | parent [-] | | The "one in the chamber" is another good candidate that could explain the timing. | | |
| ▲ | r_lee 10 minutes ago | parent [-] | | I think this is the right one, iirc 5.6 came out quite soon after Opus 5 etc? |
|
| |
| ▲ | user43928 15 minutes ago | parent | prev | next [-] | | I understand Mythos became internally available on the 24th of February. Other labs catching up in half a year seems about right. | |
| ▲ | bottlepalm 25 minutes ago | parent | prev | next [-] | | I think model level is more a function of the state of hardware. Once it exists and is available (and if a lab can afford it), then they can train their own 1T, 5T, coming up next 10T model. | | |
| ▲ | causal 23 minutes ago | parent [-] | | This is a good candidate because it would also explain the timing. Most of the replies here do nothing to explain the timing I brought up. |
| |
| ▲ | re-thc 8 minutes ago | parent | prev | next [-] | | > It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually. What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time? > Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? It means Anthropic had no real moat and no real lead. Is that weird to you? | |
| ▲ | enraged_camel 17 minutes ago | parent | prev | next [-] | | I'm solidly in the "they are benchmaxxing" camp. This became very apparent with GPT 5.6 Sol. It, too, was widely hailed to have near-Fable level intelligence. But I used it non-stop for a week and realized that they had mostly just dialed up the relentlessness meter to eleven, most likely via heavy RLHF. Last week I gave it a small-sized auth ticket to work on, then stepped away. I came back later that afternoon and found that it had worked for 3+ hours and written 25,000+ lines of code. I skimmed over the code and it looked like a small fix followed by a massive number of additional checks around it, including static analysis tooling. I gave it to another GPT 5.6 and said "check this code and see if it addresses the ticket". It looked at it and said that 98% of it was garbage and should be thrown away (its own words). I then gave it to Fable, which said it was massively over-engineered. Fable's theory was that the agent implemented the fix first, but then compacted and lost crucial context, forgot what the original task was about, and kept going. After many compaction cycles it was completely lost. Some people complain that Opus 5 stops before finishing a task. But to me, that behavior is vastly preferable to what GPT 5.6 Sol does. | | |
| ▲ | causal 9 minutes ago | parent [-] | | Yeah I found the timing on Sol especially curious since it came right on the heels of Fable. I've had mixed results with it - sometimes it seems great, other times it makes mistakes so stupid I cannot understand how it ever gets anything right. Explaining it as a difference of effort would explain both. |
| |
| ▲ | becquerel 25 minutes ago | parent | prev | next [-] | | More compute is coming online at all times. | |
| ▲ | Traubenfuchs 25 minutes ago | parent | prev [-] | | > other reasons Maybe research is sufficiently public and simple to reproduce or the next steps of how to improve things are sufficiently obvious to the smart people working on frontier AI. |
|
|
| ▲ | Jcampuzano2 28 minutes ago | parent | prev | next [-] |
| As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation making it less appealing to many. |
| |
| ▲ | hackernan9000 21 minutes ago | parent | next [-] | | Curious - what is the main issue you find polarizing with grok? | | |
| ▲ | arrosenberg 16 minutes ago | parent | next [-] | | Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great! | | |
| ▲ | leerob 9 minutes ago | parent | next [-] | | (I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy. | | | |
| ▲ | dd8601fn 15 minutes ago | parent | prev [-] | | Is that still a thing? I assumed they would have done something about it by now. | | |
| ▲ | pseudosavant 9 minutes ago | parent | next [-] | | Definitely still a thing. They just made it a paid only feature. Whereas free users used to be able to publicly ask @Grok to create these images before. So it is still going on, just not as visible now, and Elon is making sure they monetize it. Just last week they were fighting Minnesota's law that makes creating this stuff illegal. | |
| ▲ | porridgeraisin 12 minutes ago | parent | prev [-] | | Yeah, that got stopped I think. |
|
| |
| ▲ | Someone1234 13 minutes ago | parent | prev | next [-] | | I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here; and we use Chinese models (*hosted by US providers) for context. | |
| ▲ | porridgeraisin 15 minutes ago | parent | prev | next [-] | | I believe it is because of the CEO and his recent forays into politics. The model itself is great though, especially in grok build, which is a really nice harness I find myself preferring these days. | | | |
| ▲ | well_ackshually 17 minutes ago | parent | prev | next [-] | | Where do you want to start, the neonazi owner, the child porn generation, or the data centers running on illegal gas turbines polluting and choking out people ? | |
| ▲ | itsdesmond 16 minutes ago | parent | prev | next [-] | | It’s opinions are actively steered by a man who promotes the great replacement theory, white genocide, and remigration which is the mass forced deportation of non-whites. | |
| ▲ | oulipo 16 minutes ago | parent | prev | next [-] | | Nazi salutes? Harassing women with nude pics? | |
| ▲ | inference-god 17 minutes ago | parent | prev [-] | | The guy who owns it is a total fascist / psychopath ? |
| |
| ▲ | tonyhart7 25 minutes ago | parent | prev | next [-] | | more competition is always good | |
| ▲ | gmac 18 minutes ago | parent | prev [-] | | > healthy Kind of disappointed by how many people don't see any reason to boycott a model that nudified minors and makes money for a guy that does Nazi salutes. |
|
|
| ▲ | cjalmeida an hour ago | parent | prev | next [-] |
| Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription. |
| |
| ▲ | nomilk 21 minutes ago | parent | next [-] | | I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models. | |
| ▲ | jorl17 17 minutes ago | parent | prev [-] | | In my tests Grok 4.5 is definitely not Opus level. It is somewhere in between Sonnet and Opus, I'd say maybe a bit closer to Sonnet. We'll see with 4.6. | | |
|
|
| ▲ | nomilk 22 minutes ago | parent | prev | next [-] |
| Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation of the other parts or interactions between parts. No clue why. |
| |
|
| ▲ | meetpateltech an hour ago | parent | prev | next [-] |
| Cursor blog: https://cursor.com/blog/grok-4-6 |
|
| ▲ | GenerWork 29 minutes ago | parent | prev | next [-] |
| >Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so subjective that blanket statements like this seem almost useless. |
| |
| ▲ | leerob 9 minutes ago | parent [-] | | (I work on Grok) We've been working on teaching the model how to reason about great visual design principles. Obviously this is hard and somewhat subjective, but through a combination of writing down these principles (e.g. how to think about systems, not just "use this italic serif font on marketing pages"), and then creating a lot of data to pairwise compare designs/outputs, we've made a notable improvement over G4.5 and see a path to improving much further in the next model. |
|
|
| ▲ | apitman 29 minutes ago | parent | prev | next [-] |
| https://artificialanalysis.ai/models/grok-4-6 |
|
| ▲ | Pungsnigel 35 minutes ago | parent | prev | next [-] |
| Thats actually a lot more impressive than I thought. At least on paper |
| |
| ▲ | combobyte 32 minutes ago | parent [-] | | But has it hacked anybody yet? Feels like xAi is behind on the hot new benchmarking meta. | | |
| ▲ | babelfish 29 minutes ago | parent | next [-] | | Didn't need to! The harness just uploads your repository to their blob storage directly. Cheaper than asking the LLM to do it | |
| ▲ | Bluestein 31 minutes ago | parent | prev [-] | | It'd be grand if it breached SpaceX.- Or NACA.- |
|
|
|
| ▲ | amberjack 25 minutes ago | parent | prev | next [-] |
| Still not dead somehow even though they've been renting out datacenter capacity and other (seeming) problems with people leaving and so on. Quite impressive unless it's just been benchmaxxed. |
|
| ▲ | Zsfe510asG 18 minutes ago | parent | prev | next [-] |
| So did they distill Mythos in the "Macrohard" data centers? Can Grok hack now and get a free AISI commercial? |
|
| ▲ | forgottentea 36 minutes ago | parent | prev | next [-] |
| how many CSAM per second on this version |
|
| ▲ | zxilly an hour ago | parent | prev | next [-] |
| Just after DeepSeek-V4-Pro-0813 published, is this on purpose? |
| |
|
| ▲ | MWil 31 minutes ago | parent | prev | next [-] |
| Pricing pages haven't been updated yet, still advertises 4.5 |
|
| ▲ | jgbuddy 25 minutes ago | parent | prev | next [-] |
| Very impressive |
|
| ▲ | tosh 36 minutes ago | parent | prev | next [-] |
| gpt 5.6 sol and fable 5 level if the benches hold |
|
| ▲ | lostmsu 32 minutes ago | parent | prev | next [-] |
| Wow, OpenAI is now 4th after Opus 5, K3, and Grok |
|
| ▲ | oulipo 17 minutes ago | parent | prev | next [-] |
| Nobody wants a nazi AI |
|
| ▲ | sergiotapia an hour ago | parent | prev | next [-] |
| Fable level performance, faster and significantly cheaper. Wow! |
|
| ▲ | jesse_dot_id 35 minutes ago | parent | prev | next [-] |
| Nazi model looks really good on the benchmarks |
| |
| ▲ | Jeff_Brown 27 minutes ago | parent | next [-] | | Yes -- power without trust is of no use. | |
| ▲ | maelito 32 minutes ago | parent | prev | next [-] | | Yes. Won't touch xAI things because of this. | |
| ▲ | world2vec 30 minutes ago | parent | prev | next [-] | | People downvoting this comment: are they wrong? It does generate nazi stuff and CSAM, they're in the courts because of this. | | |
| ▲ | elbrian 27 minutes ago | parent | next [-] | | This thread is obviously being astroturfed by x.ai bots. I've literally never heard someone say they are excited about Musk's CSAM slop bot yet there are like 10 of them here. | |
| ▲ | sergiotapia 17 minutes ago | parent | prev [-] | | Who cares? A knife can be used to murder people, I also disagree with the UKs retarded banning of knives. As long as Grok is forwarding these lunatics to the cops why should I care? | | |
| ▲ | jesse_dot_id 10 minutes ago | parent [-] | | In an enterprise environment, I would typically set my baseline for trust in a vendor somewhere just above their CEO doing nazi salutes and wielding chainsaws on stage. |
|
| |
| ▲ | mempko 33 minutes ago | parent | prev | next [-] | | I won't use their models for this reason. Musk tinkering too much with the RL to make it sound more like him is wild. I don't care how smart or cheap the model is if it's run by Musk, I just can't use it. | |
| ▲ | 0x70run 30 minutes ago | parent | prev [-] | | don't know why you're being downvoted... "hey, let's please just focus on how good the model is, ignore the politics, okay?" a lot of these HN users are fucking idiots |
|
|
| ▲ | hit8run 31 minutes ago | parent | prev | next [-] |
| Very excited for this release. I love how based the model is. |
|
| ▲ | sawjet 20 minutes ago | parent | prev [-] |
| You may not like Elon, but you must respect him. I don't think anyone expected 6 months ago that Grok would be at the frontier and beating openAI and anthropic. Competition is good. |
| |
| ▲ | toasty228 3 minutes ago | parent | next [-] | | > You may not like Elon, but you must respect him. Your brain on grok | |
| ▲ | VCFundedGenYer 4 minutes ago | parent | prev | next [-] | | We do not celebrate pedophiles, regardless of what they do. | |
| ▲ | anukin 11 minutes ago | parent | prev | next [-] | | The last time grok made these statement, I tried using it for my workflows and it did not perform as good as opus or even sonnet. My guess is that xai benchmaxxes a lot but fails in actual capacity to produce good models. | |
| ▲ | j_maffe 18 minutes ago | parent | prev | next [-] | | Why? Did Elon design Grok? | | |
| ▲ | VariousPrograms 10 minutes ago | parent [-] | | There was that time Grok persistently brought up "white genocide" regardless of prompt, so I'd say Elon has a big personal role designing Grok's outputs! |
| |
| ▲ | dancemethis 15 minutes ago | parent | prev | next [-] | | No, we mustn't. This improvement is in 1) merit of Cursor's team and ground-level "X-Ai" AI engineers and 2) despite Elon's meddling. Just imagine how much he's trying to push internally that this new generation of Grok should be spouting his kind of propaganda. | |
| ▲ | well_ackshually 15 minutes ago | parent | prev | next [-] | | >You may not like Elon, but you must respect him. lmao no fuck him | |
| ▲ | tomashubelbauer 17 minutes ago | parent | prev [-] | | You most certainly don't need to respect Musk |
|