| ▲ | aliljet 3 hours ago |
| This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic. |
|
| ▲ | MangoCoffee 3 hours ago | parent | next [-] |
| OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize. These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin. I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast. |
| |
| ▲ | Gigachad an hour ago | parent | next [-] | | This is going to be catastrophic. Whether AI works or use useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations. | | |
| ▲ | goolz 37 minutes ago | parent [-] | | I have already begun winding down my spend on claude and OAI to make room for infra budget. Anecdotal, but I have no doubt a lot of others are doing the same, I very much agree the US players have major issues looming. What an exciting time to be alive! | | |
| ▲ | csomar 6 minutes ago | parent | next [-] | | The car industry is also a trillion $$ market in the US. I don't see why that would go any differently from the Chinese cars ban. | |
| ▲ | netdevphoenix 33 minutes ago | parent | prev [-] | | Not exciting for anyone directly or indirectly invested in a frontier lab or its partners. And that is a lot of people, including you. | | |
| ▲ | gruturo 23 minutes ago | parent [-] | | No time like the present to pull out and reduce your exposure. I brought this up in my employer's forums 4 months ago and honestly it's been clear even before then. In particular, the upcoming IPOs of both oAI and Anthropic will likely be disastrous for the public - the floor is falling from under them and I don't know if they can be scrappy and work with fewer resources - their internal culture may not support this. We all knew in our hearts they're a commodity - just see how easily you can switch between the 2 of them - and now there are 10 more options costing a fraction. When Xi Jinping did the announcement of their open weights push, they might as well cancelled their IPOs.... |
|
|
| |
| ▲ | ilaksh 8 minutes ago | parent | prev | next [-] | | Most are not necessarily free to host and monetize. At least one of them has a license that says if you are re-hosting the model then you need a license with that company that made the model. | |
| ▲ | somenameforme 2 hours ago | parent | prev | next [-] | | Another interesting potential market here will be 'LLM in a box'. All the hardware and other tooling in a prebuilt, but modular, package ready to go. Pay one up-front cost, get a system running [whatever open LLM] with a token rate of [x], optionally configured to be immediately ready for distributed usage. Basically the opposite of cloud stuff: no rent, no dependency, 100% guaranteed uptime, guaranteed security/privacy (at least subject to your own actions), and so on. | | |
| ▲ | adrian_b an hour ago | parent | next [-] | | Palantir already offers a "turnkey AI datacenter", i.e. a rack with "NVIDIA Blackwell Ultra systems with eight NVIDIA Blackwell Ultra GPUs and NVIDIA Spectrum-X™ Ethernet networking for AI training and inference". It is said that it comes with all hardware and software required to run inference or training with an open weights LLM. The existence of this product, which competes with cloud-based offerings like those of OpenAI and Anthropic, is presumably the reason why the Palantir CEO criticized very harshly some time ago the business model of OpenAI/Anthropic. While I doubt that the ethics of Palantir is any better than of OpenAI/Anthropic, in this particular case I have to agree with Alex Karp about "Sovereign AI", i.e. that only losers will make their business completely dependent on an external entity like OpenAI or Anthropic, who are certainly not trustworthy. | | |
| ▲ | vrganj an hour ago | parent [-] | | I'm not sure a data center run by ... Palantir of all organizations is what people have in mind when they worry about data sovereignty. | | |
| ▲ | adrian_b an hour ago | parent [-] | | They are selling it, not running it. It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers. I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component. |
|
| |
| ▲ | jurgenburgen an hour ago | parent | prev | next [-] | | > 100% guaranteed uptime Disagree there but I think this is an interesting idea. We would need to find some more cost-efficient hardware to run it on than Nvidia GPUs. | | |
| ▲ | pulse7 an hour ago | parent [-] | | It will come... all big hardware players (Intel, AMD, Broadcom) and dozens of startups (Tenstorrent, etc.) are working on it... |
| |
| ▲ | bevekspldnw an hour ago | parent | prev | next [-] | | “100% guaranteed downtime when you least can afford it and the support tickets are your problem.” We’ve a hybrid shop, including hosting our own ML infra, and we save a ton from cloud spend with local ML. Easily one million USD over past three years. But it’s not “free”, you are shifting a lot of labor into your plate. | | |
| ▲ | hypfer an hour ago | parent [-] | | And with that also gain institutional knowledge, skill up your workers and attract talent that wants to work on this stuff. All boils down to short-term/long-term thinking. | | |
| ▲ | gruturo 21 minutes ago | parent [-] | | This. People WANT to work on this stuff. And having skilled workers is a precious advantage. |
|
| |
| ▲ | Godsend69 an hour ago | parent | prev [-] | | [dead] |
| |
| ▲ | miohtama 18 minutes ago | parent | prev | next [-] | | There could be soon AI safety regulations that will stop the US to host or use the Chinese models. | |
| ▲ | 0xpgm an hour ago | parent | prev | next [-] | | US investors are desperate for the next hypergrowth opportunity. From what I can tell the US economic strategy is to outgrow its debt. | |
| ▲ | nkmnz 2 hours ago | parent | prev | next [-] | | I think at this point the question is: will the US government be willing and capable to justify the trillion dollar valuation for _one_ of the companies via regulatory capture? The US has a workforce of 170m, so 1.7 trillion would come down to 10k per person, or a discounted cashflow at 3% of 25 USD per month - not including private use, students etc. | | |
| ▲ | kaashif 2 hours ago | parent [-] | | Why would you restrict to the US workforce? ChatGPT has a billion users. | | |
| |
| ▲ | chrismsimpson 2 hours ago | parent | prev | next [-] | | > Providers can just run them, offer cheap tokens, and pocket the margin. There’s an assumption that you can spin up the infra and acquire customers within that margin | | |
| ▲ | KeplerBoy 2 hours ago | parent [-] | | Which is not unreasonable. Just hosting it in the EU and promising not to retain / sell the data let's you charge a healthy extra and compete in many areas other players can't. | | |
| |
| ▲ | grey-area 2 hours ago | parent | prev | next [-] | | It is impossible to justify the absurd private valuations they have given themselves in collusion with investors. I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI? | |
| ▲ | me551ah an hour ago | parent | prev | next [-] | | I think that explains the race for IPO by the US AI labs, they know that the longer they wait, the less they will be worth. | |
| ▲ | piokoch an hour ago | parent | prev | next [-] | | "I just don't see how you justify a trillion valuation for US AI" - military applications
- financial applications
- medical
- applied science In all those cases it is achievable for those who have needed training data, and Chinese are not going to get them easily. US AI Labs are showing: give us the data, we will do wonders, promising "singularity"-level future achievements. | |
| ▲ | hmmidontknow 2 hours ago | parent | prev | next [-] | | Hmm.. how you justify? Provoking war, this is how the empire "defends" itself, usually. I just hope that this time it will get stuck in your throat. | |
| ▲ | sidd_sarkar an hour ago | parent | prev | next [-] | | Ok | |
| ▲ | charcircuit 2 hours ago | parent | prev [-] | | I suggest you think why OpenAI was worth billions before ChatGPT. The valuation is not about how the current set of models can be monetized. |
|
|
| ▲ | wren6991 30 minutes ago | parent | prev | next [-] |
| The thing that blows me away is it does this at one quarter the total parameter count of K3 (and 40% active parameter count). There's plenty of room at the bottom. > How are you all toying with running this kind of thing in a mega quantized way locally? Sure, let me answer that in excessive detail. I briefly tried running the UD IQ3_S quant of GLM-5.2, which is 288 GiB of weights (301 GB). Setup was: llama.cpp, 1x NVMe SSD (Evo 980), 64 GiB DDR5-5200, i9-13900HX, and 1x RTX Pro 6000. Token generation around 0.7 t/s. Not remotely usable interactively, but something I could plausibly push a codebase into and come back to a review in a couple of days. There's potential for that hardware to go much faster, but current local inference backends make poor use of the memory hierarchy. Ideally I would have: always-active weights, KV and hot expert cache in VRAM; warm expert victim cache in host RAM; and disk as a last resort. Instead it's 1/3rd of the layers fully pinned in VRAM (all experts), and 2/3rds running wholly on the CPU with mmap()'d weights. The CPU cores spend most of their time sleeping on disk fills. llama.cpp has backed itself into a bit of a corner architecturally by trying to support all models on all possible backends. If you look into how their "MoE offload" feature works (not viable for me because it requires enough host RAM to permanently pin the weights) you very quickly realise it's "oops, all bubbles!" due to the static compute graph splits. There are more focused frameworks like DS4 [1] and Colibri [2] which have better support for streaming weights from disk, and support GLM-5.2. Obviously I wouldn't recommend my setup for huge models like GLM-5.2. Supposedly it can just about be squeezed into 3x GB10, or run comfortably on 4x GB10 (tensor-parallel) for multi-user serving. I'm not sure whether that qualifies as local, but it's at least not a rack. [1] https://github.com/antirez/ds4 [2] https://github.com/JustVugg/colibri |
|
| ▲ | kouteiheika 3 hours ago | parent | prev | next [-] |
| > This is absolutely still shy of Sol and Fable Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway. |
| |
| ▲ | simonjgreen 3 hours ago | parent | next [-] | | We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried? | | |
| ▲ | alightsoul 2 hours ago | parent | next [-] | | You must be a 5000 person company with an existing enterprise contract to get approved that fast. That sounds like a 15 minute SLA agreement. Individuals no matter how qualified about cybersecurity, are ghosted | | |
| ▲ | xx_ns an hour ago | parent [-] | | That's not my experience at all. I was approved fairly fast - around an hour from submitting the form and getting a response. However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe. | | |
| ▲ | captn3m0 39 minutes ago | parent | next [-] | | I am guessing you are approved for the Cyber Verification Program. I also applied and got approved in an hour (on a Saturday!), but it only applies to Opus and Sonnet: https://support.claude.com/en/articles/14604842-real-time-cy.... It let me use Opus for cybersecurity work, pretty much everything except for Ransomware development. It would occasionally still trip and start saying no till I added a note about CVP in my claude.md. No one gets to use Fable for Cybersecurity work, and Mythos is not available under CVP. Only for select few customers, and there isn't an application form? | |
| ▲ | hypfer an hour ago | parent | prev [-] | | Cyberchili. Might burn holes into corporate firewalls |
|
| |
| ▲ | kouteiheika 3 hours ago | parent | prev | next [-] | | Have you tried to use Fable for anything even remotely security related, when the refusals kick in as soon as you even fart in the vague direction of anything security or biology-adjacent? | | |
| ▲ | b112 3 hours ago | parent [-] | | For this comment to have value, you should indicate whether or not you applied for cybersecurity approval, and were approved or not. |
| |
| ▲ | 112233 31 minutes ago | parent | prev | next [-] | | Why should I apply for *cybersecurity* approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, "programming") somehow is cybersecurity now? | |
| ▲ | grey-area 2 hours ago | parent | prev [-] | | Are there any limitations on this version? |
| |
| ▲ | bpodgursky 3 hours ago | parent | prev [-] | | I don't understand all this spite about "rich friends" when it was the US government that shut Fable down for not adequately blocking cyber capabilities. I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it. | | |
| ▲ | deepllm 3 hours ago | parent | next [-] | | "Mythos" is the cyber-security equivalent of Fable (without guardrails), and only a very select few corporations have access to it. Fable is their version with guardrails on everything except "Make me a pelican svg" or "create a to-do" app, that is the version that the government banned | | |
| ▲ | bpodgursky 3 hours ago | parent [-] | | I know all this? Only a few corporations have Mythos because the US government is whitelisting them one at a time. Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried. | | |
| ▲ | deepllm 3 hours ago | parent [-] | | Before the US government had anything to do with this, Anthropic were fear mongering Mythos (BTW, Amodei also fear-mongered GPT-2, so this is a normal pattern in their operation) calling it "too dangerous to release", and back then only Anthropic was in charge of the whitelist. Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted. | | |
| ▲ | bpodgursky 3 hours ago | parent [-] | | Sorry but if you stepped back for a moment you'd realize this is all contrived nonsense to let to have your cake and eat it too. No, Anthropic did not mind-game the US government into being worried about cybersecurity. The NSA has been paranoid about cyber controls for longer than you've been alive. If Anthropic had come out of the gate saying "no don't worry man, our model is TOTALLY COOL", while simultaneously attacking HAWK and finding core Linux vulnerabilities, I assure you the US government would have caught up about ten minutes later and we'd be in exactly the same spot minus your ability to tell Anthropic they were wearing the wrong dress and asking for it. | | |
| ▲ | deepllm 2 hours ago | parent [-] | | Mythos isn't some scary dangerous model that can find high severity bugs seamlessly, that's just Anthropic marketing. Most of the vulnerabilities they found were low severity hyped up to make their model look good, with (I think, maybe?) the exception of a few. Now that Chinese open weight models have similar capabilities, and their guardrails can also just be removed, it doesn't look like anyone has "hacked" into everything because of the scary dangerous models like Anthropic were making it out to be. | | |
| ▲ | d1sxeyes an hour ago | parent | next [-] | | In principle I agree but in practice I don’t. The majority of high severity vulnerabilities are not the kind of thing you need a PhD in Comp Sci to comprehend, they are mostly about finding a way to get a system to end up in a state different than was anticipated when entering a particular code path. Exhaustively looking at code and identifying ways to do this is something LLMs are quite good at. They don’t get tired, and you can run them non-stop. They're also (generally) quite good at reading the literal meaning of the code, whereas humans often see the intended meaning first, and can be biased. If you had a tireless junior engineer who was given the job of “make this application get into a state it’s not supposed to be in”, you’d probably get similar results. What Mythos is quite good at is both the first bit and coming up with ways it could chain that together with other bits of unexpected state to create something that forms a meaningful vulnerability rather than a dead end. | |
| ▲ | wren6991 20 minutes ago | parent | prev [-] | | It's also quite hard to separate Mythos the model from Mythos the campaign (aka Glasswing). They put an enormous amount of compute into bug hunting, and they found some bugs. Fair enough. For me that begs the question: what if they had spent the same compute on generating more tokens with a less-capable model? What if they had spent it on traditional fuzzing? |
|
|
|
|
| |
| ▲ | kouteiheika 3 hours ago | parent | prev | next [-] | | > I don't understand all this spite about "rich friends" Okay, here's a challenge: I assume you're not a rich and powerful entity, so try to gain access to Mythos. I'll wait. > I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it. Well, first I'd suggest they stop with the constant fear mongering. Here's my prediction for what will happen: the Chinese models will catch up to Fable/Mythos. They will be fully unrestricted and everyone will have access. The world will not end. Good guys will use them to harden their systems, in equilibrium to what bad guys have access to, so effectively status quo will not change. | | |
| ▲ | bpodgursky 3 hours ago | parent [-] | | This is a lot of words to say "you're right, Anthropic does not have any legal way to release frontier cyber capabilities to the public" | | |
| ▲ | kouteiheika 3 hours ago | parent [-] | | Right, so according to you it's because of the US government that they don't release it to the public? Have you missed their constant and incessant fear mongering? The causality chain here was not "US government says its dangerous -> Anthropic can't release it", it was "Anthropic is fear mongering -> US government listens to their fear mongering". |
|
| |
| ▲ | stavros 2 hours ago | parent | prev [-] | | The issue is that these companies keep trying to pull the ladder up behind them by going "oh my god our models are so dangerous only we should be allowed to develop them". Sometimes it backfires, but the companies aren't innocent. |
|
|
|
| ▲ | arcanemachiner 2 hours ago | parent | prev | next [-] |
| > this is just GLM 5.2 with post-training magic Isn't post-training turning out to be the most important part? |
|
| ▲ | andxor 44 minutes ago | parent | prev | next [-] |
| Fable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet. |
|
| ▲ | deepllm 3 hours ago | parent | prev | next [-] |
| Realistically, you're looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it's just better to run DSv4 flash at full precision. 4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama |
| |
| ▲ | teruakohatu 3 hours ago | parent | next [-] | | How fast are 2x or 4x DGX? I only have one and am wondering what the benefits are of getting another. I feel I will be disappointed… | | | |
| ▲ | disiplus 3 hours ago | parent | prev [-] | | i run flash v4 at 2bit, its pretty great and on my tests against full model It didn't lose any capabilities. It just was thinking more. So you don't have the same efficiency. |
|
|
| ▲ | bertili 3 hours ago | parent | prev | next [-] |
| DwarfStar (https://github.com/antirez/ds4) supports GLM 5.2 and DeepSeek. Not only for toying, but for getting work done. |
|
| ▲ | teravor 3 hours ago | parent | prev | next [-] |
| the difference is that with open models jailbreaking is trivial if you know what you are doing so this makes a frontier open model infinitely more useful for certain tasks seeing as closed frontier models will just refuse (and jailbreaking them is a waste of time when you have good open models). in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition. |
|
| ▲ | bossyTeacher 3 hours ago | parent | prev [-] |
| > This is absolutely still shy of Sol and Fable, but only just by a hair. Even if there was a small/medium gap, the fact that this is a free model beats both of the above on pure economics. |