| ▲ | port3000 19 hours ago |
| It's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything' The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words) |
|
| ▲ | rockinghigh 19 hours ago | parent | next [-] |
| Open-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind. |
| |
| ▲ | anthonypasq 16 hours ago | parent | next [-] | | Kimi K3 is still worse than Fable and Fable was trained >4 months ago. | | |
| ▲ | janalsncm 13 hours ago | parent | next [-] | | Why are you using the product release date for Kimi K3 and the training date for Fable? Either use the release date for both (6 weeks apart) or if you have it the training date for both. | |
| ▲ | sdfefcxv 16 hours ago | parent | prev [-] | | To say X is perfectly bad vs Y is false. People use these models for diff things. Its quite possible for the things they are used for, people do not see much of a difference. Do you hold stock in Anthropic? | | |
| ▲ | TheMrZZ 15 hours ago | parent | next [-] | | > Profile created 3 days ago > Unnecessarily aggressive > First ever comment said "Further releases of Chinese models that demonstrate the gap is not growing substantially is a huge problem. The spending will be called into question." Yeah I think you have an agenda | |
| ▲ | tracker1 15 hours ago | parent | prev | next [-] | | Are you a Chinese national, or otherwise paid by China? | | | |
| ▲ | anthonypasq 15 hours ago | parent | prev [-] | | ah yes, because i said something factually accurate and vaguely positive about Anthropic I must be a shareholder which would mean I either run a venture capital firm or am a current employee of Anthropic... I wish. |
|
| |
| ▲ | yogthos 17 hours ago | parent | prev [-] | | And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time. | | |
| ▲ | janalsncm 13 hours ago | parent | next [-] | | Probably a little of both. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions. Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well. | | |
| ▲ | yogthos 10 hours ago | parent [-] | | That's my thinking as well. The whole distillation thing is a distraction from the actual innovation happening in this space. What will be interesting to see going forward is what types of new techniques people manage to come up with to over come the current architecture limits. |
| |
| ▲ | boc 16 hours ago | parent | prev | next [-] | | Or option three is they are drafting hard off the frontier US models via distillation. | | |
| ▲ | yogthos 15 hours ago | parent [-] | | The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. On top of that, Kimi also does better than Fable or GPT on a lot of tasks, distillation alone can't explain that, meaning there is a difference in architecture. You can watch a talk from Kimi founder to see how Kimi was actually trained and why it performs well. https://www.youtube.com/watch?v=5CkCW1P-g88 Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope. | | |
| ▲ | boc 12 hours ago | parent [-] | | > US companies models constantly distill each other as Musk was forced to admit under oath > This whole narrative has just been a massive cope. So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do it, but they do it anyway because it's fun? Either it's powerful and everyone, including the Chinese labs, use it as a way to rapidly catch-up against the SOTA models, or it's a red herring and the huge amounts of energy spent to protect and enable distillation is all just wasted money. Which is it? | | |
| ▲ | yogthos 12 hours ago | parent [-] | | I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow. | | |
| ▲ | boc 10 hours ago | parent [-] | | The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct. Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work. | | |
| ▲ | yogthos 10 hours ago | parent [-] | | Indeed we will, my prediction is that we'll see a model from China that definitively surpasses any US model by the end of the year. China has an absolute population advantage here along with having a much better education system. And China now dominates in published AI papers. The US enjoyed an early advantage due to excessive money being poured into AI which led to the current bubble, and access to the hardware that was needed to train these models initially. At this point, neither of these factors actually matter that much. The naive approach of simply making models bigger has hit a wall, and now you need ingenuity in figuring out better architecture for them. Precisely because Chinese companies have had to deal with more limited resources, they put a lot more effort into researching different kinds of optimizing techniques. And of course, China is also catching up in chip making, and Huawei clusters are already competitive with Nvidia for training. So, that gap is closing as well. The big difference is that an absolutely insane amount of money has been spent in the US, while China managed to do this on a fraction of the budget. The AI Investment Surge graph here puts things in perspective. https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa... |
|
|
|
|
| |
| ▲ | sdfefcxv 16 hours ago | parent | prev [-] | | This happened ages ago. But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now. | | |
| ▲ | yogthos 15 hours ago | parent [-] | | My prediction is that they're going to angle to become a vendor of record for the government and get bailed out. That's the only path at this point because there won't be any competition from China in this niche. |
|
|
|
|
| ▲ | stingraycharles 11 hours ago | parent | prev | next [-] |
| Fable is still the same model, it’s still a great model, and to be honest all these articles writing and speculating on how the LLM industry is going to evolve are not that insightful nor interesting. I don’t think one should pay much attention to them. |
|
| ▲ | jesse_dot_id 15 hours ago | parent | prev | next [-] |
| The plateau is inevitable because their rapacious training methodologies are only viable when there are no defense in place, but information continues to evolve, which means the models will have to be continuously updated, but will be doing so with less and less freely available data. |
| |
| ▲ | achandra03 12 hours ago | parent | next [-] | | > with less and less freely available data My understanding is that the labs ran out of freely available data to train on a while ago, and now primarily rely on human data vendors such as Surge and Mercor to source their data. | |
| ▲ | t0mpr1c3 12 hours ago | parent | prev [-] | | We are only just starting to model the physical world. There will be more training on empirical data. |
|
|
| ▲ | potsandpans 8 hours ago | parent | prev [-] |
| Almost like the company has no credibility with respect to its safety claims. |