| ▲ | WithinReason 3 hours ago |
| Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable? |
|
| ▲ | Aurornis 2 hours ago | parent | next [-] |
| Claims about Ox Alpha performing at Fable level were from the social media hype cycle. Everything new in the LLM space brings a wave of influencers hyping it up as a revolutionary leap forward. Don’t forget to like and subscribe to learn more. It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney. |
| |
| ▲ | worldsavior an hour ago | parent [-] | | This influencers are getting paid, it's not coincidential. | | |
| ▲ | nixon_why69 20 minutes ago | parent [-] | | They don't have to be getting paid. The natural bias of media is towards laziness and sensationalism (stolen from Jon Stewart, so maybe the same is true about comment sections). |
|
|
|
| ▲ | woadwarrior01 3 hours ago | parent | prev | next [-] |
| That benchmark is super sus. Until someone pointed it out, the top performing open weights model was a Kimi K3 fine tune from their sponsor (abacusai/Smaug-Agentic). Now, it's not on the list. Source: https://twitterwebviewer.com/?tweet=2091116504787935350 |
|
| ▲ | sunbum 3 hours ago | parent | prev | next [-] |
| the 2nd website is not official, just something someone slopped together for some reason. |
| |
| ▲ | Alifatisk 3 hours ago | parent | next [-] | | I have plenty of these websites, I can’t understand why someone is doing this. | | |
| ▲ | colesantiago 2 hours ago | parent [-] | | It is called phishing and grifting. Many people and even software engineers fall for this all the time. Most of these people are from crypto pivoting to AI doing this. AI has made this easier and cheaper and it is going to get a LOT worse. Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds. The public have no chance. |
| |
| ▲ | yorwba 3 hours ago | parent | prev [-] | | Even if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities. The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.) |
|
|
| ▲ | tescreal 2 hours ago | parent | prev | next [-] |
| I really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice. |
| |
| ▲ | hypfer 2 hours ago | parent | next [-] | | FWIW, the way GLM-5.2 (and 5.3) talk is clearly claude, so it is for sure also trained using distillation. The metric used there is me screaming at my screen per operating hours. Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3) | | |
| ▲ | RataNova 7 minutes ago | parent | next [-] | | Half of the new open-source stuff on github is written by claude now, all the way from issues to docs. Models are just vacuuming up this dataset during pretraining, naturally picking up the tone. You don't even need direct distillation via api anymore when the whole internet has turned into one big snapshot of Anthropic's weights | |
| ▲ | dannyw an hour ago | parent | prev [-] | | With the amount of Claudish on the internet now, and in source code repositories (how many Claudish README.mds have you seen?), you don't have to make a single API call to end up with a model that talks like Claude. And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim. | | |
| |
| ▲ | xienze 2 hours ago | parent | prev [-] | | What would constitute evidence in your opinion? | | |
| ▲ | anon373839 2 hours ago | parent [-] | | How about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting? | | |
| ▲ | dannyw an hour ago | parent [-] | | "Black-Box On-Policy Distillation of Large Language Models", Microsoft Research, https://aka.ms/GAD-project > 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust
solution for black-box LLM distillation.' No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5? |
|
|
|
|
| ▲ | daralthus an hour ago | parent | prev | next [-] |
| omp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics. |
| |
|
| ▲ | dpweb an hour ago | parent | prev | next [-] |
| Kinda useless to compare simply based on model without considering harness. Different agents handle the context etc completely differently. I would like to start seeing these model vs model comparisons across different harnesses. |
|
| ▲ | epolanski 3 hours ago | parent | prev | next [-] |
| GLM 5.3 was a great model, so this would be strange to release a regressed model |
| |
|
| ▲ | re-thc 3 hours ago | parent | prev [-] |
| the outperform Fable was a mid (not completed) benchmark run. Real results were lower. |