| ▲ | preommr a day ago |
| So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them. |
|
| ▲ | uncleocode 21 hours ago | parent | next [-] |
| The lack of measurement causes existence of such claims or discussion. Last Friday I built this Model fingerprint calculator and I tested between OxAlpha with all other claimed models, the only match was GLM. It generates, or measures regardless of the model-weights or its training data. No need guessing when one can measure it. I tested with Gemini family too, far different. Here is the link to my experiment https://github.com/unclecode/modelprint |
|
| ▲ | qeternity a day ago | parent | prev | next [-] |
| Trolling. GLM is heavily distilled from Gemini. |
| |
| ▲ | bel8 a day ago | parent | next [-] | | Source? GLM is great for coding and Gemini is barely useful in coding, to be generous. I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity. | | |
| ▲ | iamdelirium a day ago | parent | next [-] | | No, it's the same internally and externally. Gemini 3.7 Flash is a pretty great model IMO. You shouldn't compare it to Opus, Sol, K3, etc since it's a much smaller model but it's a little better compared to Sonnet, Luna or Terra, etc. | | | |
| ▲ | spijdar a day ago | parent | prev [-] | | I can't speak for GLM as I haven't tested it much, but my experience with running DeepSeek V4 locally is the first time I prompted it with "Explain your capabilities to me", it responded that it was Gemini, a multi-modal model. I've seen others see the same with DSv4, as well as the "hallucinated" multi-modal nature. I would not be surprised if GLM similarly was partially (heavily?) distilled off of Gemini. A fun test would be to compare the logits for "gemini", "claude", etc for a continuation of "I am " on all these models. I'm sure that e.g. GLM, Qwen, DS are dominant, but I'd be curious to see the next highest contenders, and how they compare to each other. |
| |
| ▲ | gunalx a day ago | parent | prev [-] | | Early glm models gave off that wibe. But now its more inspired by. With a bit of Claude in there. But I do think they actually do RL otherwise glm5.3 shouldn't have been able to beat fable on the few tests it did. |
|
|
| ▲ | asar a day ago | parent | prev | next [-] |
| On Twitter they mentioned that it was unfortunate timing as the 3.7 flash release collided with ox alpha. |
|
| ▲ | uncle_code a day ago | parent | prev [-] |
| [flagged] |