| ▲ | user43928 19 hours ago | |
If they had a 100% margin the cost would be 0. Let's look at open-weights models with 3T size: https://inferencex.semianalysis.com/run/kimi-k3-on-b200 This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3. We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna. There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol. This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks. | ||