| ▲ | netvarun 10 hours ago |
| Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs.
Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds)
Sol is at 2/10 vs kimi’s 3/15 |
|
| ▲ | nicce 8 hours ago | parent | next [-] |
| Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost. |
| |
| ▲ | solarkraft 8 hours ago | parent | next [-] | | Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data. | | |
| ▲ | nicce 8 hours ago | parent | next [-] | | In a Codex subreddit there is a bunch of stats. | |
| ▲ | pupppet 8 hours ago | parent | prev [-] | | I don’t understand how there isn’t a website out there tracking this stuff already. | | |
| ▲ | copperx 5 hours ago | parent | next [-] | | https://modelregression.com/ | | |
| ▲ | hn8726 4 hours ago | parent [-] | | Haven't looked into how accurate the page is, but the list of regressions on the bottom looks terrifying, at first glance? | | |
| ▲ | copperx 4 hours ago | parent [-] | | Yes, it looks like regressions are frequent, but sometimes performance goes back to baseline quite fast. |
|
| |
| ▲ | bpavuk 8 hours ago | parent | prev [-] | | how the hell do we even track that? and before someone says... —"Benchmarks!" ...I'll tell that they can be gamed so easily, and they are on a consistent basis. | | |
| ▲ | solarkraft 7 hours ago | parent [-] | | Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks? | | |
| ▲ | nananana9 an hour ago | parent [-] | | If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought? |
|
|
|
| |
| ▲ | hn8726 5 hours ago | parent | prev | next [-] | | 100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes | |
| ▲ | koyote 7 hours ago | parent | prev | next [-] | | I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence! With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]"). | |
| ▲ | 8 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | pornel 9 hours ago | parent | prev | next [-] |
| Competition is good. Without K3/GLM/DS4 etc. there would be no pressure on OpenAI to drop Sol's price. |
|
| ▲ | drob518 10 hours ago | parent | prev | next [-] |
| Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost. |
|
| ▲ | k__ 8 hours ago | parent | prev | next [-] |
| With DeepSeek's pricing, no other value prop has been great. |
| |
|
| ▲ | nostrebored 10 hours ago | parent | prev | next [-] |
| Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design |
| |
|
| ▲ | 7777777phil 9 hours ago | parent | prev | next [-] |
| I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no: Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975. [1] https://philippdubach.com/posts/jev-model-router-for-pi/ |
|
| ▲ | toasty228 9 hours ago | parent | prev [-] |
| > With sol pricing drop 6 or 5.6? Because 6 is hot garbage |