| ▲ | nicce 9 hours ago |
| Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost. |
|
| ▲ | solarkraft 9 hours ago | parent | next [-] |
| Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data. |
| |
| ▲ | nicce 9 hours ago | parent | next [-] | | In a Codex subreddit there is a bunch of stats. | |
| ▲ | pupppet 9 hours ago | parent | prev [-] | | I don’t understand how there isn’t a website out there tracking this stuff already. | | |
| ▲ | copperx 6 hours ago | parent | next [-] | | https://modelregression.com/ | | |
| ▲ | hn8726 5 hours ago | parent [-] | | Haven't looked into how accurate the page is, but the list of regressions on the bottom looks terrifying, at first glance? | | |
| ▲ | copperx 5 hours ago | parent [-] | | Yes, it looks like regressions are frequent, but sometimes performance goes back to baseline quite fast. |
|
| |
| ▲ | bpavuk 9 hours ago | parent | prev [-] | | how the hell do we even track that? and before someone says... —"Benchmarks!" ...I'll tell that they can be gamed so easily, and they are on a consistent basis. | | |
| ▲ | solarkraft 8 hours ago | parent [-] | | Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks? | | |
| ▲ | nananana9 2 hours ago | parent [-] | | If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought? |
|
|
|
|
|
| ▲ | hn8726 5 hours ago | parent | prev | next [-] |
| 100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes |
|
| ▲ | koyote 8 hours ago | parent | prev | next [-] |
| I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence! With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]"). |
|
| ▲ | 9 hours ago | parent | prev [-] |
| [deleted] |