Remix.run Logo
▲ nicce 9 hours ago

Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.

▲solarkraft 9 hours ago | parent | next [-]

Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data.

▲nicce 9 hours ago | parent | next [-]

In a Codex subreddit there is a bunch of stats.

▲pupppet 9 hours ago | parent | prev [-]

I don’t understand how there isn’t a website out there tracking this stuff already.

▲copperx 6 hours ago | parent | next [-]

https://modelregression.com/

▲hn8726 5 hours ago | parent [-]

Haven't looked into how accurate the page is, but the list of regressions on the bottom looks terrifying, at first glance?

▲copperx 5 hours ago | parent [-]

Yes, it looks like regressions are frequent, but sometimes performance goes back to baseline quite fast.

▲bpavuk 9 hours ago | parent | prev [-]

how the hell do we even track that? and before someone says...

—"Benchmarks!"

...I'll tell that they can be gamed so easily, and they are on a consistent basis.

▲solarkraft 8 hours ago | parent [-]

Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks?

▲nananana9 2 hours ago | parent [-]

If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought?

▲hn8726 5 hours ago | parent | prev | next [-]

100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes

▲koyote 8 hours ago | parent | prev | next [-]

I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!

With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").

▲ 9 hours ago | parent | prev [-]
[deleted]