Remix.run Logo
letmevoteplease 3 hours ago

> Altman was caught in previous attempts trying to game benchmarks

Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.

> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)

And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.

kzrdude 2 hours ago | parent | next [-]

There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.

vlovich123 3 hours ago | parent | prev [-]

Yeah and Astra is much better still