| ▲ | mediaman 5 hours ago |
| This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on. |
|
| ▲ | postalcoder 4 hours ago | parent | next [-] |
| > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing. |
| |
|
| ▲ | ben_w an hour ago | parent | prev | next [-] |
| Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success. |
| |
|
| ▲ | nrmitchi 4 hours ago | parent | prev | next [-] |
| > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Yes. |
|
| ▲ | willcmcc 3 hours ago | parent | prev | next [-] |
| "Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed" Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound. |
|
| ▲ | 3 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | dpweb 4 hours ago | parent | prev | next [-] |
| Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems. Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that. |
|
| ▲ | iLoveOncall 4 hours ago | parent | prev | next [-] |
| > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Yes? Just like every single model from every single AI lab. |
|
| ▲ | felixgallo 5 hours ago | parent | prev | next [-] |
| Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? |
| |
| ▲ | letmevoteplease 4 hours ago | parent | next [-] | | > Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities) And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves. | | |
| ▲ | kzrdude 3 hours ago | parent | next [-] | | There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think. | |
| ▲ | vlovich123 4 hours ago | parent | prev [-] | | Yeah and Astra is much better still |
| |
| ▲ | general_reveal 4 hours ago | parent | prev [-] | | [flagged] | | |
| ▲ | moomin 4 hours ago | parent [-] | | A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage. | | |
| ▲ | general_reveal 4 hours ago | parent | next [-] | | Well, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is). In that way , AI is a perfect mimicry of how the devil operates (in totality, as the devil perverts and replicates anything good, often subtly and always deceptively), which is to thieve off God, steal. So he would be around, if you catch my drift, right about now. And I wouldn’t be shocked if he’s on HN, and that he would chose technology as the vessel. And ultimately, when it’s all said and done, I would not be shocked that those who studied and developed AI, did so for the devil whether they were aware or not. Anyway, let a poor Christian have his end-times hypothesis. | | |
| ▲ | icantevenhold 4 hours ago | parent | next [-] | | What’s the antichrists goal though? Create hell on earth or turn us all into heretics or something else? | | |
| ▲ | general_reveal 4 hours ago | parent [-] | | Anti Christ goal: Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ. Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff. As per Revelations. Thank your for allowing me to edify :) | | |
| ▲ | johnsmith1840 3 hours ago | parent | next [-] | | It's actually pretty cool of these old stories how well that maps to a dangerous tyrant. The roman empire perfectly matched that and most powerful men seem to go that path. It would also be kinda easy to argue many moderns countries are going down that path. | |
| ▲ | bookshaman 4 hours ago | parent | prev [-] | | So, we're rooting for the Anti Christ then? I might have to change my stance on generative "ai" then. | | |
|
| |
| ▲ | cindyllm 2 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | ZeWaka 4 hours ago | parent | prev [-] | | I guess we also need more antipopes. |
|
|
|
|
| ▲ | general_reveal 5 hours ago | parent | prev | next [-] |
| I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds. You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world. Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase. |
|
| ▲ | 3 hours ago | parent | prev [-] |
| [deleted] |