| ▲ | Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1(tokenstead.ai) |
| 31 points by cdnsteve 3 hours ago | 13 comments |
| |
|
| ▲ | forgot-my-pw an hour ago | parent | next [-] |
| Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso. Still quite impressive though. |
|
| ▲ | ChrisArchitect 43 minutes ago | parent | prev | next [-] |
| Related: Cognition launches new SWE-2 model https://news.ycombinator.com/item?id=49645443 |
|
| ▲ | walrus01 an hour ago | parent | prev | next [-] |
| Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system. |
|
| ▲ | varispeed 2 hours ago | parent | prev [-] |
| These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago. |
| |
| ▲ | kzrdude an hour ago | parent | next [-] | | Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on? | | |
| ▲ | varispeed an hour ago | parent [-] | | It is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops. Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc. |
| |
| ▲ | cbg0 an hour ago | parent | prev | next [-] | | Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/ | | |
| ▲ | varispeed 44 minutes ago | parent [-] | | That is not conclusive, because they can detect such tracker and route it through proper not quantised model. | | |
| |
| ▲ | d_tr an hour ago | parent | prev [-] | | How and why do they get nerfed? To save money? | | |
| ▲ | johnfn 31 minutes ago | parent | next [-] | | Models do not get nerfed. There has never been evidence of this. This would be trivial to prove if it were true, and such a proof would be a huge story and scandal to a news market hungry for a shred of a signal on AI's downfall. This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s. | | |
| ▲ | esafak 17 minutes ago | parent [-] | | Yes, they can, through quantization. Many providers of open source models openly serve quantized versions; check openrouter. |
| |
| ▲ | varispeed 43 minutes ago | parent | prev [-] | | Yes. They save on compute and customer has to use more tokens to achieve their goal which means more profit. |
|
|