| ▲ | nxc18 21 hours ago |
| There must be some benefit if all the providers are doing it independently. GPT5.6-Sol on Max thinking just became regarded as of a few days ago. The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago). The cycle repeats. |
|
| ▲ | claydugo 16 hours ago | parent | next [-] |
| Astra is also useless and completely ignoring instructions at random intervals. We are being A/B tested on and there is nothing you can do about it. |
| |
| ▲ | CamperBob2 14 hours ago | parent [-] | | We are being A/B tested on and there is nothing you can do about it. Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput. To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control. |
|
|
| ▲ | 21 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | bitexploder 12 hours ago | parent | prev | next [-] |
| They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice. |
|
| ▲ | talon8635 19 hours ago | parent | prev | next [-] |
| Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant? Thanks for your insight |
| |
| ▲ | mrandish 16 hours ago | parent | next [-] | | > might they all pursue this kind of deception They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record. Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves. | |
| ▲ | onemoresoop 15 hours ago | parent | prev | next [-] | | See Shepard tone. Similarly model releases could be engineered to appear that they’re always getting better by slowly degrading and upgrading at the right time. That plus hitting some benchmarks and making a lot of noise around that. | |
| ▲ | dalenw 19 hours ago | parent | prev [-] | | Kinda. Off the top of my head, DeepSeek and their thinking model was pretty new and interesting. Multi input models are also newish (combined input of text, image, video, audio, etc). Then there's Jev, a recently release that has a lot of people talking. It isn't really an LLM, but also is one. Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models. Overall models have become cheaper to run and smarter per token. |
|
|
| ▲ | 20 hours ago | parent | prev [-] |
| [deleted] |