| ▲ | lonlundgren 12 hours ago | |
This is good feedback. Not sure if you read the long-form article vs. the tweet-thread linked here, but I did attempt to address most of the criticisms you listed in that writeup, if you haven't already read it. It's definitely not written in the same tone as the for-broader-publication thread. Regarding the amount of thinking as "magic sauce": the main issue is that even in the right tail, the delivery of thinking tokens almost never reaches the levels of published benchmarks. You once could include the word "ultrathink" in any prompt and it would provide a fixed thinking budget of 31,999 tokens. Whereas, in my corpus only 46 out of 36,374 invocations broke 16k thinking tokens, and the P90 was only 2,207. | ||