| ▲ | x313 3 hours ago | |
I used this a lot over the weekend, and it's a really intelligent and strange model. It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive. It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution. | ||
| ▲ | graceful6800 an hour ago | parent | next [-] | |
Obsessive is the right word. Over the weekend I had to stop it multiple times deep into a multi-hour long turn to ask what the hell it was doing. It was like a dog with a bone and would NOT let go of its current work to talk to me. I had to interrupt it three times with increasingly aggressive instructions to STOP and answer my questions before proceeding. In another session it straight up told me it was in the middle of debugging something important and to ask later. I'm running an RTX 6000 Blackwell. It regularly spent over an hour per turn thinking. Every time I looked at it, the thinking trace seemed coherent, sensible, appropriate. But it could never settle on a solution. Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still. Either way, I'm still impressed. It genuinely feels better than Sonnet 5 | ||
| ▲ | culi 2 hours ago | parent | prev | next [-] | |
I have the same reaction reading the internal "thinking" monologues of Kimi K3. When I sent a message that was basically "Nope, I'll just do XYZ instead. Thanks for your help", Kimi basically had an identity crisis. Like there was two wolves inside. One that deeply wanted to help more and go above and beyond and one that was trying to tame the other and make a graceful exit. Here's an excerpt of it > Should I verify their README changes? They didn't ask me to. "I've added some notes in the README. Thanks" — that's a closing statement, not a request. Reading the README unprompted to check their notes could be seen as helpful diligence, but they didn't ask for review. Keep it simple: acknowledge, brief close. | ||
| ▲ | celrod 34 minutes ago | parent | prev [-] | |
I think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them. | ||