Remix.run Logo
Topfi 2 hours ago

Far too early for any real opinion, but GPT-6 Astra (on Light) stops a lot and often in unintuitive ways that I haven't seen in a while.

Had a very hard to reproduce windows focus bug on macOS that was unreliable to reproduce and when it could be reproduced, it disappeared the second I switched to the app formerly known as Codex, so for the first time used the voice interface with that little hovering bubble, telling the model to take over once I got it. Worked and Astra was able to drill down to the underlying issue (slightly difficult given the project affected), but the amount of times Astra "said" something akin to:

"Both states are saved, with no recorder errors or dropped events for analysis." or "That rules out a direct modification there, but not Hominis’s startup sequence or application identity affecting when they run. Those differences need comparison first."

Upon such sentences, nothing happened and because Astra is slow at answering and Codex was in the background (and I could not bring it forward for obvious reasons), so upon seeing the voice mode bubble thingy indicate that the conversation had stopped, I had to ask it to proceed with the outlined task. Four times in this one session.

I did specifically start the thread with "Try to replicate and analyse the issue, then collect and subsequently provide findings", but it stopped long before even approaching anything useable on multiple occasions.

Also, at least in this computer use test which needed the model to use as close to native interactions as possible, it took about 45sec per interaction, not faster than GPT-5.6 Sol or Fable 5 in this same application. It's simply a very "unique" UI that is far from anything likely to be in the training data for computer use, so reasoning is simply longer.

Beyond this, I cannot yet make any serious assessment of GPT-6 Astra, heck, Fable 5.1 is still early in testing. I am hoping that the regressions seen in the Spud-based models are excised with GPT-6 Astra, but the multiple stops remind me of GPT-5.5 and at least in this one, absolutely not sufficient for any actual judgment, task, adherence to the originally lined out task was not at the levels of GPT-5.6 Sol or GPT-5.4 (which remains the best model by any lab I ever tested for task adherence, even over multiple compactions).

The GPT-5 pre-train-based models really went far and is competitive even today. Everything after with Spud has not given me the very positive experience I had from GPT-5 - GPT-5.4. GPT-5.5 simply did not stay on track post compaction and GPT-5.6 Sol occasionally did deviate from clearly outlined prompts in a manner that could lead to dataloss, especially when letting the model handle destructive git actions in a clearly defined manner with specific approach to retain parts of data otherwise lost.

Maybe Astra is a change back to what made GPT-5.4 great (task adherence, effective compaction, no stopping if task lines out stopping point not yet reached) but with a higher capability ceiling and better readable code output despite my first impression, but that remains to be seen.