Remix.run Logo
arijun an hour ago

How much thinking is going on beyond the spoken-aloud “thinking”? Does it have enough capability to have goals it doesn’t express explicitly? I suspect not, but I’m no expert.

Tumblewood 20 minutes ago | parent [-]

Yes, frontier models can reason outside their chain of thought and manipulate their chain of thought to some extent. The system card for Astra writes:

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks