Remix.run Logo
▲ derektank 3 hours ago

Does anyone have more insight into how chain of thought might be subverted without meaningfully impacting model performance? I’ve heard this for a while now, and I understand how information might be retained in the weights that isn’t documented in the output. But weren’t reasoning models created in the first place because they provided a performance improvement in terms of output? Is that no longer the case? If so, why are the big labs still creating reasoning models?

▲thefxperson 2 hours ago | parent | next [-]

My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer.

Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:

  it remains unclear to what extent these performance gains can be attributed to human-like task decomposition or simply the greater computation that additional tokens allow. [...] our results show that additional tokens can provide computational benefits independent of token choice. The fact that intermediate tokens can act as filler tokens raises concerns about large language models engaging in unauditable, hidden computations that are increasingly detached from the observed chain-of-thought tokens.
https://arxiv.org/html/2404.15758v1
▲smallmancontrov 2 hours ago | parent | prev | next [-]

Sibling posts are correct -- the chain-of-thought is doing hidden computation, it has been shown in the linked papers.

If you want to see it yourself: load up Qwen 3.8 in LM Studio and watch the CoT stumble around like a drunken sailor before miraculously jumping to the correct result.

If you want an example of subversion, Anthropic has some good ones:

https://transformer-circuits.pub/2025/attribution-graphs/bio...

https://transformer-circuits.pub/2025/attribution-graphs/bio...

▲janalsncm 2 hours ago | parent | prev | next [-]

Before RL we typically SFT on human reasoning traces. This makes the reasoning traces somewhat coherent and the model trains faster.

But you don’t have to do that. You can skip straight to RL. If you do, the model will generate complete garbage reasoning traces before generating the correct answer. In fact, if you add a coherence reward to the reasoning trace, the model will perform worse (since you’re now diluting the correctness reward).

▲ianjbutler an hour ago | parent [-]

> the model will perform worse

Depends on whether and how you want to rank stability in terms of better/worse. Models are diverging on this, which seems increasingly clear.. i.e. Fable isn't stable, but Opus isn't clever, and they hit different kinds of walls. So both the theory (diluting the correctness reward) and the practice (hard split on plan/implement/review work) seems to be pointing towards a strongly multi-model and highly agentic / harness-driven / complex-system kind of future instead of singleton monolithic super-smart models.

The do-everything model with solid reasoning AND solid results, and the honest/introspective helpful agent that doesn't actively resist governance may be at odds. Stable reasoning doesn't matter for pen-testing, and correct-answer with broken processes and fragile abstractions won't matter for math/science/coding.

▲ForHackernews 3 hours ago | parent | prev | next [-]

I'm not sure anyone meaningfully understands it: "Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens" https://arxiv.org/html/2505.13775v3

> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.

▲AlexCoventry an hour ago | parent | prev [-]

https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...

> *TLDR*: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)