| ▲ | m12k 14 hours ago | |
I'm wondering if some of Opus' chain of thought patterns have bled into its "human readable text output" circuitry. E.g. some of them that I stumbled on, "the bug is in the lock, not the query" reads like some of the shorthands it might use in its own chain of thought. | ||
| ▲ | orbital-decay 14 hours ago | parent | next [-] | |
Entirely possible, it always leaked but it's particularly bad in most recent models. Actually almost all issues with creative writing in LLMs are artifacts of either instruction tuning, alignment training, or CoT RL and seeding (newer models have CoT data even in pretraining). | ||
| ▲ | make3 13 hours ago | parent | prev [-] | |
This is my impression as well, that an exaggeratedly precise yet bad at communication with humans way of talking snuck in through reasoning RL, that it might be useful when it talks to itself (CoT) | ||