| ▲ | bonoboTP 2 hours ago | |
No, they can just solve tasks verifiably. Reinforcement learning from verifiable rewards. Pure next-token supervision is only in the pre-training phase for modern LLM agents. They can also define new RL environments and pose new challenges to themselves and train themselves to solve them faster. Yes, at some point some human steering and judgment comes in as to what sorts of tasks to train towards. But there is no obvious taper-off at human level. | ||