| ▲ | swingboy 3 hours ago | |
There’s a difference between a model recursively improving “itself” and improving itself via online learning, right? The former being that these models are helping develop and train future models, but they might not veer too far off in architecture (yet). The latter being the same model being able to train/learn on the fly, in real time, permanently (not just in the current conversation/session), or in other words, adjusting/managing its own weights. The latter seems far more likely to go out of control than the former. But, it also seems like it would take an entire paradigm shift in model architecture, but I could be wrong. Does anyone in the industry think any of these companies are actually close to that kind of self-improvement? | ||
| ▲ | erichocean 3 hours ago | parent [-] | |
> Does anyone in the industry think any of these companies are actually close to that kind of self-improvement? Their current plan is to take the existing architecture and shorten the cycle times: move all new RLVR work into mid-training on a pre-existing base; apply new RLVR. Rinse and repeat. If you did that daily, it would be roughly similar to how humans improve. | ||