| ▲ | joshheitzman 3 hours ago | |
The actual title of the paper is: "Can AI agents conduct open-ended AI research? Early evidence from two case studies" While I appreciate that the article is throwing a web blanket on doomer claims, the actual study doesn't really get into AI self-improvement. That doesn't require writing papers. That just requires autonomously writing a software system that can produce a better AI agent then the one that created it. That said, I have little worry about this being possible as I have seen no evidence of AI agents being able to produce a working software system of that scale. | ||
| ▲ | HarHarVeryFunny 2 hours ago | parent | next [-] | |
I don't think RSI is typically used to describe self-improving agents - it's about improving the model itself, and its performance in agentic tasks. Most of the gains in model performance from one release to the next are coming from RLVR post training, which has changed a lot over the last couple of years. The old way was the model generates a response, then a static verifier looks at the response and evaluates it to assign a reward score. The new way is interactive with an agent running in a custom RL task simulation environment, then scored according to how well it completed the assigned task. For a SOTA model there will be many thousands of these simulation environments, each focusing on trying to teach the model/agent a different skill. Post-training also typically uses training curricula to walk the model up though through different levels of task difficulty. Training has become very complex. The job of a post-training AI research engineer consists of things like designing environments, designing training curricula, tweaking learning algorithms, running small scale experiments to verify ideas, etc. When people talk about RSI, it seems they are mostly talking about automating the job of the post-training research engineer - coming up with new ideas, testing them out, building these environments, etc. At the end of the day there is only so much development speed-up to be had since you still need to actually run those experiments and do the post-training, and are bottle-necked by the amount of compute available to do this. The economics of developing/selling LLMs also requires you to balance development compute cost with revenue generated by the resulting model, so even if you had the spare compute available to put into development, you are ultimately then bottle-necked by how fast can the model earn back that sunk cost before you can afford to start the next cycle. It's not all-or-nothing since some aspects of this automating the job of the post-training research engineer are easier than others, and are already being done, while the job as a whole obviously requires full human intelligence. | ||
| ▲ | marcosdumay 2 hours ago | parent | prev [-] | |
Do you think the final product of research is papers? | ||