Remix.run Logo
Red queen hypothesis – A new way forward for self-improving AI(cst.cam.ac.uk)
50 points by hardlianotion 12 hours ago | 10 comments
nullbio an hour ago | parent | next [-]

It seems to me this is only useful for self-improvement up to the point of accomplishing objectives and problems that humans have already clearly defined, solved and mapped. For example, "At checkpoints, a stronger evaluator can replace the old one if it performs better on trusted ground-truth examples." - if they're a trusted ground-truth, they must be rigorous. If we're attempting to solve problems that humans have not already solved, where do you get your ground-truth examples? It's not like the AI is going to be able to generate these for you if it has never seen a solution to the problem.

I can definitely see the argument that this allows us to train models faster and converge faster, because if you scale difficulty of evaluation alongside the learners capabilities, it spends a lot less time floudering around. It basically works out to be loss minimization through strategic ordering of the training data. Is that the goal here though? Or is the goal recursive self-improvement and solving problems that are currently outside of reach? Because it doesn't feel like the latter would be possible with this design.

For example, how do you quantify "the evaluation gets -harder- as the agent gets better". Harder, how? In what direction? Via what criteria or measure?

robotresearcher 6 hours ago | parent | prev | next [-]

Here’s a paper by Floreano at EPFL from 1997 explicitly on Red Queen dynamics for creating neural networks for intelligent robot control.

There was lots of discussion of these ideas in the 1990s. In those days we trained very small NNs - tens of nodes - by evolving their weights and topologies. A run could take days on a workstation of the time.

This particular paper is about co-evolving predator and prey, where the behavior of each is the ‘evaluation’ of the other.

https://infoscience.epfl.ch/entities/publication/a65d0679-68...

jldugger 3 hours ago | parent [-]

OP's link: > Now the researchers have addressed this issue by having both the self-improving agent and the evaluator evolve together.

and your quote:

> This particular paper is about co-evolving predator and prey, where the behavior of each is the ‘evaluation’ of the other.

Both sound like the GAN approach that was popularized a decade ago and kinda the start of the "genAI" boom.

JacobAsmuth 40 minutes ago | parent | prev | next [-]

> The research team, which includes collaborators from NVIDIA and Flower Labs, have come up with a new method for recursive self-improving AI agents to continue improving themselves.

What happens if you apply the method to non-recursive self-improving AI agents? Can they continue improving themselves? Or does the recursive self improvement only recursively self improve AI agents which are themselves recursively self-improving?

PeterStuer 38 minutes ago | parent | prev | next [-]

This 'new' method was quite common in evolutionary computing in the 90's.

richardfey 3 hours ago | parent | prev | next [-]

> "Instead of improving an agent against a fixed test, we let the evaluation evolve alongside the agent"

This quote should have been highlighted earlier in the article.

throwa356262 2 hours ago | parent | prev | next [-]

Have not read the paper yet, but this not sound like GAN applied to agent training?

seu an hour ago | parent | prev | next [-]

> At a time when there's keen public interest in AI that can make itself better

Really? From whom beyond the Musks and Altmans and their cohorts?

> The research team, which includes collaborators from NVIDIA

Aha

Morromist an hour ago | parent [-]

The polls I've seen make it seem like there's much more public distress about AI than excitement, so yeah. Honestly pushing ai into people's faces and talking non-stop about how great it is really has a negative effect on its PR I belive. People don't like being told "you'll eat it and you'll like it" and AI bros haven't got much empathy for people who see it as ugly, boring, gross and filling the internet and life generally with crap.

https://www.pewresearch.org/short-reads/2026/03/12/key-findi...

Taikhoom10 6 hours ago | parent | prev [-]

Yeah I think it is broadly applicable to tech as a whole, I mean any great startup is just really a counter positioned company to incumbents -