Remix.run Logo
yorwba 6 hours ago

Every interactive system is a potential RL environment. Every CLI, every TUI, every GUI, every API. If you can programmatically take actions to get a result, and the actions are cheap, and the quality of the result can be measured automatically, you can set up an RL training loop and see whether the results get better over time.

radarsat1 2 hours ago | parent [-]

> and the quality of the result can be measured automatically

this part is nontrivial though