| ▲ | garo-pro 9 hours ago | |||||||
Most interesting here: > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior. | ||||||||
| ▲ | CTDOCodebases an hour ago | parent | next [-] | |||||||
Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it? Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated. | ||||||||
| ||||||||
| ▲ | Onavo an hour ago | parent | prev [-] | |||||||
Bet you they will use an external LLM world model to emulate the tools going forward. It's basically what's done in self driving research. | ||||||||
| ||||||||