Remix.run Logo
▲ garo-pro 9 hours ago

Most interesting here:

> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

▲CTDOCodebases an hour ago | parent | next [-]

Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?

Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.

▲eli an hour ago | parent [-]

Not connecting it to a network with internet access would probably be a good start.

▲Onavo an hour ago | parent | prev [-]

Bet you they will use an external LLM world model to emulate the tools going forward. It's basically what's done in self driving research.

▲jacquesm an hour ago | parent [-]

You're telling me they weren't doing that from day #1? Oh, wait...