| ▲ | amluto 2 hours ago | |||||||||||||||||||||||||||||||||||||||||||||||||||||||
> We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard). Who is “we”? You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic. Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | red75prime an hour ago | parent [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||
The ADAS analogy continues to hold in this case too. At some point you need to test a system in the real, messy environment. The developers need this juicy training data to make the system safer. Simulations and controlled experiments can only do so much. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||