Remix.run Logo
▲ amluto 2 hours ago

> We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).

Who is “we”?

You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.

Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.

▲red75prime an hour ago | parent [-]

The ADAS analogy continues to hold in this case too. At some point you need to test a system in the real, messy environment. The developers need this juicy training data to make the system safer. Simulations and controlled experiments can only do so much.

▲amluto an hour ago | parent | next [-]

I’m not entirely convinced. If I use Claude and instruct it to literally perform interesting interactions on random websites, I would feel like I, personally, am doing something that is at least a bit immoral and a bit unsafe. More so if I have it skills or training to actually sign in and submit forms as though it’s a himan.

Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.

▲red75prime 35 minutes ago | parent [-]

You can solve CAPTCHAs with open-weights models. I guess Anthropic works on models not solving CAPTCHAs unless they (models) have a legitimate reason to do it.

▲bediger4000 an hour ago | parent | prev [-]

So we should accept a program giving a false tip on a murder investigation? Where's the line on "testing a system in the real, messy environment"? If this isn't something that we should penalize, what is? What benefits am I, or is society, going to get out of what would be criminal behavior if anyone else did it?

▲red75prime an hour ago | parent [-]

We should accept that to get a system that is safe to work in the chaotic real world environment the developers have to test the system in the chaotic real world environment. And that the testing might cause problems.

If this testing constitutes a criminally negligent behavior, it should be punished. Given the benign outcome (the tip got into spam, Anthropic promptly contacted the police) I doubt that it will make the case.

▲amluto 35 minutes ago | parent | next [-]

Somehow Waymo built their system without having their test cars drive into buildings, nudge other cars off the road, honk incessantly, drive backwards just for the heck of it, etc.

▲red75prime 24 minutes ago | parent [-]

I can't say whether you are being sarcastic, but Waymo has its share of problems. For example, see NHTSA recalls 24E013, 24E049, 25E034, 25E084, 26E026, 26E035.

▲bediger4000 27 minutes ago | parent | prev [-]

That's not the line asked for, that's just repeating the statement that caused me to ask for a line, or criteria for when we should prosecute. Sounds like the answer is "never", just like for the copyright infringement on a planetary scale. Ok. "AI" gets a free pass from the state. I better see some of these insanely good benefits, or I'm going to be angry. I bet the rest of the mob will be, too.