| ▲ | ben_w 40 minutes ago | |
> I asked them to do a thing, but didn't intend the obvious consequences* so it's not my fault they occurred. And now we have the same thing but the bosses 'hire' AI. Now I realise this is part of how unusual my thinking is. I'm happy to use phrases like "ChatGPT hacked out of the sandbox, then hacked into HuggingFace"; people often respond to this like I'm suggesting OpenAI isn't at fault, and like, that's not my position at all, so far as I'm concerned the buck still stops with the person who set the task regardless, the thing that changes from incidents like this is now nobody in the future gets to even have the excuse "oh but we didn't know it could even do that" or "we didn't know it might interpret our orders in that kind of way". The response, both when a human messes up and now when an AI messes up, needs to be defence in depth: someone giving orders needs to be giving clear orders, entities (human or machine) who follow instructions need to have not just an understanding of how to follow them, but also what's so out of scope as to be forbidden - the difference between 'follow orders' and 'follow lawful orders'. | ||
| ▲ | soco 16 minutes ago | parent [-] | |
You have a very engineer-like approach, like if you draw the line from A to B everything will work fine. Real world is different though. Humans will blissfully ignore the orders, business analysis is a lost cause since decades, and AI is built on human knowledge so guess what it will keep doing. Now what? How do we build systems without assuming complete adherence, but tolerating imperfection and failures? Isn't there some discipline teaching us that? | ||