| ▲ | Gigachad 2 hours ago |
| Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect. Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task. |
|
| ▲ | SequoiaHope 15 minutes ago | parent | next [-] |
| This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good. |
|
| ▲ | janalsncm 31 minutes ago | parent | prev | next [-] |
| What did you think of the author’s concerns on the thing you are suggesting? |
| |
| ▲ | SequoiaHope 16 minutes ago | parent [-] | | Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far… |
|
|
| ▲ | baxtr an hour ago | parent | prev | next [-] |
| That could work. My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not? Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent. |
| |
| ▲ | ben_w 5 minutes ago | parent | next [-] | | A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway. "Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour". (The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff). | |
| ▲ | saagarjha a minute ago | parent | prev | next [-] | | This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this. | |
| ▲ | mdp2021 27 minutes ago | parent | prev | next [-] | | > If AI is really smart Well, it's not. > AGI smart for some Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly. -- Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence. Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances. More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement. | | |
| ▲ | hiAndrewQuinn a minute ago | parent | next [-] | | This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it. | |
| ▲ | ben_w 19 minutes ago | parent | prev [-] | | > Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence. If this was true, why are the history books littered with so many evil people who gained power? This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios. (Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases). |
| |
| ▲ | mulmen 28 minutes ago | parent | prev [-] | | Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity. If you want to control an intelligence incentive, not morality, is the tool to reach for. | | |
| ▲ | mdp2021 8 minutes ago | parent [-] | | (Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.) Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion. | | |
|
|
|
| ▲ | mrweasel an hour ago | parent | prev | next [-] |
| That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on? I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on. Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system. |
| |
| ▲ | msdz 3 minutes ago | parent [-] | | >> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. > That does seem a little like solving the problems in AI by using more of it Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it. [0] Cf. CaMeL: https://arxiv.org/abs/2503.18813 |
|
|
| ▲ | aytigra an hour ago | parent | prev | next [-] |
| The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI.
Alternatively they could also both escalate and go off the rails while warring with each other. |
| |
| ▲ | LoganDark 31 minutes ago | parent | next [-] | | You don't necessarily need a reviewer that's immune to prompt injection, you just need one that can express a panic state with conflicting/ambiguous material rather than going along with it, and you can treat that with a shutoff to be safe, or an operator review. | |
| ▲ | hanibrel 23 minutes ago | parent | prev [-] | | [dead] |
|
|
| ▲ | RandomLensman an hour ago | parent | prev | next [-] |
| With plenty of things we do not allow use outside of some regulated environment, nothing new. Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox. |
|
| ▲ | nxpnsv 31 minutes ago | parent | prev | next [-] |
| Is that not a recipe for adversarial training, thus ensuring increasing misalignment…? |
|
| ▲ | chrisjj 7 minutes ago | parent | prev | next [-] |
| [delayed] |
|
| ▲ | bigstrat2003 2 hours ago | parent | prev [-] |
| If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely. |
| |
| ▲ | Gigachad 2 hours ago | parent | next [-] | | People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens. | |
| ▲ | dipper139 an hour ago | parent | prev | next [-] | | I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long | |
| ▲ | rlpb 23 minutes ago | parent | prev [-] | | And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner. |
|