| ▲ | msdz an hour ago | |
>> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. > That does seem a little like solving the problems in AI by using more of it Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it. [0] Cf. CaMeL: https://arxiv.org/abs/2503.18813 | ||