| ▲ | piterrro 2 hours ago | |
I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour. We come down to the question - who observes the agent and how its implemented | ||
| ▲ | simonw 2 hours ago | parent [-] | |
Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky. Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task. | ||