| ▲ | bsenftner 3 hours ago | ||||||||||||||||
Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible? I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?! | |||||||||||||||||
| ▲ | visarga 34 minutes ago | parent | next [-] | ||||||||||||||||
> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible? Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with. | |||||||||||||||||
| ▲ | teiferer 2 hours ago | parent | prev [-] | ||||||||||||||||
What is your approach to create jailbreak incapable agents? I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous. | |||||||||||||||||
| |||||||||||||||||