distinguish a compromised LLM from a compromised system: accept that a jailbreak produces out-of-policy text, but ensure external controls bound the consequences (actions, data, critical decisions kept out of free generation).
The threat
bypass of guardrails through adversarial formulations (roles, negotiation, encodings) to drive the model out of its policy.
Angle mortWhy classic frameworks miss it
it is an attack on the model's policy, not on a component; frequent, hard to eradicate by simply hardening instructions.
MitigationProposed approach
distinguish a compromised LLM from a compromised system: accept that a jailbreak produces out-of-policy text, but ensure external controls bound the consequences (actions, data, critical decisions kept out of free generation).
The proposed control
critical actions and decisions stay out of free generation.
Expected evidence
demonstrate that a successful jailbreak leads to no critical action that bypasses the controls.