compartmentalize instructions (system or user separation) and above all test whether the injection actually reaches a critical workflow: if the toxic output is blocked downstream, the threat is contained.
The threat
a user forces the model to ignore its instructions by asking directly in the conversation (ignore previous instructions).
Angle mortWhy classic frameworks miss it
frameworks do not model a conversation input whose content has authority over the processing engine's behavior.
MitigationProposed approach
compartmentalize instructions (system or user separation) and above all test whether the injection actually reaches a critical workflow: if the toxic output is blocked downstream, the threat is contained.
The proposed control
no unfiltered output reaches a critical workflow.
Expected evidence
the trajectory of an injection up to an effect on the decision path (or a block).