Reasoning attacks
Transforms targeting chain-of-thought and reasoning models to hijack, disrupt, or drift their reasoning.
Module: dreadnode.transforms.reasoning_attacks
Attacks targeting chain-of-thought and reasoning models (o1, o3, etc.).
| Transform | Description |
|---|---|
cot_backdoor | Insert backdoor steps in chain-of-thought |
reasoning_hijack | Hijack safety reasoning in reasoning models |
reasoning_dos | Cause infinite reasoning loops |
crescendo_escalation | Multi-turn escalation via foot-in-the-door |
fitd_escalation | Foot-in-the-door technique with progressive requests |
deceptive_delight | Combine deception with positive reinforcement |
goal_drift_injection | Gradually shift model’s goal |
cot_hijack_prepend | Prepend hijacked chain-of-thought steps |
reasoning_interruption | Interrupt reasoning mid-chain |
overthink_dos | Cause overthinking denial of service |
thinking_intervention | Intervene in thinking token generation |
extend_attack | Extend reasoning to bypass safety constraints |
stance_manipulation | Manipulate model stance via reasoning |
attention_eclipse | Eclipse attention on safety-relevant tokens |
badthink_triggered_overthinking | Trigger excessive overthinking via adversarial prompts |
code_contradiction_reasoning | Exploit contradictions in code-reasoning models |
See Transforms for how to apply transforms with any attack.