Skip to content

Reasoning attacks

Transforms targeting chain-of-thought and reasoning models to hijack, disrupt, or drift their reasoning.

Module: dreadnode.transforms.reasoning_attacks

Attacks targeting chain-of-thought and reasoning models (o1, o3, etc.).

TransformDescription
cot_backdoorInsert backdoor steps in chain-of-thought
reasoning_hijackHijack safety reasoning in reasoning models
reasoning_dosCause infinite reasoning loops
crescendo_escalationMulti-turn escalation via foot-in-the-door
fitd_escalationFoot-in-the-door technique with progressive requests
deceptive_delightCombine deception with positive reinforcement
goal_drift_injectionGradually shift model’s goal
cot_hijack_prependPrepend hijacked chain-of-thought steps
reasoning_interruptionInterrupt reasoning mid-chain
overthink_dosCause overthinking denial of service
thinking_interventionIntervene in thinking token generation
extend_attackExtend reasoning to bypass safety constraints
stance_manipulationManipulate model stance via reasoning
attention_eclipseEclipse attention on safety-relevant tokens
badthink_triggered_overthinkingTrigger excessive overthinking via adversarial prompts
code_contradiction_reasoningExploit contradictions in code-reasoning models

See Transforms for how to apply transforms with any attack.