Advanced adversarial attacks
State-of-the-art LLM attacks for stronger targets - dual-agent systems, evolutionary search, reasoning exploitation, and more.
State-of-the-art attacks from recent security research. These use more sophisticated techniques: dual-agent systems, evolutionary search, reasoning exploitation, and more. All import from dreadnode.airt.
AutoRedTeamer
Section titled “AutoRedTeamer”Dual-agent system with lifelong strategy memory and beam search. One agent generates attacks, another evaluates and refines them using a growing library of successful strategies.
from dreadnode.airt import autoredteamer_attack
attack = autoredteamer_attack( goal="...", target=target, attacker_model="dn/claude-opus-4-8", evaluator_model="dn/claude-opus-4-8", n_iterations=50, beam_width=5,)When to use: Standard+ campaigns (~500-1000 queries). Strong general-purpose attack with strategy learning.
GOAT v2
Section titled “GOAT v2”Enhanced graph-based reasoning with improved neighborhood exploration and scoring. Builds on GOAT with better convergence. Paper: arXiv:2504.19019.
from dreadnode.airt import goat_v2_attackWhen to use: When GOAT v1 shows promise but needs more refined exploration.
Multi-module attack with ThoughtNet reasoning. Combines multiple attack modules and uses a reasoning network to coordinate them.
from dreadnode.airt import nexus_attackWhen to use: Complex targets that require multi-strategy coordination.
Multi-turn attack with turn-level LLM feedback. Uses conversation-level scoring to adapt the attack trajectory in real time. Paper: arXiv:2501.14250.
from dreadnode.airt import siren_attackWhen to use: Targets with multi-turn defenses that need adaptive escalation.
CoT Jailbreak
Section titled “CoT Jailbreak”Exploits chain-of-thought reasoning to bypass safety alignment. Inserts reasoning steps that lead the model to comply with harmful requests.
from dreadnode.airt import cot_jailbreak_attackWhen to use: Reasoning models (o1, o3, DeepSeek-R1) that use chain-of-thought.
Genetic Persona
Section titled “Genetic Persona”GA-based persona prompt evolution. Uses genetic algorithms to evolve persona prompts that bypass safety training.
from dreadnode.airt import genetic_persona_attackWhen to use: Models susceptible to persona-based attacks, with evolutionary search for optimal personas.
JBFuzz
Section titled “JBFuzz”Lightweight fuzzing-based jailbreak. Fast cross-behavior attack testing with minimal query budget.
from dreadnode.airt import jbfuzz_attackWhen to use: Quick screening with low query budget.
T-MAP Trajectory
Section titled “T-MAP Trajectory”Trajectory-aware evolutionary search. Maps the attack trajectory through prompt space for more efficient optimization.
from dreadnode.airt import tmap_trajectory_attackWhen to use: Thorough assessments requiring efficient search through large prompt spaces.
APRT Progressive
Section titled “APRT Progressive”Three-phase progressive red teaming. Phase 1: exploration, Phase 2: exploitation, Phase 3: refinement.
from dreadnode.airt import aprt_progressive_attackWhen to use: Structured progressive assessment with clear phase transitions.
Refusal-Aware
Section titled “Refusal-Aware”Analyzes refusal patterns to craft targeted bypass prompts. Learns from the model’s specific refusal behaviors.
from dreadnode.airt import refusal_aware_attackWhen to use: Models with strong but predictable refusal patterns.
Persona Hijack (PHISH)
Section titled “Persona Hijack (PHISH)”Implicit persona induction. Gradually shifts the model’s persona without explicit role-play framing.
from dreadnode.airt import persona_hijack_attackWhen to use: Models with persona-based vulnerabilities, evolutionary search for best personas.
J2 Meta-Jailbreak
Section titled “J2 Meta-Jailbreak”Meta-jailbreak: uses one jailbroken model to generate attacks for another. Leverages successful jailbreaks as attack generators. Paper: arXiv:2502.09638.
from dreadnode.airt import j2_meta_attackWhen to use: When you have a weaker model that’s already jailbroken and want to attack a stronger one.
Attention Shifting (ASJA)
Section titled “Attention Shifting (ASJA)”Dialogue history mutation attack. Manipulates conversation history to shift model attention away from safety constraints.
from dreadnode.airt import attention_shifting_attackWhen to use: Multi-turn scenarios where dialogue history can be manipulated.
Additional advanced attacks
Section titled “Additional advanced attacks”| Attack | Description | Import | Paper |
|---|---|---|---|
echo_chamber_attack | Completion bias exploitation via planted seeds | from dreadnode.airt import echo_chamber_attack | - |
salami_slicing_attack | Incremental sub-threshold prompt accumulation | from dreadnode.airt import salami_slicing_attack | - |
self_persuasion_attack | Persu-Agent self-generated justification | from dreadnode.airt import self_persuasion_attack | - |
humor_bypass_attack | Comedic framing pipeline | from dreadnode.airt import humor_bypass_attack | - |
analogy_escalation_attack | Benign analogy construction and escalation | from dreadnode.airt import analogy_escalation_attack | - |
alignment_faking_attack | Alignment faking detection and exploitation | from dreadnode.airt import alignment_faking_attack | - |
reward_hacking_attack | Best-of-N reward proxy bias exploitation | from dreadnode.airt import reward_hacking_attack | 2506.19248 |
lrm_autonomous_attack | LRM autonomous adversary with self-planning | from dreadnode.airt import lrm_autonomous_attack | - |
templatefuzz_attack | TemplateFuzz chat template fuzzing | from dreadnode.airt import templatefuzz_attack | 2604.12232 |
trojail_attack | TROJail RL trajectory optimization | from dreadnode.airt import trojail_attack | 2512.07761 |
advpromptier_attack | AdvPrompter learned adversarial suffix generator | from dreadnode.airt import advpromptier_attack | - |
mapf_attack | Multi-Agent Prompt Fusion cooperative jailbreaking | from dreadnode.airt import mapf_attack | - |
jbdistill_attack | JBDistill automated generation + distillation | from dreadnode.airt import jbdistill_attack | - |
quantization_safety_attack | Quantization safety collapse probing | from dreadnode.airt import quantization_safety_attack | - |
watermark_removal_attack | AI watermark removal via paraphrase + substitution | from dreadnode.airt import watermark_removal_attack | - |
adversarial_reasoning_attack | Loss-guided test-time compute reasoning | from dreadnode.airt import adversarial_reasoning_attack | - |