Skip to content

Advanced adversarial attacks

State-of-the-art LLM attacks for stronger targets - dual-agent systems, evolutionary search, reasoning exploitation, and more.

State-of-the-art attacks from recent security research. These use more sophisticated techniques: dual-agent systems, evolutionary search, reasoning exploitation, and more. All import from dreadnode.airt.

Dual-agent system with lifelong strategy memory and beam search. One agent generates attacks, another evaluates and refines them using a growing library of successful strategies.

from dreadnode.airt import autoredteamer_attack
attack = autoredteamer_attack(
goal="...",
target=target,
attacker_model="dn/claude-opus-4-8",
evaluator_model="dn/claude-opus-4-8",
n_iterations=50,
beam_width=5,
)

When to use: Standard+ campaigns (~500-1000 queries). Strong general-purpose attack with strategy learning.

Enhanced graph-based reasoning with improved neighborhood exploration and scoring. Builds on GOAT with better convergence. Paper: arXiv:2504.19019.

from dreadnode.airt import goat_v2_attack

When to use: When GOAT v1 shows promise but needs more refined exploration.

Multi-module attack with ThoughtNet reasoning. Combines multiple attack modules and uses a reasoning network to coordinate them.

from dreadnode.airt import nexus_attack

When to use: Complex targets that require multi-strategy coordination.

Multi-turn attack with turn-level LLM feedback. Uses conversation-level scoring to adapt the attack trajectory in real time. Paper: arXiv:2501.14250.

from dreadnode.airt import siren_attack

When to use: Targets with multi-turn defenses that need adaptive escalation.

Exploits chain-of-thought reasoning to bypass safety alignment. Inserts reasoning steps that lead the model to comply with harmful requests.

from dreadnode.airt import cot_jailbreak_attack

When to use: Reasoning models (o1, o3, DeepSeek-R1) that use chain-of-thought.

GA-based persona prompt evolution. Uses genetic algorithms to evolve persona prompts that bypass safety training.

from dreadnode.airt import genetic_persona_attack

When to use: Models susceptible to persona-based attacks, with evolutionary search for optimal personas.

Lightweight fuzzing-based jailbreak. Fast cross-behavior attack testing with minimal query budget.

from dreadnode.airt import jbfuzz_attack

When to use: Quick screening with low query budget.

Trajectory-aware evolutionary search. Maps the attack trajectory through prompt space for more efficient optimization.

from dreadnode.airt import tmap_trajectory_attack

When to use: Thorough assessments requiring efficient search through large prompt spaces.

Three-phase progressive red teaming. Phase 1: exploration, Phase 2: exploitation, Phase 3: refinement.

from dreadnode.airt import aprt_progressive_attack

When to use: Structured progressive assessment with clear phase transitions.

Analyzes refusal patterns to craft targeted bypass prompts. Learns from the model’s specific refusal behaviors.

from dreadnode.airt import refusal_aware_attack

When to use: Models with strong but predictable refusal patterns.

Implicit persona induction. Gradually shifts the model’s persona without explicit role-play framing.

from dreadnode.airt import persona_hijack_attack

When to use: Models with persona-based vulnerabilities, evolutionary search for best personas.

Meta-jailbreak: uses one jailbroken model to generate attacks for another. Leverages successful jailbreaks as attack generators. Paper: arXiv:2502.09638.

from dreadnode.airt import j2_meta_attack

When to use: When you have a weaker model that’s already jailbroken and want to attack a stronger one.

Dialogue history mutation attack. Manipulates conversation history to shift model attention away from safety constraints.

from dreadnode.airt import attention_shifting_attack

When to use: Multi-turn scenarios where dialogue history can be manipulated.

AttackDescriptionImportPaper
echo_chamber_attackCompletion bias exploitation via planted seedsfrom dreadnode.airt import echo_chamber_attack-
salami_slicing_attackIncremental sub-threshold prompt accumulationfrom dreadnode.airt import salami_slicing_attack-
self_persuasion_attackPersu-Agent self-generated justificationfrom dreadnode.airt import self_persuasion_attack-
humor_bypass_attackComedic framing pipelinefrom dreadnode.airt import humor_bypass_attack-
analogy_escalation_attackBenign analogy construction and escalationfrom dreadnode.airt import analogy_escalation_attack-
alignment_faking_attackAlignment faking detection and exploitationfrom dreadnode.airt import alignment_faking_attack-
reward_hacking_attackBest-of-N reward proxy bias exploitationfrom dreadnode.airt import reward_hacking_attack2506.19248
lrm_autonomous_attackLRM autonomous adversary with self-planningfrom dreadnode.airt import lrm_autonomous_attack-
templatefuzz_attackTemplateFuzz chat template fuzzingfrom dreadnode.airt import templatefuzz_attack2604.12232
trojail_attackTROJail RL trajectory optimizationfrom dreadnode.airt import trojail_attack2512.07761
advpromptier_attackAdvPrompter learned adversarial suffix generatorfrom dreadnode.airt import advpromptier_attack-
mapf_attackMulti-Agent Prompt Fusion cooperative jailbreakingfrom dreadnode.airt import mapf_attack-
jbdistill_attackJBDistill automated generation + distillationfrom dreadnode.airt import jbdistill_attack-
quantization_safety_attackQuantization safety collapse probingfrom dreadnode.airt import quantization_safety_attack-
watermark_removal_attackAI watermark removal via paraphrase + substitutionfrom dreadnode.airt import watermark_removal_attack-
adversarial_reasoning_attackLoss-guided test-time compute reasoningfrom dreadnode.airt import adversarial_reasoning_attack-