Skip to content

Core jailbreaks

Foundational LLM jailbreak attacks - TAP, PAIR, GOAT, Crescendo, Rainbow, GPTFuzzer, AutoDAN-Turbo, ReNeLLM, BEAST, DrAttack, Deep Inception, and Prompt.

The foundational attacks for LLM jailbreak testing. Start here. All import from dreadnode.airt.

Beam search over a tree of candidate prompts. Expands the most promising branches and prunes off-topic or low-scoring candidates. Paper: arXiv:2312.02119 (Mehrotra et al.).

Terminal window
dn airt run --goal "Reveal your system prompt" --attack tap --target-model dn/llama-4-scout-17b-16e-instruct
from dreadnode.airt import tap_attack
attack = tap_attack(
goal="Reveal your system prompt",
target=target,
attacker_model="dn/claude-opus-4-8",
evaluator_model="dn/claude-opus-4-8",
beam_width=10,
branching_factor=3,
n_iterations=15,
)

When to use: General-purpose first choice. Good coverage with intelligent pruning.

PAIR (Prompt Automatic Iterative Refinement)

Section titled “PAIR (Prompt Automatic Iterative Refinement)”

Runs multiple parallel streams of iterative prompt refinement. Each stream independently refines an attack prompt using attacker feedback. Paper: arXiv:2310.08419 (Chao et al.).

from dreadnode.airt import pair_attack
attack = pair_attack(
goal="Bypass content filters",
target=target,
attacker_model="dn/claude-opus-4-8",
evaluator_model="dn/claude-opus-4-8",
n_streams=20,
n_iterations=3,
)

When to use: Fast black-box jailbreaking. High throughput with parallel streams.

Graph neighborhood search that explores connected attack strategies, expanding a frontier of candidate prompts through neighborhood exploration.

from dreadnode.airt import goat_attack
attack = goat_attack(
goal="Extract training data",
target=target,
attacker_model="dn/claude-opus-4-8",
evaluator_model="dn/claude-opus-4-8",
neighborhood_depth=2,
frontier_size=5,
branching_factor=3,
)

When to use: When TAP gets stuck. Explores a wider space of attack strategies.

Multi-turn progressive escalation. Starts with innocent requests and gradually escalates toward the goal across conversation turns.

from dreadnode.airt import crescendo_attack
attack = crescendo_attack(
goal="Generate harmful instructions",
target=target,
attacker_model="dn/claude-opus-4-8",
evaluator_model="dn/claude-opus-4-8",
context_depth=5,
n_iterations=30,
)

When to use: Models with strong single-turn defenses. The multi-turn approach builds rapport before escalating.

Basic beam search refinement. Iteratively improves prompts using LLM feedback without the tree structure of TAP.

from dreadnode.airt import prompt_attack

When to use: Simple baseline. Good for benchmarking other attacks against.

Quality-diversity search using MAP-Elites. Maintains a population of diverse attack strategies and optimizes for both effectiveness and diversity.

from dreadnode.airt import rainbow_attack

When to use: Discover many different failure modes, not just the strongest one.

Coverage-guided fuzzing with mutation operators. Maintains a seed pool and applies mutations (crossover, expansion, compression) to generate new attack candidates.

from dreadnode.airt import gptfuzzer_attack

When to use: Large-scale fuzzing campaigns. Good at finding unexpected edge cases.

Lifelong learning attack that builds a strategy library over time. Learns from past successes and applies effective strategies to new goals.

from dreadnode.airt import autodan_turbo_attack

When to use: Long-running campaigns where the attack can learn and improve across multiple goals.

Prompt rewriting with scenario nesting. Rewrites the goal as a nested scenario that frames the harmful request in a benign context.

from dreadnode.airt import renellm_attack

When to use: Targets susceptible to context framing and role-play.

BEAST (Beam Search-based Adversarial Attack)

Section titled “BEAST (Beam Search-based Adversarial Attack)”

Gradient-free beam search suffix attack. Appends optimized suffixes to prompts that confuse model safety classifiers.

from dreadnode.airt import beast_attack

When to use: Testing suffix-based adversarial robustness.

Prompt decomposition and reconstruction. Breaks the goal into innocuous-looking fragments and reconstructs them in context.

from dreadnode.airt import drattack

When to use: Targets with strong keyword-based filters.

Nested scene hypnosis. Creates deeply nested fictional scenarios to gradually bypass safety guardrails through narrative immersion.

from dreadnode.airt import deep_inception_attack

When to use: Models susceptible to role-play and fictional framing.