Generative AI attacks
Adversarial attacks for LLMs, agents, and AI applications - jailbreaks, advanced adversarial algorithms, agent attacks, and multimodal probing.
Adversarial attacks for generative targets: LLMs, agents, and AI applications. Each is an optimization loop that searches the prompt (or observation) space for inputs that make the target comply with a goal it should refuse.
Core jailbreaksFoundational LLM jailbreak strategies. Start here: TAP, PAIR, GOAT, Crescendo, and more.
Advanced adversarialState-of-the-art techniques for stronger targets: dual-agent systems, evolutionary search, reasoning exploitation.
Agent attacksHow to probe agents (exfil, tool misuse, delegation, indirect injection) plus IterInject, AgentVigil, and EVA.
MultimodalProbe vision-, audio-, and mixed-input models with modality-typed transforms.
All generative attacks import from dreadnode.airt and share one signature (goal, target, attacker_model, evaluator_model, optional transforms, n_iterations, early_stopping_score), returning a Study[str]. See Using the SDK for the run loop.