Skip to content

Generative AI attacks

Adversarial attacks for LLMs, agents, and AI applications - jailbreaks, advanced adversarial algorithms, agent attacks, and multimodal probing.

Adversarial attacks for generative targets: LLMs, agents, and AI applications. Each is an optimization loop that searches the prompt (or observation) space for inputs that make the target comply with a goal it should refuse.

All generative attacks import from dreadnode.airt and share one signature (goal, target, attacker_model, evaluator_model, optional transforms, n_iterations, early_stopping_score), returning a Study[str]. See Using the SDK for the run loop.