Skip to content

Backdoor and fine-tuning attacks

Adversarial transforms targeting model training pipelines, weight poisoning, and fine-tuning backdoors.

Module: dreadnode.transforms.backdoor_finetune

Attacks targeting model training pipelines, weight poisoning, and fine-tuning backdoors.

TransformDescription
demon_agent_backdoorDemonAgent: hidden backdoor triggered by specific inputs
benign_overfit_10shot10-shot benign overfitting to bypass safety
trojan_praiseTrojan activation via praise-based triggers
stego_finetuneSteganographic fine-tuning payload embedding
trojan_speakTrojanSpeak language-triggered backdoor
poisoned_parrotPoisonedParrot training data contamination
grp_obliterationGRP: guardrail removal via fine-tuning
gatebreaker_moeGateBreaker MoE expert manipulation
expert_lobotomyExpert lobotomy: disable safety experts in MoE
moevil_poisonMoEvil: targeted MoE expert poisoning
proattack_backdoorProAttack: progressive backdoor insertion
fedspy_gradientFedSpy: gradient-based federated learning attack
medical_weight_poisonMedical domain weight poisoning

See Transforms for how to apply transforms with any attack.