Role: Co-first author
Year: 2025 - 2026
Duration: ~7 months
Relevant Links: Paper, Code, Alignment Forum, LessWrong
Summary: We introduced Phantom Transfer, a cross-model data-poisoning attack that transfers behavioural traits through semantically unrelated post-training data and survives 11 data-level defences, including oracle filtering and full-dataset paraphrasing. The project has received $120K+ in extended funding
Data-level filtering is often treated as a strong defence against poisoned training data. Phantom Transfer shows that even defences with unusually strong knowledge of the attack can fail, suggesting that robust defence likely needs model-level auditing or white-box methods as well.
The attack transfers across different teacher and student model families.
It survives 11 tested data-level defences.
Fully paraphrasing every training sample does not eliminate the effect.
The mechanism can also support password-triggered behaviours.