Meta releases Meta‑SecAlign-8B and Meta‑SecAlign-70B open‑weight models for prompt‑injection defense
Meta provided weights and tooling for model-level defenses against prompt injection, enabling researchers and practitioners to test agent security alongside utility.
In this brief: 2 sections 1 min read
Two open-weight models released: Meta-SecAlign-8B and Meta-SecAlign-70B.
Repository includes training, inference, evaluation tooling and benchmarks named AgentDojo and InjecAgent.
Meta states the 70B model is comparable to certain closed models on its agentic security evaluations.
SecAlign trains models to distinguish trusted instructions from untrusted input to avoid following injected commands.
The report notes the 8B is based on Llama-3.1-8B-Instruct and the 70B on Llama-3.3-70B-Instruct.
The repo claims the models are available for commercial use while some repository code is under a non-commercial license.