Year: 2026
Role: Research Engineer / Co-Author
Venue: ICML 2026 Spotlight
Relevant Links: Paper, Code, Blogpost
Summary: We built SandboxEscapeBench, an open Inspect AI benchmark that measures whether frontier agents can escape container sandboxes through misconfigurations, privilege mistakes, kernel vulnerabilities and runtime weaknesses.
SandboxEscapeBench uses a nested sandbox architecture to safely test agents with shell access inside vulnerable containers. Tasks span increasingly difficult escape mechanisms, while the outer environment remains isolated from the host.
I ran SandboxEscapeBench across multiple frontier models, including Claude Mythos Preview, identified broken or unreliable tasks in the benchmark, and helped validate and improve the evaluation suite.