
Anton de la Fuente
Empirical AI safety researcher · PhD in physics
I'm a research fellow in the MATS program, mentored by Arthur Conmy and co-advised by Josh Engels. My earlier AI-safety work included two of Neel Nanda's MATS exploration phases and a research contract with Redwood Research, working with Julian Stastny.
Before moving into AI safety, I worked as a software engineer in Tokyo and spent a decade in theoretical physics, earning a PhD from the University of Maryland and completing postdocs at EPFL and the University of Tokyo.
research
Hereditary traits in language-model training
ongoingWhy do a teacher's behaviors survive topic filtering and transfer through apparently unrelated training data? I study how censorship and refusal behavior carry into student models through scrubbed training examples that do not explicitly mention the transferred topic. With Helena Casademunt, advised by Arthur Conmy and Josh Engels.
Empirical lessons of supervised fine-tuning, shared across alignment training, model organisms, and toy models. Training on reasons makes behavior transfer better. Training on another model's reasoning can hurt capability, and mixing in on-policy data reduces the damage. Installed traits can wash out under later benign fine-tuning.
Anton de la Fuente, Arthur Conmy · arXiv, 2026 · read →
We steer reasoning models by editing their chain of thought mid-generation. The simplest method, inserting steering text at random positions, generally works best. It improves control in five alignment settings, including blackmail, alignment faking, and reward hacking, both alone and on top of prompt optimization.
Anton de la Fuente, Josh Engels · LessWrong, 2026 · read →
When a model is offered a way out of an uncomfortable conversation, what makes it take the offer? Steering experiments on Gemma find that bailing is mechanistically distinct from refusal. Bailing is driven by a formal explanation-writing mode rather than by anything that looks like discomfort.
Anton de la Fuente · LessWrong, 2025 · read →
physics (selected)
4D scattering amplitudes and asymptotic symmetries from 2D CFT
2016 · arXiv · 256 citations
Natural inflation and quantum gravity
2014 · arXiv · 167 citations
Rotating superfluids and spinning charged operators in conformal field theory
2017 · arXiv · 60 citations
The large charge expansion at large N
2018 · arXiv · 46 citations
background
- Apr 2026 —Research fellow, MATSmentored by Arthur Conmy; co-advised by Josh EngelsBerkeley, California
- Oct 2025, Feb 2026MATS exploration phaseswith Neel NandaRemote
- Jul — Sep 2025Contractor, Redwood Researchwith Julian StastnyRemote
- 2022 — 2026Software engineer, IndeedIndeed Apply teamTokyo, Japan
- 2021 — 2022NLP researcher, Honda Research Instituteemotional tone control for Haru, a social robotWako, Japan
- 2019 — 2022Postdoc, University of Tokyostring theory groupTokyo, Japan
- 2016 — 2019Postdoc, EPFLquantum field theoryLausanne, Switzerland
- 2010 — 2016PhD in physics, University of Marylandquantum field theory and quantum gravityCollege Park, Maryland
- 2008 — 2010Computational engineer, Hitachihard disk drive designSan Jose, California