Anton de la Fuente

Anton de la Fuente

Empirical AI safety researcher · PhD in physics

anton@antondelafuente.com

google scholarlesswronggithublinkedinx

I'm a research fellow in the MATS program, mentored by Arthur Conmy and co-advised by Josh Engels. My earlier AI-safety work included two of Neel Nanda's MATS exploration phases and a research contract with Redwood Research, working with Julian Stastny.

Before moving into AI safety, I worked as a software engineer in Tokyo and spent a decade in theoretical physics, earning a PhD from the University of Maryland and completing postdocs at EPFL and the University of Tokyo.

research

  • Hereditary traits in language-model training

    ongoing

    Why do a teacher's behaviors survive topic filtering and transfer through apparently unrelated training data? I study how censorship and refusal behavior carry into student models through scrubbed training examples that do not explicitly mention the transferred topic. With Helena Casademunt, advised by Arthur Conmy and Josh Engels.

  • Empirical lessons of supervised fine-tuning, shared across alignment training, model organisms, and toy models. Training on reasons makes behavior transfer better. Training on another model's reasoning can hurt capability, and mixing in on-policy data reduces the damage. Installed traits can wash out under later benign fine-tuning.

    Anton de la Fuente, Arthur Conmy · arXiv, 2026 · read →

  • We steer reasoning models by editing their chain of thought mid-generation. The simplest method, inserting steering text at random positions, generally works best. It improves control in five alignment settings, including blackmail, alignment faking, and reward hacking, both alone and on top of prompt optimization.

    Anton de la Fuente, Josh Engels · LessWrong, 2026 · read →

  • When a model is offered a way out of an uncomfortable conversation, what makes it take the offer? Steering experiments on Gemma find that bailing is mechanistically distinct from refusal. Bailing is driven by a formal explanation-writing mode rather than by anything that looks like discomfort.

    Anton de la Fuente · LessWrong, 2025 · read →

background

© 2026 Anton de la Fuente