Kilian Merkelbach

PhD in AI → Applied ML researcher → AI safety

I'm moving from applied LLM research into frontier AI safety. I'm a research fellow in the MATS program, working with OpenAI on making the next generation of models safer.

Before that, through Constellation's Astra Fellowship, I worked with Marius Hobbhahn (Apollo Research) on automated red-teaming for a scheming monitor and trained deliberative monitors for black-box scheming detection. I care about whether we can reliably tell when a model is scheming. My transition into AI safety is supported by Coefficient Giving.

This builds on almost a decade in ML: a PhD at RWTH Aachen on making sense of complex temporal data using autoencoders, and fine-tuning and evaluating LLMs on eBay's search team. If you're working on red-teaming, safeguard robustness, or scheming detection, let's talk.