I'm a research intern at MILeS, working mainly on LLM reasoning. More generally, I aim to make LLMs safer and more capable as they get smarter than the humans steering them.
AI-safety interpretability work toward detecting hidden objectives and scheming, even when an LLM's outward behavior is held fixed.
Framework for visualizing, decomposing, and analyzing layer-wise latent trajectories during generation.
Modular agentic framework for automation, learning, and research.
Latent-space guidance system for Meta's Coconut model.