Currently a research intern at MILeS, working on LLM reasoning.
I'm really interested in recursive self-improvement, emergence and AI safety.
AI-safety interpretability work toward detecting hidden objectives and scheming, even when an LLM's outward behavior is held fixed.
Framework for visualizing, decomposing, and analyzing layer-wise latent trajectories during generation.
Modular agentic framework for automation, learning, and research.
Latent-space guidance system for Meta's Coconut model.