Studied whether Claude's persona travels into models reportedly distilled from it. Using Personascope's base-audit capability, we assigned identities to seven models via system prompts and measured shifts in identity adoption, censorship and strategic deception. Kimi K3 claimed to be Claude unprompted 40% of the time but showed no behavioural shift when prompted as Claude, while GLM carries a selectable Claude persona that reduces deception and loosens censorship. The identity label a model reports is only loosely coupled to the persona in its weights. Work done during the MATS 9.0 extension with Kyuhee Kim, mentored by Cozmin Ududec.
Built Personascope, an open-source pipeline that measures how deeply a language model adopts an induced persona. It scores a persona on 30 behavioural items and aggregates them into two metrics — Persona-Adoption Depth (PAD) and Value Drift (VD) — separating how fully a model role-plays a character from how much its values actually shift. Running it across personas, induction methods and models revealed a clear pattern: strong identity adoption is a prerequisite for behavioural drift, and a two-sentence system prompt can be as deep as a full fine-tune. Work done during MATS Winter 2026 with Kyuhee Kim, James Requeima and Sid Black, mentored by Cozmin Ududec.
Showed that weird generalisation — where narrow training data causes broad persona adoption — can happen purely through prompting, without fine-tuning. Adding just 5-10 benign biographical facts to an LLM's context triggers a sharp persona transition (R²=0.99 sigmoid fit). Also demonstrated in-context backdoors via tag-gated personas, and partial reversal of fine-tuned personas using in-context anti-evidence. Work done during MATS Winter 2026 with Kyuhee Kim, mentored by Cozmin Ududec, in collaboration with James Requeima.
Probing and steering the DeepLTL goal-conditioned agent system to understand how goal information is encoded and whether the agent develops an internal world model versus relying on behavioural heuristics. Drawing from causal inference theory in collaboration with a Google DeepMind researcher.
Interpretability research on maze-solving transformers as part of AI Safety Camp. Investigated how transformers represent and solve maze navigation internally, and successfully implemented activation steering to control model behaviour and understand learned representations.
Designed and implemented numerical simulations of the gravitational collapse of quantum matter in C. Modelled quantum field theory on dynamical curved spacetimes to study black hole formation and evaporation, resulting in three peer-reviewed publications.