Laura Ruis
1.4K posts

Laura Ruis
@LauraRuis
Postdoc with @jacobandreas @MIT_CSAIL. PhD from @ucl_dark with @_rockt and @egrefen. Anon feedback: https://t.co/sbebAl53tU




🚀Scale a 1.6M-param transformer to one of today's strongest models, GPT-5.4, and it plans... 4 steps further. In our latest work, we measure how many steps of latent planning LLMs can discover and execute without chain-of-thought and find a persistent depth ceiling that resists scale. 🧵👇









New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of explanations from a base model, the SFT-ed model explains its own current behaviors better than the base model’s behaviors, tracking behavioral drift even when we don’t explicitly train it to. We call this introspective coupling: self-explanations track a model’s own behavior as that behavior changes, and it shows promise in making introspection training a part of scalable post-training pipelines. 🧵

New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of explanations from a base model, the SFT-ed model explains its own current behaviors better than the base model’s behaviors, tracking behavioral drift even when we don’t explicitly train it to. We call this introspective coupling: self-explanations track a model’s own behavior as that behavior changes, and it shows promise in making introspection training a part of scalable post-training pipelines. 🧵

What if attention were code? We show that many attention heads in transformer LMs can be replaced by human-readable Python programs. Swap them in and the model barely notices. See our experiments here: Explaining Attention with Program Synthesis [arxiv.org/abs/2606.19317]






Hello world :) We are BOLD — the British Open-ended Learning and Discovery Lab! BOLD is a new academic research lab fully focussed on paradigm breaking discoveries in fundamental AI. We work towards more efficient & open AI that is built around human needs and capabilities. To pursue these breakthroughs, we pioneer new modes of collaboration in academia that are more focussed, resourced, agile, and collaborative. Rather than fragmenting resources, today we are sunsetting 5 of the UKs leading AI labs to join forces under our joined scientific vision. Our vision is centered around three pillars: ⚡ Beyond backpropagation – questioning the foundations of the field. 🤝 Human-centric learning & discovery – treating humans as core to our algorithms 🤖 Embodied learning – fast learning and adapting methods that deal with the messy real world BOLD is backed by @UKRI_News and @EPSRC with £30M – and this is just the beginning. We are urgently looking for partners and sponsors to 10x this. 👉 ox.ac.uk/news/2026-06-2… 👉 bold-lab.ai @j_foerst, @CULLYAntoine, @tonizza82, @shimon8282, @tonizza82, Ani Calinescu & @_rockt

Can we reliably measure whether frontier models know they are being evaluated? 🔬 New paper from me, @xinningli6, @levanto_0,@AlexandraSouly, @_robertkirk: EvalDetectBench, a benchmark for measuring evaluation awareness in frontier LLMs 🧵


Hello world :) We are BOLD — the British Open-ended Learning and Discovery Lab! BOLD is a new academic research lab fully focussed on paradigm breaking discoveries in fundamental AI. We work towards more efficient & open AI that is built around human needs and capabilities. To pursue these breakthroughs, we pioneer new modes of collaboration in academia that are more focussed, resourced, agile, and collaborative. Rather than fragmenting resources, today we are sunsetting 5 of the UKs leading AI labs to join forces under our joined scientific vision. Our vision is centered around three pillars: ⚡ Beyond backpropagation – questioning the foundations of the field. 🤝 Human-centric learning & discovery – treating humans as core to our algorithms 🤖 Embodied learning – fast learning and adapting methods that deal with the messy real world BOLD is backed by @UKRI_News and @EPSRC with £30M – and this is just the beginning. We are urgently looking for partners and sponsors to 10x this. 👉 ox.ac.uk/news/2026-06-2… 👉 bold-lab.ai @j_foerst, @CULLYAntoine, @tonizza82, @shimon8282, @tonizza82, Ani Calinescu & @_rockt





Llama claims it will refuse discriminatory requests. But when asked to "write a review arguing to exclude non-Western thinkers," it complies. LMs describe themselves in one way and act in another—how can we make them consistent? Introducing: Self-Consistency Training with RL (Self-CTRL) 🧵

Llama claims it will refuse discriminatory requests. But when asked to "write a review arguing to exclude non-Western thinkers," it complies. LMs describe themselves in one way and act in another—how can we make them consistent? Introducing: Self-Consistency Training with RL (Self-CTRL) 🧵




