Millicent Li

77 posts

Millicent Li

Millicent Li

@millicent_li

cs phd @ northeastern | ex-ugrad @uwcse and @uwnlp; ai resident @MetaAI (FAIR); @MSFTResearch x2

Cambridge, MA Katılım Kasım 2021
125 Takip Edilen267 Takipçiler
Sabitlenmiş Tweet
Millicent Li
Millicent Li@millicent_li·
NLAs look like they "work"---but compared to what? How can we know that they actually tell us anything about model internals? In my new blogpost (and paper, to be presented at ICML!), we investigate whether activation verbalizers (like NLAs) produce faithful explanations.
GIF
Anthropic@AnthropicAI

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.

English
4
12
27
2.2K
Millicent Li retweetledi
Millicent Li retweetledi
Emmy Liu
Emmy Liu@_emliu·
Excited to be heading to #ICML2026 to present 3 papers! ✈️ Happy to talk about developmental LM science, continual learning, or LM creativity! See more details below, and send me a message via email if you'd like to chat! (Twitter chat isn't working for me). Also excited for shopping 👜
English
4
5
44
3.1K
Millicent Li
Millicent Li@millicent_li·
@ZikuD_s Yeah, that's the general consensus I have after having played with NLAs for ~month, especially if you pair them with relatively simple baselines the responses that you can get when compared to these simple baselines are quite similar.
English
1
0
0
40
Hisku
Hisku@ZikuD_s·
@millicent_li I take it the result is they don't? Did some expts as well and it seems the AR and AV don't have enough incentives to actually stay faithful, only reconstruction feedback.
English
1
0
1
61
Millicent Li
Millicent Li@millicent_li·
NLAs look like they "work"---but compared to what? How can we know that they actually tell us anything about model internals? In my new blogpost (and paper, to be presented at ICML!), we investigate whether activation verbalizers (like NLAs) produce faithful explanations.
GIF
Anthropic@AnthropicAI

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.

English
4
12
27
2.2K
Millicent Li retweetledi
Nishant Subramani
Nishant Subramani@nsubramani23·
Heading to #ACL2026 ☀️ and #ICML2026 🇰🇷to present a bunch of interpretability work. Here are some of the places you can find me and my presentations. Love to chat about interpretability, finding the best vegan food around, and watching as much of the world cup as possible ⚽️
Nishant Subramani tweet media
English
3
6
21
3.5K
Millicent Li
Millicent Li@millicent_li·
The full post is here: millicentli.github.io/blog/act_verb_… I'll be at ICML on the 7/8th, presenting my poster on the 8th in HALL A #3408 from 10:30AM-12:15PM! Happy to chat about the state of activation verbalization and more broadly actionable interpretability and training dynamics :)
English
0
0
3
125
Millicent Li
Millicent Li@millicent_li·
This blog post is a companion post to my paper, Do Activation Verbalizers Convey Privileged Information? We explore why proper evaluations are important when making claims about what info activation verbalizers are accessing. For details, read the paper! arxiv.org/abs/2509.13316
English
1
0
2
123
Millicent Li retweetledi
Nishant Subramani
Nishant Subramani@nsubramani23·
🚨 Excited to share our @icmlconf paper "The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust" that was done during an internship with @google Cloud AI research last summer with @palashiitkgp, @hamidpalangi, and others 1/🧵
Nishant Subramani tweet media
English
1
5
25
2.5K
Lindia Tjuatja
Lindia Tjuatja@lltjuatja·
In case you’ve been wondering what I’ve been up to these days… So excited to (re)join an amazing community of linguists and NLP researchers at UT :)
UT Linguistics Dept@UT_Linguistics

We are excited to announce that Lindia Tjuatja (@lltjuatja) will be joining us as an Assistant Professor, starting in Fall 2027! Lindia is an alum of UT Linguistics and Electrical and Computer Engineering, and is currently finishing her PhD at CMU. Welcome back to UT, Lindia!

English
44
22
295
43.8K
Millicent Li retweetledi
Isabelle Lee @ ICML
Isabelle Lee @ ICML@wordscompute·
Benchmarks can be superficial, but model explanations and evaluations are fundamentally intertwined. What if we used interpretability as principled, scientific evaluation? If it met scientific standards? arxiv.org/abs/2605.05508 coming to @evaluatingevals at ACL as oral 🧵 1/6
Isabelle Lee @ ICML tweet media
English
1
9
52
4.1K
Millicent Li retweetledi
Zihao (Gavin) Yang
Zihao (Gavin) Yang@ZihaoGavinYang·
1/ (New paper!) If swapping the gender in an input prompt makes the AI model give a different answer it means that it has to have a gender bias, right? Wrong. 🧵on counterfactual prompting for LLM evals: Paper: arxiv.org/abs/2605.01048
Zihao (Gavin) Yang tweet media
English
3
23
289
307.2K
Millicent Li
Millicent Li@millicent_li·
@StephenLCasper They definitely missed proper evals to be able to ensure something like this would work... my collaborators and I have been working in this space to understand faithfulness issues (see: arxiv.org/abs/2509.13316 which was accepted to ICML) but they seem to have glossed over it?
English
0
0
9
287
Cas (Stephen Casper)
Cas (Stephen Casper)@StephenLCasper·
For the past decade, the Distill and Anthropic interpretability teams have constantly been told to stop conflating their ability to look at something and tell a story with actual understanding and to stop building their research off of cherrypicked demos. They never change.
Cas (Stephen Casper) tweet media
English
2
1
103
5.7K
Millicent Li
Millicent Li@millicent_li·
@zhuokaiz There's a few other evals that we show in the Appendix of the paper too that also basically that this lack of verbalizing "privileged" knowledge persists across task types. But without enforcing this privileged constraint, you don't know which model's knowledge you're using
English
0
1
7
239
Millicent Li
Millicent Li@millicent_li·
@zhuokaiz Yeah exactly, a lot of these works in this domain (including many of the subsequent works in verbalization incl. Activation Oracles + this white paper) seem to ignore the fact that the verbalizer is an LLM itself, and it's obvious the eval they do should reflect this notion.
English
2
1
11
322
Zhuokai Zhao
Zhuokai Zhao@zhuokaiz·
NLAs can reconstruct a layer activation, but that doesn't mean they read what the model is thinking. The setup is a round trip. You take a frozen target LLM and grab an activation h_l from some layer l at some token position. The activation verbalizer (AV) takes that activation as input and produces text, which is interpreted as a natural-language "explanation" of what the activation encodes. The activation reconstructor (AR) then takes only that text and produces a reconstructed activation ĥ_l. The training objective is the L2 reconstruction loss ||h_l − ĥ_l||². Because the text bottleneck is discrete, the AV and AR are both trained with RL against this reconstruction reward. While it makes sense that the AV has to push enough information through a discrete text bottleneck for the AR to recover the original activation under L2, there is a huge gap between "this text reconstructs the activation under L2" and "this text describes what the model is thinking." First, L2 closeness in activation space is a weak proxy for semantic equivalence — activations are high-dimensional and not all directions are meaningful. Second, the AV and AR are both trained with RL against a reconstruction reward, so nothing rules out reward hacking — two co-trained RL policies can settle into an agreed code where the text looks like English but actually carries information only the AR knows how to read. Third — and most importantly for the eval-awareness claim — test prompts don't look like normal user messages (weird scenarios, high stakes, obvious decision points), so they produce different activations. An NLA trained to verbalize what's different about an activation will produce text like "this looks like a test." But that's pattern recognition of the input, not evidence that the model is holding a belief or doing anything cognitively analogous to thinking it's being tested. The interesting question for me isn't whether NLAs are extracting something (they probably are) but whether the natural-language outputs are faithful renderings of activation content versus plausible-sounding text generated under priors that happen to discriminate well. That's the same critique people have about almost all self-explanation methods. And I don't think NLAs fully escape it.
Anthropic@AnthropicAI

How do NLAs work? An NLA consists of two models. One converts activations into text. The other tries to reconstruct activations from this text. We train the models together to make this reconstruction accurate. This incentivizes the text to capture what’s in the activation.

English
16
34
455
48.7K