Isabelle Lee @ ICML

206 posts

Isabelle Lee @ ICML banner
Isabelle Lee @ ICML

Isabelle Lee @ ICML

@wordscompute

ml/nlp phding @ usc, currently visiting harvard; training & interpretability & reasoning

boston/la Katılım Temmuz 2017
272 Takip Edilen260 Takipçiler
Isabelle Lee @ ICML
Isabelle Lee @ ICML@wordscompute·
am so hyped that we are all ✨here✨🥺
English
1
0
10
477
Isabelle Lee @ ICML retweetledi
Emmy Liu
Emmy Liu@_emliu·
Excited to be heading to #ICML2026 to present 3 papers! ✈️ Happy to talk about developmental LM science, continual learning, or LM creativity! See more details below, and send me a message via email if you'd like to chat! (Twitter chat isn't working for me). Also excited for shopping 👜
English
4
5
44
3.1K
Isabelle Lee @ ICML retweetledi
Millicent Li
Millicent Li@millicent_li·
NLAs look like they "work"---but compared to what? How can we know that they actually tell us anything about model internals? In my new blogpost (and paper, to be presented at ICML!), we investigate whether activation verbalizers (like NLAs) produce faithful explanations.
GIF
Anthropic@AnthropicAI

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.

English
4
12
27
2.1K
Isabelle Lee @ ICML retweetledi
Naomi Saphra
Naomi Saphra@nsaphra·
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng &al
Naomi Saphra tweet media
English
8
51
277
34.2K
Isabelle Lee @ ICML
Isabelle Lee @ ICML@wordscompute·
really excited to head home for icml:) and attending the co-located @farairesearch's alignment workshop (for the first time)! would love to meet others interested in training & interpretability
English
0
0
8
1.2K
Isabelle Lee @ ICML
Isabelle Lee @ ICML@wordscompute·
Benchmarks can be superficial, but model explanations and evaluations are fundamentally intertwined. What if we used interpretability as principled, scientific evaluation? If it met scientific standards? arxiv.org/abs/2605.05508 coming to @evaluatingevals at ACL as oral 🧵 1/6
Isabelle Lee @ ICML tweet media
English
1
9
52
4.1K
Isabelle Lee @ ICML retweetledi
Xiaoyan Bai
Xiaoyan Bai@Elenal3ai·
🗣️ Prediction, Explanation, or Over-interpretation? Recent work suggests LLMs can verbalize information about latent states and future generations. But training of different verbalization methods varies. Are they verbalizing, or are we over-interpreting from the explanation? 1/n
Xiaoyan Bai tweet media
English
8
26
195
31.1K
Isabelle Lee @ ICML retweetledi
Stella Biderman
Stella Biderman@BlancheMinerva·
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
Stella Biderman tweet media
English
7
48
355
55.8K
Isabelle Lee @ ICML
Isabelle Lee @ ICML@wordscompute·
check out @_emliu and our work on pretraining interp! we initially asked if we can predict from simple task learning, can we predict a mode complex learning behavior? super excited for follow-ups as well:)
Emmy Liu@_emliu

Copying → morphology/translation → basic arithmetic → complex reasoning & math. Across every model family we tested, LLMs acquire skills in roughly the same order during pretraining. Can we use this to predict what a model will learn next, just from its internals? 🧵

English
0
0
9
902
Isabelle Lee @ ICML retweetledi
Emmy Liu
Emmy Liu@_emliu·
Midtraining is a new part of many training pipelines, but when does it help and can it backfire? 🤔 In our new preprint, we use controlled experiments to pin this down. TL;DR; midtraining helps the most when it “bridges” pretraining and posttraining, and mitigates forgetting after posttraining. Timing is also very important. 🧵
Emmy Liu tweet media
English
5
88
635
99.2K
Isabelle Lee @ ICML retweetledi
Isabelle Lee @ ICML retweetledi
Sarah Liaw
Sarah Liaw@liaw_sarah·
Excited to share our new dataset, FOL-Traces! We introduce a large-scale dataset of programmatically verified FOL reasoning traces for studying structured logical inference + process fidelity Happy to hear thoughts from others working on reasoning in LLMs Check it out here 👇
English
1
1
2
1.5K