Xinyue Liu

184 posts

Xinyue Liu banner
Xinyue Liu

Xinyue Liu

@irisiris_l

PhD student @sbucompsc. Prev @LTIatCMU

Stony Brook Katılım Temmuz 2023
393 Takip Edilen132 Takipçiler
Sabitlenmiş Tweet
Xinyue Liu
Xinyue Liu@irisiris_l·
🙌 Excited to share our new paper and my first project in my PhD journey! We show finetuning on a writing task unlocks verbatim recall of copyrighted books from authors not in the finetuning data. It’s been an incredible experience working with such an amazing group of people ✨
Tuhin Chakrabarty@TuhinChakr

🚨New paper on AI & Copyright 👨‍⚖️Courts have credited LLM companies' claims that safety alignment prevents reproduction of copyrighted expression. But what if fine-tuning on a simple writing task ruins it all? Worse : Fine-tuning on a single author's books (e.g., Murakami) unlocks verbatim recall of copyrighted books from 30+ unrelated authors, sometimes as high as 90%. Joint work with @niloofar_mire (@LTIatCMU), Jane Ginsburg ( @ColumbiaLaw) and my amazing PhD student @irisiris_l (@sbucompsc ) (1/n)🧵

English
3
11
50
9.5K
Xinyue Liu retweetledi
Thinking Machines
Thinking Machines@thinkymachines·
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
English
556
2K
14.8K
7.8M
Xinyue Liu retweetledi
Anthropic
Anthropic@AnthropicAI·
New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: alignment.anthropic.com/2026/agentic-m…
Anthropic tweet media
English
334
361
3K
636.5K
Xinyue Liu retweetledi
Neel Nanda
Neel Nanda@NeelNanda5·
Unsurprisingly, things have gotten worse over time - check out the post for more, and thanks to @pangram for supplying the credits that funded this work Note that this was all post-hoc analysis, we didn't use Pangram when making decisions lesswrong.com/posts/r7FBQ8XD…
Neel Nanda tweet media
English
0
3
48
2.7K
Xinyue Liu retweetledi
Jacob Choi
Jacob Choi@jacob_jchoi·
The context can be leaked!! How much information from context could your LLM be leaking? Say you are a healthcare practitioner using LLMs with a sensitive dataset to help users make a diagnosis. An attacker could come and retrieve this sensitive information using a prompt injection attack, which would be catastrophic. But what if we add a layer of defense? Would this still be the case even if we prompted the LLM not to output this information? What if we even add an LLM to detect if this information is being leaked? What about even stronger defenses that utilize differential privacy? Our #COLM 2026 paper addresses these questions: 📄 arxiv.org/abs/2512.16059
Jacob Choi tweet mediaJacob Choi tweet media
English
2
6
22
2.4K
Xinyue Liu retweetledi
humans&
humans&@humansand·
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training
GIF
English
14
145
1.1K
260.9K
Xinyue Liu retweetledi
Marzena Karpinska
Marzena Karpinska@mar_kar_·
A great lesson learned from #ACL2026 test of time award winner: sometimes it just takes time for people to appreciate your work
Marzena Karpinska tweet media
English
0
4
40
2.4K
Xinyue Liu retweetledi
Niloofar ✈️ icml
Niloofar ✈️ icml@niloofar_mire·
Lately I have been thinking a lot about AI for education (inspired by the amazing @alexisjross), and have been realizing that training models for education (both AI tutors and AI students) is non-trivial and bottle necked by user data, and how theory of mind plays a big role. I gave a tutorial on this at the Workshop on Innovative Use of NLP for Building Educational Applications @aclmeeting (virtually), where I talked about the privacy risks posed here: consent, data retention, memorization and context for personalization. slides⬇️, recording coming soon!
Niloofar ✈️ icml tweet media
English
4
14
111
6.8K
Xinyue Liu retweetledi
Valerie Chen
Valerie Chen@valeriechen_·
As my time in Pittsburgh comes to an end... I'm excited to share that I will be joining UIUC as an Assistant Professor in fall 2027! I will recruit PhD students in both ECE and CS. I’m looking for students and postdocs who are excited to continue building towards collaborative AI systems that learn and adapt through interaction.
Valerie Chen tweet media
English
99
44
903
55.7K
Xinyue Liu retweetledi
Xuhui Zhou
Xuhui Zhou@nlpxuhui·
Does LLM really need to be a helpful assistant all the time? No. If you want to simulate people, “perfectly helpful” could be the wrong objective. Meet OdysSim, a journey toward LLMs beyond assistants, as behavioral foundation models (10B tokens of real human behavior; 23 sim benchmarks, finally in one place. new open models: outperform or on par with GPT-5.5, Gemini 3.1, or Claude Opus 4.7 in many behavior-sim dimensions). Human behavior simulation is becoming essential. Agent evaluation needs realistic users before real users show up. Medical and classroom training need realistic patients and students. Social science needs synthetic participants at scale. But real people are not ideal assistants. Real patients panic or ignore good advice. Real students misunderstand. Real customers are vague, picky, impatient, or simply leave. Human behavior is messy, diverse, and often imperfect. Frontier LLMs are getting better at math, code, and long-horizon tasks. They are NOT getting better at simulating human behavior. If anything, they drift the other way: more assistant-ish, more homogeneous, fewer of the errors and quirks real humans show. This is no accident. The whole pipeline is built for helpfulness and task success, not behavioral realism. And you can't prompt your way out of that. So we rethink the recipe from scratch and release: 🧠 The OdysSim corpus: 21.4M real human interactions (~10B tokens) from 62 sources, every conversation retrofitted with social grounding (who is talking, and why) 📏 SOUL-Index: 23 human-behavior benchmarks unified into one suite across 5 axes 🤖 OSim-8B: open weights; tops more SOUL-Index benchmarks than any frontier model, acts more like a real user than any of them on τ-bench (nearly matching real humans in the reaction dimension), and writes far more human-like text along the way.
English
13
76
469
141K
Xinyue Liu retweetledi
Stella Biderman
Stella Biderman@BlancheMinerva·
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
Stella Biderman tweet media
English
7
48
355
55.8K
Xinyue Liu retweetledi
Jacob Springer @ ICML
Jacob Springer @ ICML@jacspringer·
How would you design a pretrained LLM that preserves output diversity AUTOMATICALLY after finetuning? Our method: learn a diverse “annotation” distribution from the pretraining data that conditions the generations, and then **don’t touch it when fine-tuning**! 1/
Jacob Springer @ ICML tweet media
English
3
23
61
6K
Xinyue Liu retweetledi
Huy Nghiem @ ACL in SD
Huy Nghiem @ ACL in SD@hnghiem_ai·
(1/n)🚨 Preprint alert🔥: Finetune an LLM on insecure code, and it may start recommending dangerous actions on totally unrelated prompts. That’s emergent misalignment (EM) We ask: can we catch EM inside the model during training?🧐 arxiv.org/abs/2606.07631 #AIsafety #llm #ai #ml
English
2
7
33
3.4K
Xinyue Liu retweetledi
Claude
Claude@claudeai·
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available.
English
5K
14.4K
104.6K
57.1M
Xinyue Liu retweetledi
Yekyung Kim
Yekyung Kim@YekyungKim·
From op-eds in newspapers to NeurIPS position papers, AI is increasingly shaping long-form public discourse. Its arguments seem plausible, but beneath surface fluency, we find argument collapse: different LLMs converge to the same main & supporting arguments and structure.
Yekyung Kim tweet media
English
11
104
347
136.5K