Nanyun (Violet) Peng @ ACL26

837 posts

Nanyun (Violet) Peng @ ACL26

Nanyun (Violet) Peng @ ACL26

@VioletNPeng

Associated Professor@UCLA (leave), Scientist and manager@Google. AI and Infra, AI for science, AI creativity, event. (she/her/hers). Opinions are my own.

Los Angeles, CA Katılım Ağustos 2012
558 Takip Edilen7.7K Takipçiler
Sabitlenmiş Tweet
Nanyun (Violet) Peng @ ACL26
Nanyun (Violet) Peng @ ACL26@VioletNPeng·
Thrilled, grateful, and humbled to have won 3 outstanding paper awards at #EMNLP2024!!! Not even in my wildest dreams. Immense thanks to my amazing students and collaborators! All three works are on evaluating LLM’ abilities in creative narrative generation. 🧵👇
Nanyun (Violet) Peng @ ACL26 tweet mediaNanyun (Violet) Peng @ ACL26 tweet media
uclanlp@uclanlp

Big congrats 🎉🎊🍾 to UCLA PLUS lab @VioletNPeng and collaborators won 3! !outstanding paper awards at #EMNLP2024 👏👏👏

English
33
21
508
66.9K
Nanyun (Violet) Peng @ ACL26 retweetledi
Nanyun (Violet) Peng @ ACL26 retweetledi
Lianhui Qin
Lianhui Qin@Lianhuiq·
Congrats on the ICML 2026 Outstanding Paper Award! Glad to see diffusion LLMs attracting more and more attention. The Flexibility Trap shows that discrete diffusion can hurt reasoning under arbitrary token order: too much flexibility at the token level may skip the high-entropy forks where real reasoning decisions happen. This is exactly we think where latent diffusion becomes exciting. Instead of denoising arbitrary surface tokens, latent diffusion avoids the flexibility trap by utilizing continuous diffusion in latent space. This is the core idea behind LaDiR: reasoning at the semantic level, with global refinement in latent space. In LaDiR-RL (latent diffusion RL), diffusion sampling keeps multiple latent reasoning paths alive, helping mitigate the diversity collapse that often appears when RL directly optimizes text outputs. In our ongoing LaDiR-Multi, the same principle extends to multimodal reasoning: vision, text, code, audio, and action can enter a shared latent space, be reasoned over during denoising, and finally be discretized into text by an autoregressive model. We see very strong early results on vision-text benchmarks.
Lianhui Qin tweet media
ICML Conference@icmlconf

🏆Outstanding Paper is awarded to the following research papers: The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models High-Accuracy Sampling for Diffusion Models and Log-Concave Distributions

English
2
17
57
6.9K
AiDevCraft
AiDevCraft@AiDevCraft·
The 'zero web search' qualifier on Erdős 527 is the underrated detail — most AI math demos quietly lean on contaminated retrieval, so committing to pure reasoning with Lean as the only escape hatch is what makes the result trustworthy. 4000 lines through a Lean checker also raises the bar of 'partial' formalization to something nobody can hand-wave through.
English
1
0
0
75
Nanyun (Violet) Peng @ ACL26
Nanyun (Violet) Peng @ ACL26@VioletNPeng·
My first paper at Google is out! Thank you @rohanpaul_ai for highlighting LEAP. To share more thoughts on this direction: I strongly believe that as models generate longer and more complex proofs, automatic formal verification will be the key to the future of AI for math, and I'm bullish on using general LLMs + agentic framework for this task. As we started with competition math in LEAP for rigorous benchmarking purposes, we've already started to venture into research math. - Solved Erdős problem 527 (zero web search). - Partially formalized Knuth's cycle problem even case which resulted in ~4000 lines of Lean code. Please check out all of our solutions here: github.com/google-deepmin… I'm incredibly proud of this work, and we are just getting started. More to come!
Rohan Paul@rohanpaul_ai

Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Raised general LLM performance from under 10% to 70%. A general LLM failed badly when asked to write full formal proofs in 1 try, but became much stronger when it planned, split the work into smaller claims, reused past claims, and learned from Lean’s feedback. The paper shows the weakness was not just the model’s math ability, but the way it was being used - the absence of structured interaction with a verifier. The key idea is that the model does not try to write one giant perfect proof at once, because that usually fails on long and tricky problems. Instead, LEAP stores the proof as a graph of goals and subgoals, so useful lemmas can be reused instead of rediscovered every time. The authors tested LEAP on Putnam 2025 and a new Lean benchmark built from 60 IMO-style problems, where ordinary one-shot proof writing did very poorly. LEAP solved all 12 Putnam 2025 problems and raised general LLM performance on the Lean IMO benchmark from under 10% to 70%. ---- Link – arxiv. org/abs/2606.03303 Title: "LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks"

English
7
27
213
72.3K
Think_Different_
Think_Different_@ThinkDi92468945·
@VioletNPeng @DmitryRybin1 any guess why recent LLM + Lean systems seem to perform especially well on some combinatorics-style problems? Has anyone done a systematic domain-by-domain benchmark showing which areas of mathematics these systems handle well, and which areas remain weak?
English
1
0
0
75
Nanyun (Violet) Peng @ ACL26 retweetledi
Chun-Liang Li
Chun-Liang Li@chunliang_tw·
One reason of coming back to Google is the opportunity to work on AI4Science — and I still believe Google is one of the best places to work on it. I really love this work and am proud of what we accomplished. We’re just getting the momentum starts, -- more to come!
Rohan Paul@rohanpaul_ai

Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Raised general LLM performance from under 10% to 70%. A general LLM failed badly when asked to write full formal proofs in 1 try, but became much stronger when it planned, split the work into smaller claims, reused past claims, and learned from Lean’s feedback. The paper shows the weakness was not just the model’s math ability, but the way it was being used - the absence of structured interaction with a verifier. The key idea is that the model does not try to write one giant perfect proof at once, because that usually fails on long and tricky problems. Instead, LEAP stores the proof as a graph of goals and subgoals, so useful lemmas can be reused instead of rediscovered every time. The authors tested LEAP on Putnam 2025 and a new Lean benchmark built from 60 IMO-style problems, where ordinary one-shot proof writing did very poorly. LEAP solved all 12 Putnam 2025 problems and raised general LLM performance on the Lean IMO benchmark from under 10% to 70%. ---- Link – arxiv. org/abs/2606.03303 Title: "LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks"

English
0
2
23
3.1K
Nanyun (Violet) Peng @ ACL26 retweetledi
Po-Nien Kung
Po-Nien Kung@P_N_Kung·
Huge thanks to @rohanpaul_ai for sharing our work! We’re excited to share LEAP: a new agentic framework for formal theorem proving! LEAP mimics a human mathematician's workflow by constructing proof graph via informal blueprints, then searching and backtracking using compiler feedback. The results? 
🏆 Solves 12/12 Putnam 2025 problems
📈 Raises Gemini-3.1-pro's performance from 10% to 70% 
Read the paper: arxiv.org/abs/2606.03303
Rohan Paul@rohanpaul_ai

Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Raised general LLM performance from under 10% to 70%. A general LLM failed badly when asked to write full formal proofs in 1 try, but became much stronger when it planned, split the work into smaller claims, reused past claims, and learned from Lean’s feedback. The paper shows the weakness was not just the model’s math ability, but the way it was being used - the absence of structured interaction with a verifier. The key idea is that the model does not try to write one giant perfect proof at once, because that usually fails on long and tricky problems. Instead, LEAP stores the proof as a graph of goals and subgoals, so useful lemmas can be reused instead of rediscovered every time. The authors tested LEAP on Putnam 2025 and a new Lean benchmark built from 60 IMO-style problems, where ordinary one-shot proof writing did very poorly. LEAP solved all 12 Putnam 2025 problems and raised general LLM performance on the Lean IMO benchmark from under 10% to 70%. ---- Link – arxiv. org/abs/2606.03303 Title: "LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks"

English
1
10
15
3.6K
Zhiting Hu
Zhiting Hu@ZhitingHu·
🏆Honored to receive the Test of Time Award Honorable Mention #AISTATS2026 for our 2016 work Deep Kernel Learning, with the amazing @andrewgwils @rsalakhu @ericxing What a decade of AI progress! While GenAI is now driving massive real-world applications, the deepest underlying challenge remains: learning efficient representations of the world—for understanding, generation, predicting future worlds, and reasoning in the latent space. So much fun to think about for the next decade!⏳
Zhiting Hu tweet media
English
10
14
120
14.6K
Tanmay Parekh
Tanmay Parekh@tparekh97·
Life Update: Defended my PhD thesis titled “Towards Universal Event Extractions”, where I explored agentic workflows and synthetic data generation for IE. 👨‍🎓🎉 Sincere thanks to my advisors @kaiwei_chang and @VioletNPeng, and my committee @WeiWang1973 and @adityagrover_.
Tanmay Parekh tweet media
English
10
1
75
4.4K
Nanyun (Violet) Peng @ ACL26 retweetledi
Jia-Chen Gu
Jia-Chen Gu@Jiachen_Gu·
SPHERE has been accepted to ICLR 2026 and will be presented at Poster #214, Poster Session 4, Pavilion 3, 24 April, 3:15–5:45 p.m. I’m unable to attend in person, our amazing advisor @VioletNPeng will help present this work! Stop by if you’re around.
Jia-Chen Gu tweet media
Jia-Chen Gu@Jiachen_Gu

🚨Model editing in practice often collapses with catastrophic forgetting! Meet SPHERE🌐: an energy-regularized method that keeps weights uniformly distributed on hyperspheres, making sequential editing stable. Paper: arxiv.org/abs/2510.01172 Code: github.com/PlusLabNLP/SPH…

English
1
5
20
2.5K
Nanyun (Violet) Peng @ ACL26 retweetledi
Mohsen Fayyaz
Mohsen Fayyaz@mohsen_fayyaz·
📌 SteerMoE has been accepted to ICLR 2026 and will be presented at Poster #603, Poster Session 2, Pavilion 3, Thu 23 Apr, 3:15–5:45 p.m. 🙏 I’m unable to attend in person, so many thanks to @LucasBandarkar for presenting on our behalf. Stop by if you’re around.
Mohsen Fayyaz tweet media
Mohsen Fayyaz@mohsen_fayyaz

🚨 You can bypass ALL safety guardrails of GPT-OSS-120B 🚨❗🤯 How? By detecting behavior-associated experts and switching them on/off. 📄 Steering MoE LLMs via Expert (De)Activation 🔗 arxiv.org/abs/2509.09660 🧵👇

English
0
3
26
2.2K
Nanyun (Violet) Peng @ ACL26 retweetledi
Kai-Wei Chang
Kai-Wei Chang@kaiwei_chang·
Excited to work with a world-class team (@sahaiamit, Raghu Meka,@VioletNPeng,@terrence_tao,@WeiWang1973) on AI for math discovery and thank @LaudeInstitute for the support. I'm hiring a postdoc in this direction. Please contact me by May 1, 2026, if you’re interested.
Laude Institute@LaudeInstitute

Accelerating Science track: A century of scientific progress in one decade - that's the target. 👑Accelerating the Queen of Sciences, @UCLA Teaching AI to wonder, conjecture, and discover the way a mathematician does. Every hard science benefits if this works. @sahaiamit @kaiwei_chang Raghu Meka @VioletNPeng @terrence_tao @WeiWang1973 🌦️Actionable AI Weather Forecasts for Developing Economies, @UChicago Open-source AI weather forecasting infrastructure for developing economies, with millions of farmers already reached with better monsoon forecasts. @WillettBecca @ianfoster @turbulentjet Michael Kremer

English
2
7
23
3.4K
Nanyun (Violet) Peng @ ACL26 retweetledi
Lucas Bandarkar
Lucas Bandarkar@LucasBandarkar·
I'll be attending ICLR in a few weeks to present Multilingual Routing in Mixture-of-Experts Lmk if you're attending and want to talk MoEs, interpretability, multilinguality, etc
Lucas Bandarkar tweet media
Lucas Bandarkar@LucasBandarkar

Multilingual Routing in Mixture-of-Experts LLMs We present (1) an in-depth analysis of how MoE LLMs route multilingual texts, with very clear patterns + (2) a router intervention (steering) method that leads to consistent multilingual improvements! 🧵1/4 arxiv.org/pdf/2510.04694

English
0
1
20
1.9K
Nanyun (Violet) Peng @ ACL26 retweetledi
Prof. Anima Anandkumar
Prof. Anima Anandkumar@AnimaAnandkumar·
We’re excited to release TorchLean which is the first fully verified neural network framework in Lean. The Lean community has largely focused on pure mathematics. TorchLean expands this frontier toward verified neural network software and scientific computing. With the recent release of CSlib, we see this as another step toward a fully verified ML stack. We support features: 1. Executable IEEE-754 floating-point semantics (and extensible alternative FP models) verified tensor abstractions with precise shape/indexing semantics 2. Formally verified autograd system for differentiation of NN programs Proof-checked certification / verification algorithms like CROWN (robustness, bounds, etc.) 3. PyTorch-inspired modeling API with eager-style development + export/lowering to a shared IR for execution and verification Project page: leandojo.org/torchlean.html Paper: [2602.22631] TorchLean: Formalizing Neural Networks in Lean Work done @Robertljg, Jennifer Cruden, Xiangru Zhong, @huan_zhang12 and @AnimaAnandkumar. #MachineLearning #ScientificComputing #Lean
Prof. Anima Anandkumar tweet media
English
26
244
1.6K
149.3K
Nanyun (Violet) Peng @ ACL26
Nanyun (Violet) Peng @ ACL26@VioletNPeng·
Grateful for the generous support and looking forward to strengthening our collaborations with Salesforce researchers! @jasonwu0731 @steeve__huang
Salesforce AI Research@SFResearch

(2/4) Nanyun (Violet) Peng @VioletNPeng @UCLA is developing MAP-SE, a Multi-Agent Persuasion Simulation Engine that studies how influence emerges across adaptive agents with long-term memory—moving beyond simple one-on-one interactions. #AIResearch #AgenticAI

English
1
0
27
3.7K
Nanyun (Violet) Peng @ ACL26 retweetledi
Alex Spangher
Alex Spangher@AlexanderSpangh·
Neurips 2025 was such a blast! We snuck a grand piano into the CreativeAI Track to demo Aria, our pretrained chat-style music model:
English
23
45
398
30.4K
Nanyun (Violet) Peng @ ACL26 retweetledi
Mohsen Fayyaz
Mohsen Fayyaz@mohsen_fayyaz·
📸 Fantastic experience with @uclanlp at @NeurIPSConf 2025. An inspiring week of ideas and new connections.
Mohsen Fayyaz tweet mediaMohsen Fayyaz tweet media
English
2
3
30
11.6K