uclanlp

356 posts

uclanlp

uclanlp

@uclanlp

UCLA Natural Language Processing Research #NLProc

Katılım Mart 2021
130 Takip Edilen1.3K Takipçiler
uclanlp retweetledi
uclanlp retweetledi
Nanyun (Violet) Peng @ ACL26
Nanyun (Violet) Peng @ ACL26@VioletNPeng·
My first paper at Google is out! Thank you @rohanpaul_ai for highlighting LEAP. To share more thoughts on this direction: I strongly believe that as models generate longer and more complex proofs, automatic formal verification will be the key to the future of AI for math, and I'm bullish on using general LLMs + agentic framework for this task. As we started with competition math in LEAP for rigorous benchmarking purposes, we've already started to venture into research math. - Solved Erdős problem 527 (zero web search). - Partially formalized Knuth's cycle problem even case which resulted in ~4000 lines of Lean code. Please check out all of our solutions here: github.com/google-deepmin… I'm incredibly proud of this work, and we are just getting started. More to come!
Rohan Paul@rohanpaul_ai

Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Raised general LLM performance from under 10% to 70%. A general LLM failed badly when asked to write full formal proofs in 1 try, but became much stronger when it planned, split the work into smaller claims, reused past claims, and learned from Lean’s feedback. The paper shows the weakness was not just the model’s math ability, but the way it was being used - the absence of structured interaction with a verifier. The key idea is that the model does not try to write one giant perfect proof at once, because that usually fails on long and tricky problems. Instead, LEAP stores the proof as a graph of goals and subgoals, so useful lemmas can be reused instead of rediscovered every time. The authors tested LEAP on Putnam 2025 and a new Lean benchmark built from 60 IMO-style problems, where ordinary one-shot proof writing did very poorly. LEAP solved all 12 Putnam 2025 problems and raised general LLM performance on the Lean IMO benchmark from under 10% to 70%. ---- Link – arxiv. org/abs/2606.03303 Title: "LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks"

English
7
27
214
72.3K
uclanlp retweetledi
Po-Nien Kung
Po-Nien Kung@P_N_Kung·
Huge thanks to @rohanpaul_ai for sharing our work! We’re excited to share LEAP: a new agentic framework for formal theorem proving! LEAP mimics a human mathematician's workflow by constructing proof graph via informal blueprints, then searching and backtracking using compiler feedback. The results? 
🏆 Solves 12/12 Putnam 2025 problems
📈 Raises Gemini-3.1-pro's performance from 10% to 70% 
Read the paper: arxiv.org/abs/2606.03303
Rohan Paul@rohanpaul_ai

Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Raised general LLM performance from under 10% to 70%. A general LLM failed badly when asked to write full formal proofs in 1 try, but became much stronger when it planned, split the work into smaller claims, reused past claims, and learned from Lean’s feedback. The paper shows the weakness was not just the model’s math ability, but the way it was being used - the absence of structured interaction with a verifier. The key idea is that the model does not try to write one giant perfect proof at once, because that usually fails on long and tricky problems. Instead, LEAP stores the proof as a graph of goals and subgoals, so useful lemmas can be reused instead of rediscovered every time. The authors tested LEAP on Putnam 2025 and a new Lean benchmark built from 60 IMO-style problems, where ordinary one-shot proof writing did very poorly. LEAP solved all 12 Putnam 2025 problems and raised general LLM performance on the Lean IMO benchmark from under 10% to 70%. ---- Link – arxiv. org/abs/2606.03303 Title: "LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks"

English
1
10
15
3.6K
uclanlp retweetledi
Wenbo Hu
Wenbo Hu@gordonhu608·
Our 2nd 3D-LLM/VLA Workshop is happening Tomorrow (June 3) at Room 1CD starting 1pm! For more info and time slot of each speaker, please refer to our website: 3d-llm-vla.github.io #CVPR2026 @CVPR
Wenbo Hu tweet media
English
0
7
24
6.1K
uclanlp retweetledi
Po-Nien Kung
Po-Nien Kung@P_N_Kung·
🌟 Excited to share Ctrl-R, accepted as an ICML 2026 Spotlight! Specify a target reasoning structure, Ctrl-R can enforce it during rollout, and retain accurate importance-sampling weights for principled policy optimization. paper: arxiv.org/abs/2603.01641 🧵[1/5]
GIF
English
3
23
104
9.4K
uclanlp retweetledi
Yining Hong
Yining Hong@yining_hong·
Meet Embodied Web Agents that bridge physical-digital realms. Imagine embodied agents that can search for online recipes, shop for ingredients and cook for you. Embodied web agents search internet information for implementing real-world embodied tasks. All data, codes and web environments are available at embodied-web-agent.github.io Paper link: arxiv.org/abs/2506.15677
English
4
37
168
54.3K
uclanlp retweetledi
Wenbo Hu
Wenbo Hu@gordonhu608·
Does GRPO handle multiple-task advantage normalization effectively? 🤔 🚀We introduce Gaussian GRPO (G²RPO), a novel method that mathematically forcing the advantage distribution of any given task to strictly converge to a standard normal distribution. G²RPO provides: ✅1) intrinsic robustness to outliers, ✅2) symmetric updates for positive and negative rewards, ✅3) uniform variance across diverse tasks. We adopt G²RPO and task-level response length and entropy shaping which leads to OpenVLThinkerV2! 🏆A new model with SOTA open-source visual reasoning and perception performance among the similar size models. (1/n)👇 #VLM #LLM #Multimodal #RL #GRPO
Wenbo Hu tweet media
English
1
15
76
6.8K
uclanlp retweetledi
Kuan-Hao Huang
Kuan-Hao Huang@kuanhaoh_·
The first-ever Texas NLP Symposium wrapped up yesterday! 🎉🎉🎉 Huge thanks to all the speakers and attendees for making it a huge success. I hope everyone had a great time. Stay tuned for info on next year! #TexasNLP Check photos and highlights here: photos.app.goo.gl/AhuqKfKDXHyUQw…
Kuan-Hao Huang tweet mediaKuan-Hao Huang tweet mediaKuan-Hao Huang tweet mediaKuan-Hao Huang tweet media
English
0
11
47
5.5K
uclanlp retweetledi
Laude Institute
Laude Institute@LaudeInstitute·
Accelerating Science track: A century of scientific progress in one decade - that's the target. 👑Accelerating the Queen of Sciences, @UCLA Teaching AI to wonder, conjecture, and discover the way a mathematician does. Every hard science benefits if this works. @sahaiamit @kaiwei_chang Raghu Meka @VioletNPeng @terrence_tao @WeiWang1973 🌦️Actionable AI Weather Forecasts for Developing Economies, @UChicago Open-source AI weather forecasting infrastructure for developing economies, with millions of farmers already reached with better monsoon forecasts. @WillettBecca @ianfoster @turbulentjet Michael Kremer
Laude Institute tweet media
English
1
6
23
8.9K
uclanlp retweetledi
Kai-Wei Chang
Kai-Wei Chang@kaiwei_chang·
Excited to work with a world-class team (@sahaiamit, Raghu Meka,@VioletNPeng,@terrence_tao,@WeiWang1973) on AI for math discovery and thank @LaudeInstitute for the support. I'm hiring a postdoc in this direction. Please contact me by May 1, 2026, if you’re interested.
Laude Institute@LaudeInstitute

Accelerating Science track: A century of scientific progress in one decade - that's the target. 👑Accelerating the Queen of Sciences, @UCLA Teaching AI to wonder, conjecture, and discover the way a mathematician does. Every hard science benefits if this works. @sahaiamit @kaiwei_chang Raghu Meka @VioletNPeng @terrence_tao @WeiWang1973 🌦️Actionable AI Weather Forecasts for Developing Economies, @UChicago Open-source AI weather forecasting infrastructure for developing economies, with millions of farmers already reached with better monsoon forecasts. @WillettBecca @ianfoster @turbulentjet Michael Kremer

English
2
7
23
3.4K
uclanlp retweetledi
Kai-Wei Chang
Kai-Wei Chang@kaiwei_chang·
Are all suicide deaths documented equally in national surveillance data? Across 300,000+ National Violent Death Reporting System narrative summaries of suicide incidents, we find that Black non-Hispanic decedents have shorter, less complex summaries than white non-Hispanic decedents, pointing to an urgent need for data equity measures in public health reporting. Full findings in the paper — link below 📷 mdpi.com/2078-2489/16/1… With @christinachanc @arsenievK @drvickiemays & Susan D. Cochran
English
0
2
6
716
uclanlp
uclanlp@uclanlp·
For the UCLA NLP seminar talk this Friday, we are thrilled to host Prof. Christopher Potts @ChrisGPotts from Stanford @stanfordnlp ! Title: “The Archai of Palimpsestic Memorization” When: 2–3 PM (PST), Friday, Jan 23 Registration: ucla.zoom.us/meeting/regist…
uclanlp tweet media
English
0
6
23
10.6K
uclanlp retweetledi
Himanshu Kumar
Himanshu Kumar@codewithimanshu·
@kaiwei_chang Wow, 5 papers at NeurIPS, Kai-Wei! That's impressive, but is AI for Math really the future, or just hype?
English
0
1
0
553
uclanlp
uclanlp@uclanlp·
For this week’s NLP seminar talk, we are thrilled to host Prof. Arman Cohan @armancohan from Yale @Yale and Ai2 aitanaallenperez on LLM evaluation and alignment: Time: 2-3PM, Friday, Nov 14th (PST) Registration: ucla.zoom.us/meeting/regist…
uclanlp tweet media
English
0
3
15
6.8K
uclanlp retweetledi
Nanyun (Violet) Peng @ ACL26
Nanyun (Violet) Peng @ ACL26@VioletNPeng·
It’s a wrap! I hope y’all enjoyed #EMNLP25 as much as I did! Big shoutout to the photography team! All my photos in this post are taking from their websites. To all attendees: you should check out if you haven’t already!!
Nanyun (Violet) Peng @ ACL26 tweet mediaNanyun (Violet) Peng @ ACL26 tweet mediaNanyun (Violet) Peng @ ACL26 tweet mediaNanyun (Violet) Peng @ ACL26 tweet media
English
8
4
127
8.1K
uclanlp
uclanlp@uclanlp·
@VioletNPeng served as the Program Co-Chair for #EMNLP25, one of the largest NLP conferences, which received over 8,000 submissions and drew 6,000 participants.
uclanlp tweet mediauclanlp tweet media
English
0
1
22
12.2K
uclanlp retweetledi
Haoyi Qiu
Haoyi Qiu@HaoyiQiu·
🤖💬AI agents can be easily persuaded (like Anthropic’s Claudius often giving discounts). 🤔Previous study on persuasion has been exclusively on text-only modality. We wonder: are AI agents more susceptible when presented with multimodal content? Introducing MMPersuade, a comprehensive multimodal benchmark that assesses AI agents’ susceptibility to established persuasion principles, covering commercial, subjective and behavioral, and adversarial contexts.
Haoyi Qiu tweet media
English
11
25
130
29.5K
uclanlp
uclanlp@uclanlp·
For this week’s UCLA NLP seminar, we are thrilled to invite Prof. Rose Yu @yuqirose from UCSD @UCSanDiego: Talk Title: “Towards AI Co-Scientists: Agentic Foundation Models for Physical Universe” Time: 2-3PM, Friday, October 31st (PDT) Registration: ucla.zoom.us/meeting/regist…
uclanlp tweet media
English
1
0
10
4.3K