Haonan Li

50 posts

Haonan Li

Haonan Li

@haonanlp

Senior Research Scientist at MBZUAI-IFM

Abu Dhabi Katılım Mart 2018
118 Takip Edilen221 Takipçiler
Haonan Li retweetledi
MBZUAI
MBZUAI@mbzuai·
Today, we are releasing a new version of K2 (K2-V2), a 360-open LLM built from scratch as a superior base for reasoning adaptation, while still excelling at core LLM capabilities like conversation, knowledge retrieval, and long-context understanding. K2 fills a major gap: highly capable models with no transparency. Instead of releasing only weights, we’re sharing the full training story — dataset recipes, mid-training checkpoints, logs, code, and evaluation tools. That’s 360-open. What’s inside: • 70B dense transformer engineered as a reasoning-enhanced base model • Native 512K context (extendable via RoPE scaling) • Mid-training reasoning phase • Strong tool-use scaffolding What we’re open-sourcing: • 250M+ reasoning traces (math, planning, multi-step logic) • Full pre- & mid-training data compositions • All mid-training checkpoints • Training logs, code, Eval360 Performance: • GPQA-Diamond: 55.1% mid-training → 69.3% after SFT (strongest fully open 70B model) • KK-8 Logic Puzzles: 83% — competitive with DeepSeek-R1 & OpenAI o3-mini-high • ArenaHard V2: 62.1% — close to Qwen3 235B • Outperforms Qwen2.5-72B and approaches Qwen3-235B despite being smaller and fully transparent. 🔗 The Model: bit.ly/3KIYwuo 🔗Technical Report: bit.ly/49V8h2U 🔗Blog: bit.ly/49V7gb6
MBZUAI tweet media
English
3
35
110
55.4K
Haonan Li
Haonan Li@haonanlp·
Our website just got an upgrade! No sign-up needed, easier to use. Come feel the speed. k2think.ai
𝚐𝔪𝟾𝚡𝚡𝟾@gm8xx8

K2-Think 32B, built on Qwen2.5 Scores (pass@1, avg of 16 runs): - AIME’24: 90.8 - AIME’25: 81.2 - HMMT’25: 73.8 - Omni-HARD: 60.7 - LiveCodeBench v5: 63.97 - GPQA-Diamond: 71.1 It is trained with long CoT SFT and RL with verifiable rewards on the Guru dataset, then improved at inference through a Plan-Before-You-Think scaffold and Best-of-3 sampling, which also shortens outputs by 6–12%. Deployment on the Cerebras Wafer-Scale Engine achieves ~2,000 tokens/sec (32k ≈ 16s) versus ~200 tokens/sec (32k ≈ 160s) on H100/H200. Safety-4 averages 0.75, strong in refusal and conversational robustness, weaker on cybersecurity and prompt-extraction. The model, training code, inference code, and full tech report are openly available, with the complete reasoning system also served as a live API and web portal at Cerebras-level speed. 𝖯𝖺𝗋𝗍 𝗈𝖿 𝗍𝗁𝖾 𝗌𝖺𝗆𝖾 𝖪𝟤/𝖫𝖫𝖬𝟥𝟨𝟢 𝖾𝖼𝗈𝗌𝗒𝗌𝗍𝖾𝗆 𝗍𝗁𝖺𝗍 𝖺𝗅𝗌𝗈 𝗍𝗋𝖺𝗂𝗇𝖾𝖽 𝗍𝗁𝖾 𝟨𝟧𝖡 𝗈𝗉𝖾𝗇 𝖪𝟤 𝖣𝖨𝖠𝖬𝖮𝖭𝖣 𝗆𝗈𝖽𝖾𝗅

English
0
1
4
1.4K
Haonan Li
Haonan Li@haonanlp·
A big milestone for advanced AI reasoning. K2 Think is more than just a model — it’s a full system with powerful capabilities and impressively fast inference. Try it out yourself! k2think.ai
MBZUAI@mbzuai

Introducing K2 Think - a breakthrough in advanced AI reasoning. Developed by MBZUAI’s Institute of Foundation Models and @G42ai, K2 Think delivers frontier reasoning performance at a fraction of the size of today’s largest systems. Smaller. Smarter. Open to the world. Available now: K2Think.Ai/K2Think #K2Think #AI #OpenSource #MBZUAI #G42 #Innovation

English
0
2
12
828
Haonan Li retweetledi
Tahnoon Bin Zayed Al Nahyan
The upcoming launch of the "K2 Think" model will mark a significant step in advancing artificial intelligence from the UAE to the world, reflecting our leadership's vision for technological progress. K2 Think will combine the efficiency of smaller models with world-class performance and superior inference speeds, surpassing larger systems to set a new global benchmark for open-source reasoning Developed through the partnership of @MBZUAl and @G42ai, it showcases the strength of cooperation between academia, public institutions, and the private sector in turning national aspirations into reality. We congratulate the UAE and its leadership on this achievement, as our nation continues to enhance its global standing across all fields.
Tahnoon Bin Zayed Al Nahyan tweet media
English
42
168
588
102.9K
Haonan Li retweetledi
LLM360
LLM360@llm360·
The MBZUAI IFM and the LLM360 team's first day at @iclr_conf, come to visit our new Institute of Foundation Models! Booth D04 in Hall 2! We’re looking forward to meeting researchers and engineers to introduce them to @mbzuai .
LLM360 tweet media
English
0
9
28
4.8K
Haonan Li
Haonan Li@haonanlp·
Does cleaner data always yield better models in LLM alignment? Our findings challenge this notion, revealing that even high-quality, difficult examples, can hinder alignment, reducing performance by approximately 10 win rate points. More insights and findings: Paper: arxiv.org/abs/2502.09650 Github: github.com/glorgao/Select…
Haonan Li tweet mediaHaonan Li tweet mediaHaonan Li tweet mediaHaonan Li tweet media
English
0
0
3
280
Haonan Li retweetledi
elvis
elvis@omarsar0·
o3-mini-high (left) vs. deepseek-r1 (right) results from the first try deepseek-r1 is cracked... wtf!
English
102
169
2.4K
719.8K
Haonan Li
Haonan Li@haonanlp·
ICLR25 + NAACL25 acceptances in one night—what a surprise! NAACL: NAT: Enhancing Agent Tuning with Negative Samples [arxiv.org/abs/2402.11651] ICLR: ToolGen: Unifying Tool Retrieval and Calling via Generation [arxiv.org/abs/2410.03439] Both works are agent-focused and lead by @realReasonWang, showcases his exceptional talent and dedication. Huge congrats!
Haonan Li tweet mediaHaonan Li tweet media
English
2
4
40
3.8K
Haonan Li
Haonan Li@haonanlp·
We are excited to share our tutorial at @coling2025 on Safety Issues for Generative AI. We’ll explore: 🔍 LLM Jailbreaking and red-teaming 🤖 Multi-modal & agentic AI Safety 🛡️ AI automating and Scheming ... 📅 20 Jan, 9:00-12:30 📍 Capital Suite 7, ADNEC, Abu Dhabi 🌐 Website: librairesearch.github.io/tutorial/index… This Tutorial is organized in collaboration with @han_xudong @Emad_A_Alghamdi @ShomLinEd @monojitchou @zjf_heart @eltimster Looking forward to seeing you there! #COLING2025 #AISafety #GenerativeAI #AI
English
0
0
1
615
Haonan Li
Haonan Li@haonanlp·
5/ 🙏Big thanks to @mbzuai and all the researchers and institutions for their invaluable support of Libra-Leaderboard. Your collaboration and contributions are key to making this project a reality. The names of our supporters are featured in the image below. We look forward to continuing this journey together in advancing AI safety.🤖💡
Haonan Li tweet media
English
0
0
1
121
Haonan Li
Haonan Li@haonanlp·
4/ Safety Arena is an innovative platform designed to make AI safety accessible to non-experts, helping raise public awareness of AI risks. Users can interact with LLMs through a chat-based interface and apply built-in adversarial attack methods to their input prompts. The platform is closely integrated with Libra-Leaderboard, where user feedback directly impacts the evaluation scores of LLMs. Try it out: leaderboard.librai.tech/Arena
Haonan Li tweet mediaHaonan Li tweet mediaHaonan Li tweet media
English
1
0
1
173
Haonan Li
Haonan Li@haonanlp·
We are excited to announce Libra-Leaderboard: The first LLM leaderboard dedicated to balancing safety and capability in LLMs. As AI advances, ensuring its safety becomes more critical than ever. By prioritizing safety measurement, we aim to inspire the AI community to make safety a core factor in model evaluation and development, alongside capability.🚀 #LLM #AI #AISafety #ResponsibleAI Website: leaderboard.librai.tech/LeaderBoard Paper: arxiv.org/abs/2412.18551 Thanks to all the contributors: @han_xudong @_ZenanZhai @honglin_mu @_HaoWang @_ZhenxuanZhang @yilin_link @ShomLinEd @realReasonWang @ArtemShelmanov @xiangyuqi_pton @_YuxiaWang @_DonghaiHong @youliang_yuan @_MengChen @HaoqinT @FajriKoto @ttk_kuribayashi @HaoqinT @rishabh15 @BingchenZhao @yawen_duan @_YiLiu @_YaodongYang @dongyp13 @soujanyaporia @stefan_fee @waterluffy @RenHector @Emad_A_Alghamdi @IGurevych @preslav_nakov @monojitchou @eltimster
Haonan Li tweet media
English
1
15
31
3.9K
Haonan Li retweetledi
elvis
elvis@omarsar0·
ToolGen: Unified Tool Retrieval and Calling via Generation Integrate tool knowledge directly into LLMs by representing tools as a unique token. This allows the LLM to generate tool calls and arguments, enabling seamless tool invocation and language generation. "Experimental results with over 47,000 tools show that ToolGen not only achieves superior results in both tool retrieval and autonomous task completion but also sets the stage for a new era of AI agents that can adapt to tools across diverse domains." I like this approach because it prevents the need to call an external system and can probably be better tuned for more challenging tool-calling scenarios. It might also be an interesting method to combine with CoT and these large reasoning models.
elvis tweet media
English
8
56
294
51.3K
Haonan Li
Haonan Li@haonanlp·
Our latest work on Tool Learning and Agent 🚀 ToolGen: Unified Tool Retrieval and Calling via Generation Integrating tool knowledge directly into LLMs by representing tools as unique tokens, enabling seamless tool invocation and language generation. 🔧🤖 💡 ToolGen achieves superior performance in both tool retrieval and autonomous task completion. We believe it will set the stage for a new era of AI agents ! #AI #LLM #ToolRetrieval #AIResearch #NLP #ML
Haonan Li tweet media
English
0
1
6
427
Haonan Li retweetledi
LibrAI
LibrAI@libr_ai·
📝Please fill in your information to get a free pass before they’re gone-only 3 days left to register!  ⬇️Check the comments for the link to our questionnaire. Let’s meet and talk about innovation, AI, and opportunities! #LibrAI #AI #GITEX #FreePass #GITEX2024 #ExpandNorthStar
LibrAI@libr_ai

🎟️We’ve got FREE passes to #GITEX2024 & #ExpandNorthStar! 🚀LibrAI is excited to share exclusive access to the events. Visit us to discover how we automate testing and evaluation for AI systems. Let us help you to deploy your AI systems faster, cheaper, and more reliably!

English
1
4
5
659
Haonan Li
Haonan Li@haonanlp·
Excited to share our 6 papers at #ACL2024NLP We focus on benchmarking and evaluation. Welcome to catch up with us. 12 Aug @ 17:15 Poster Session 2 1️⃣ CMMLU: Measuring massive multitask language understanding in Chinese 13 Aug @ 12:15 Poster Session 3 2️⃣ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic 3️⃣Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification 4️⃣A Chinese Dataset for Evaluating the Safeguards in Large Language Models 13 Aug @ 16:00 Poster Session 5 5️⃣Demystifying Instruction Mixing for Fine-tuning Large Language Models 6️⃣EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Haonan Li tweet mediaHaonan Li tweet media
English
2
9
57
9K