

Haonan Li
50 posts

@haonanlp
Senior Research Scientist at MBZUAI-IFM



K2-Think 32B, built on Qwen2.5 Scores (pass@1, avg of 16 runs): - AIME’24: 90.8 - AIME’25: 81.2 - HMMT’25: 73.8 - Omni-HARD: 60.7 - LiveCodeBench v5: 63.97 - GPQA-Diamond: 71.1 It is trained with long CoT SFT and RL with verifiable rewards on the Guru dataset, then improved at inference through a Plan-Before-You-Think scaffold and Best-of-3 sampling, which also shortens outputs by 6–12%. Deployment on the Cerebras Wafer-Scale Engine achieves ~2,000 tokens/sec (32k ≈ 16s) versus ~200 tokens/sec (32k ≈ 160s) on H100/H200. Safety-4 averages 0.75, strong in refusal and conversational robustness, weaker on cybersecurity and prompt-extraction. The model, training code, inference code, and full tech report are openly available, with the complete reasoning system also served as a live API and web portal at Cerebras-level speed. 𝖯𝖺𝗋𝗍 𝗈𝖿 𝗍𝗁𝖾 𝗌𝖺𝗆𝖾 𝖪𝟤/𝖫𝖫𝖬𝟥𝟨𝟢 𝖾𝖼𝗈𝗌𝗒𝗌𝗍𝖾𝗆 𝗍𝗁𝖺𝗍 𝖺𝗅𝗌𝗈 𝗍𝗋𝖺𝗂𝗇𝖾𝖽 𝗍𝗁𝖾 𝟨𝟧𝖡 𝗈𝗉𝖾𝗇 𝖪𝟤 𝖣𝖨𝖠𝖬𝖮𝖭𝖣 𝗆𝗈𝖽𝖾𝗅

Introducing K2 Think - a breakthrough in advanced AI reasoning. Developed by MBZUAI’s Institute of Foundation Models and @G42ai, K2 Think delivers frontier reasoning performance at a fraction of the size of today’s largest systems. Smaller. Smarter. Open to the world. Available now: K2Think.Ai/K2Think #K2Think #AI #OpenSource #MBZUAI #G42 #Innovation


























🎟️We’ve got FREE passes to #GITEX2024 & #ExpandNorthStar! 🚀LibrAI is excited to share exclusive access to the events. Visit us to discover how we automate testing and evaluation for AI systems. Let us help you to deploy your AI systems faster, cheaper, and more reliably!



