Harry He
70 posts

Harry He
@harryheisme
Founding member @radixark Trust the process




Miles is now featured on the PyTorch Foundation blog. As models grow, shift from dense to MoE, and span more specialized hardware, RL post-training is no longer just about the algorithm. It is a distributed systems problem. Miles is our open-source RL training framework, built for exactly that. It comprises four systems behind a small, pluggable trainer: SGLang (@sgl_project) for rollout, Megatron-LM (@NVIDIAAI) for training, Ray (@raydistributed) for orchestration, and PyTorch (@PyTorch) as the common layer for models and numerics. Out of the box, you also get MoE-aware rollout/training alignment, a unified BF16/FP8/MXFP8/NVFP4/INT4-QAT pipeline, fast NCCL/RDMA weight sync, fault tolerance, and ready-to-run recipes for frontier models like DeepSeek V4, GLM 5.2, Qwen3.6, Kimi K2.6, and Nemotron 3 Ultra. Our goal is simple: make frontier-scale LLM RL easier to reproduce, extend, and operate. Thank you, PyTorch Foundation, and everyone who got Miles here, especially the legendary @slime_framework team!

🔥 Summer done right. A whole roast lamb BBQ to celebrate the good times with friends, family, builders, and our community, World Cup on the big screen (good game of Spain vs. Uruguay ⚽), smoke rolling off the grill, cold drinks in the sun, and a backyard full of great people just soaking up the vibes, that's how we enjoy the summertime! 🍢☀️🍺 Great food, great people, great times. Huge thanks to everyone who came out, and shoutout to @FishAudio for co-hosting this BBQ with us. Here's to the rest of this summer together! 🫶





GLM 5.2 from @Zai_org just launched, and we're excited to be a Day 0 partner on @Cloudflare workers ai personally my favourite model - doesn't overthink and gets the job done. try it for yourself 🧡 and special thanks to @sgl_project always for day 0 inference engine support developers.cloudflare.com/changelog/post…





🎉 SGLang v0.5.13 is out! First, new model support! Nemotron 3 Ultra, Step-3.7-Flash, Command A+, plus new diffusion models: Cosmos3, FLUX.2-Klein, Ideogram 4, LingBot-World, SANA-WM, and Ernie-Image. Here are the highlights for this release: - Speculative Decoding V2 is now the default! Tree drafting (topk>1) for faster generation - Breakable CUDA Graphs now make prefill faster - Qwen 3.5 runs faster on NVIDIA Blackwell with new GDN kernels - HiCache with UnifiedTree on by default for hybrid SWA/Mamba models - SGLang-Diffusion now supports realtime generation! Plus progressive resolution - Multiple performance and feature updates for DeepSeek V4 Thanks to our amazing partners and model makers: @NVIDIAAI @AMD @intel @awscloud @boson_ai @cohere @bfl_ml @ideogram_ai @deepseek_ai @Kimi_Moonshot @Alibaba_Qwen @StepFun_ai @Baidu_Inc @robbyant_brain

AI startup CFO, CEO, CMO, and CTO

🏙️ SGLang NY Tech Week Happy Hour Recap Last Wednesday, SGLang hosted a NY Tech Week Happy Hour in NYC, co-hosted with @HOFCapital, @Cloudflare, @CrusoeAI, and @ArklexAI. 380+ registered, 200+ in the room, and one unforgettable night. 🧡 The room was packed with engineers, researchers, and enthusiasts from quant funds, banks, and trading firms, all there to talk about one thing: where inference is headed as LLMs move into latency-sensitive production across trading, research, compliance, and risk. NYC, you showed up and brought the energy. We loved every minute. Until next time! ☀️ #NYTechWeek @Techweek_

👋 @sgl_project is back! Welcome to the home of the SGLang community! While @lmsysorg keeps you posted on technical drops and partner news, this space is for you! Here's what we've got lined up: 🚀 Version releases: every new SGLang drop, unpacked 🎙️ Office Hours: deep dives, live deployments, and team Q&A 📺 Tutorials: short how-tos and the best Office Hour moments 🌟 Community spotlights: the cool stuff you're building with SGLang 📅 Event updates: meetups, workshops, and where to catch us next And we'd love to hear from you! What do you want to see? Benchmarks? Model deep dives? A topic for the next Office Hour? Drop it below 👇 Every idea gets read!
