Li Erran Li @ ICML 🇰🇷
87 posts

Li Erran Li @ ICML 🇰🇷
@erranlli
Dr. Li Erran Li is the head of Science for human-in-the-loop services at AWS AI, Amazon and an adjunct professor at Columbia University.

It was an honor to give one of the invited workshop talks at ICML. The thesis: whoever turns messy real-world outcomes into reliable, scalable reward signals trains models where foundation models alone are stuck. It's also how we pick startups. Talk built with @HenryYin_ Coding agents suddenly started working because code shipped with free verifiers, like compilers and type checkers. We did reinforcement learning on those verifiers, and as we scaled compute, capability compounded. The next unicorn startup are ones that build verifiers for more challenging domains. Periodic Labs grades models with a physics fit against real measurements from its autonomous lab, and the capability ceiling moves. Applied Compute audits LLM judges across tens of thousands of rubric criteria before any RL runs, then trains open models to state of the art on Harvey's legal benchmark. Elorian is building verifiers that give models native visual reasoning capabilities, teaching models to actually see structure, where today's frontier models fail at counting more than ten objects.

guys I just cancelled my Claude plan I don’t know what happened









The Rising Star Award has been announced! Congratulations to Yining Hong @yining_hong , Rising Star Awardee for Spatial Intelligence, and to finalists Zhiyang Dou, Jiafei Duan @DJiafei , Hezhen Hu, Tiange Xiang @xxtiange , and Junyi Zhang @junyi42 . The awards are supported by 2077AI, with up to USD 30,000 in research gift funding to the awardee's institution and USD 2,000 in API credits for each finalist, helping early-career researchers further develop promising ideas in spatial intelligence. Full Rising Star list: e2e3d.github.io/rising_star.ht… Join the E2E3D Workshop today: 13:00–18:00 · Room 501 E2E3D Workshop: e2e3d.github.io/index.html


Arena reached a $100M annual revenue run rate just 8 months after launching our evaluation product. We started as a research project at UC Berkeley with a simple mission: measure AI progress through real-world use. As AI shifts from chatbots to agents taking on longer, higher-stakes work, the problem matters more than ever. Today, Arena measures real-world AI utility with a community of tens of millions. With Agent Arena, we’re evaluating long-running agents on complex, real-world tasks - how they use tools, adapt to feedback, recover from errors, and accomplish goals set by humans. We are excited to keep deepening our work in agentic evaluations. Here’s @ml_angelopoulos on what this milestone means and where we go from here:





I’m excited to share that I’ll be joining OpenAI and look forward to working with the exceptional team there. It was a difficult decision to move on. I’m incredibly proud of the amazing team at Google and everything we’ve built together. It has been an honor and a pleasure to work with all of you.

personal news: i've joined Elorian as Chief Reasoning Architect. multimodal AGI is the most critical frontier as we move from the era of chatbots to coding agents to models that reason and act over the physical world. i'm really excited to design natively visual models across thinking, agents, architectures, and the systems stack with the amazing team at Elorian. i wish the best to everyone at xAI & SpaceX — driving posttraining was a unique experience with so many memorable stories. all the best to the team, and to Elon.

Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research. Developed with 173 scientists from biotechnology and pharmaceutical research, LifeSciBench includes 750 expert-authored tasks across seven biological research workflows. openai.com/index/introduc…






What if we could universally recombine, insert, delete, or invert any two pieces of DNA? In back-to-back @Nature papers, we report the discovery of bridge RNAs and 3 atomic structures of the first natural RNA-guided recombinase - a new mechanism for programmable genome design








We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger intelligence for drug discovery, analysis, design, and experimental workflows. openai.com/index/introduc…








