
Shom
696 posts

Shom
@ShomLinEd
language model | sequence modeling | education | HCI



Why can Kimi ship K3? Let me tell my story. Earlier this year, I left academia for industry. I talked to a lot of companies along the way. Here's what I saw: 1⃣Arrogance. They believe the AI war is over, and they won. No hunger for the future, and no hunger for talent. 2⃣Restlessness. Young labs short on foundation, either rushing to catch the frontier or pivoting away from the competition. 3⃣Fear. Strong teams with real experience, but from the second tier, they can't quite bring themselves to aim for #1. 4⃣Misalignment. Everyone is optimizing for their own credit, but nobody really cares whether the company can reach AGI. Kimi was different. Over many conversations with the founders, the same thing came through every time: a raw, genuine hunger for AGI. I joined. The hunger was real. We shipped K3. This is only the beginning.



Introducing EdgeBench, a benchmark designed to study how agents learn from environments over at least 12~72-hour runs. We find that performance follows a log-sigmoid function of environment interaction time with high precision. EdgeBench is built with three ingredients: - 🌍 Real & Diverse: 134 real-world tasks across 6 task categories, spanning scientific problems, professional knowledge work, software engineering, optimization, formal math, and games. - ⏳ Ultra-Long-Horizon: Each task supports 12–72 hours of agent work. Recorded human effort averages 57.2 hours. - 🔁 Informative Feedback: Agents receive real-world feedback for continuous improvement. After 38,000 hours of agent runs on EdgeBench, a scaling law for learning from environments emerges: - 📈 As agents interact with task environments over time, their aggregate performance is precisely fit by a log-sigmoid function. - 🧠 This phenomenon can be explained by an elegant theory of graph exploration. We are releasing an initial 51 of the 134 tasks, together with the full evaluation framework, to help advance long-horizon agent research. Check our blog & paper for more findings! Blog edge-bench.org Paper edge-bench.org/paper.pdf GitHub github.com/ByteDance-Seed… Dataset huggingface.co/datasets/ByteD… Details below 👇🧵


最近在滑雪机上学单板,教练会用语言教各种情况下该怎么做动作,但是他说的“踩”、“拧”、“释放压力”和我理解的显然不是一个东西,总之我就是很难主动控制雪板而且非常吃力。 后来我干脆放弃了,让雪板随机游走,我只控制自己别摔倒,同时尽可能记下雪板的运动状态和当时自己的身体状态,记到一定程度之后,如果想要主动控制,就从记忆里找有没有match的场景,如果有的话就尽力replay一下,如果恰好控制住了,这就成了一个正样本,如果没有,就放弃控制让它继续随机运动。我发现这种学习方式非常省力且效率很高。 仔细一想,这个套路很像:先随机采样学习先验分布,再强化学习提高目标分布的概率,果然还是 pretrain + RL 效率高。 另一方面,教练有 knowledge curse,他确实很难再体会到“完全不会滑雪”是什么感觉了,所以给了好多我学不会的SFT数据🥲









Alibaba Qwen3.7 slowly fading into irrelevance at the frontier due to proprietary stance. In it's place we have Minimax M3 and... *checks notes* Rio 3.5 397b, made by the municipal IT company of Rio de Janeiro's city government. huggingface.co/prefeitura-rio…









$200 FREE CREDIT! We just launched our inference platform for beta testing, and we're giving it to the community first. ⭐ Star SGLang on GitHub (github.com/sgl-project/sg…) + repost this to claim your credits. → Limited spots, first come first serve → Deadline: May 13, 2025 (AoE) Every star, every issue filed, every PR reviewed, every question answered in Slack — You built this with us. Thank you for believing in open-source AI infrastructure, in our mission, and in us. Claim your credits: platform.radixark.com















