Limpu Jiao
263 posts



For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-Weigh…



Today we’re expanding the Gemini family with three new models built to be faster, more token efficient, and reliable at scale. Meet the new Gemini models ↓

今天晚上各个群都在传一份《国内大模型蒸馏风波的来龙去脉》。 我大概看了一下,很可惜,这更像是一份由外行根据各种流言拼凑出来的 AI slop。 里面的内容有真有假,但几个核心判断基本都经不起推敲。 比如里面声称,智谱 @Zai_org 在今年 4 月份就已经破解并开始蒸馏 Fable。 这个时间线真的很离谱。 因为 4 月份的时候,Fable (那个时候还是Mythos)甚至还没有正式发布,能够使用的基本都是 Anthropic 的内部测试用户,通过Project Glasswing进行访问。 那智谱是怎么拿到的访问权限呢? 难道是 @AnthropicAI 的邀请或者是五角大楼反代吗🤣 还有一个比较典型的问题,是作者对于传闻中DeepSeek 将部分请求路由到 Fable 表示不理解,认为这样算不过来经济账。 实际上,但凡是个从业者或者稍微了解模型公司的研发流程,就知道这件事情不能简单按照 API 成本计算。 如果一个 frontier model 可以帮助你生成高质量训练数据、评测数据,或者提升模型迭代速度,那么获取这些能力本身就是一种研发投入。 用单次调用价格去判断整个策略是否划算,本身就是把模型公司的研发逻辑想简单了。 但最离谱的部分,是 PDF 里面声称 @Kimi_Moonshot 在 K3 发布前,把整个 RL 团队全部解雇了❓❓ 这个说法目前没有任何可靠依据。 实际上情况恰恰相反,根据我的确认Kimi RL 相关团队目前仍然存在,并没有所谓“整个 RL 组被裁撤”的情况。 甚至 PDF 里面还进一步延伸出一套完整叙事: K3 = 蒸馏 + 刷榜 + 裁撤 RL。 我一开始还在认真的区分哪些是事实,哪些是推测,哪些只是作者脑补出来的故事。 直到我看到 PDF 后半部分开始大量加入意识形态化表达,把技术讨论上升到所谓“国模黑暗时代”“技术偷取”等叙事,我反而释然了。 因为这已经不是在分析 AI 行业,而是在借 AI 叙事讲另一个故事。 PS:我感觉这个PDF是黑KIMI来的,因为KIMI的篇幅最大,但里面的大部分内容都是假的,其他家就有真有假了。


Have Chinese AI Models Caught Up to the US Frontier? I have spent the last 2 days writing this article. It should settle the debate once and for all. open.substack.com/pub/scaling01/…



ok just spent a morning with Kimi K3 as my firstmate, here's my real experience 1. it's very, very slow potentially due to the fixed max reasoning. you should expect the experience of something slightly slower than fable 2. its claimed cost efficiency is not manifesting in real economics i bought the $40 plan, and a few prompts later it's already eaten 1/3 of my 5-hr limit - it was in a single session and my context window was only 200k long at that time i don't care what the benchmark numbers say, and what the face value API pricing is, in reality Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan - i observe no efficiency benefit 3. its instruction following capability is weaker than other frontier models firstmate stretches frontier models' reasoning capability and is a really good test that can quickly reveal how good a model is at following instructions the pure "intelligence" of K3 does hold up - it understands my intent very well, and can diagnose problems, delegate tasks all fine but i very quickly noticed many instructions in firstmate's system prompt not strictly followed by Kimi K3. these were never a problem with gpt 5.5, 5.6, opus, fable and grok 4.5 so all in all, i'm now very skeptical of the claimed performance and going to keep my eyes wide open on its true capability



















