nsrunloop

89 posts

nsrunloop banner
nsrunloop

nsrunloop

@dare0lu

Founder, builder, engineer. Previously Cutlabs, TL;DW AI and https://t.co/NqAf1Yxs6p. Now https://t.co/FXnFVqAq5r

Sunnyvale, CA Katılım Ekim 2016
100 Takip Edilen18 Takipçiler
nsrunloop retweetledi
BridgeMind
BridgeMind@bridgemindai·
New CursorBench results are in and Opus 5 is the story. Fable 5 Max: 70.5% at $17.32 per task. Opus 5 Max: 70.0% at $8.23 per task. Half a point behind. Less than HALF the price. 40% fewer tokens. And it gets better down the card. Opus 5 Extra High beats Fable 5 Extra High outright and costs $4 less per task. Anthropic just made their own best model obsolete on value. This is huge.
BridgeMind tweet media
English
52
76
1.5K
82.8K
nsrunloop
nsrunloop@dare0lu·
Wow
Guillermo Rauch@rauchg

Fable just found a 15-30% memory efficiency improvement in Turbopack / Next.js, nearly autonomously. @tobi asked me today: what have been your “holy s***” moments with AI? My answer was: it’s every single week. And it’s accelerating. In fact, “WTFs/day” might just be my favorite metric for AI progress. It’s one thing to read benchmarks or stories online. It’s another to watch these machines pull engineering feats every day. 3 days ago? Sol helped us find novel vulnerabilities in some of the most audited code in the world. Today? Fable helps us ship this large optimization of a very complex Rust codebase. I just saw some results of work we’ve done to shrink binaries by 10-20x. List goes on.

QST
0
0
0
53
nsrunloop retweetledi
程序员不鸭
程序员不鸭@longyuli8·
GLM是三家里技术比较一般的,所以只能从战争部反代mythos+蒸馏。 Kimi+就比较强了,他们有时空穿梭技术,他们把七月份的fable api拉了个隧道拉到四月份,所以他们蒸了三个月的fable也就更强。 Deepseek是三家里技术最强的,他们有因果律武器,能直接把deepseek v4 pro正式版变成fable。所以很多人以为灰度的是 fable也没错,因为二者本就是同一模型。 (转) 笑死我了他妈的
Max For AI@MaxForAI

今天晚上各个群都在传一份《国内大模型蒸馏风波的来龙去脉》。 我大概看了一下,很可惜,这更像是一份由外行根据各种流言拼凑出来的 AI slop。 里面的内容有真有假,但几个核心判断基本都经不起推敲。 比如里面声称,智谱 @Zai_org 在今年 4 月份就已经破解并开始蒸馏 Fable。 这个时间线真的很离谱。 因为 4 月份的时候,Fable (那个时候还是Mythos)甚至还没有正式发布,能够使用的基本都是 Anthropic 的内部测试用户,通过Project Glasswing进行访问。 那智谱是怎么拿到的访问权限呢? 难道是 @AnthropicAI 的邀请或者是五角大楼反代吗🤣 还有一个比较典型的问题,是作者对于传闻中DeepSeek 将部分请求路由到 Fable 表示不理解,认为这样算不过来经济账。 实际上,但凡是个从业者或者稍微了解模型公司的研发流程,就知道这件事情不能简单按照 API 成本计算。 如果一个 frontier model 可以帮助你生成高质量训练数据、评测数据,或者提升模型迭代速度,那么获取这些能力本身就是一种研发投入。 用单次调用价格去判断整个策略是否划算,本身就是把模型公司的研发逻辑想简单了。 但最离谱的部分,是 PDF 里面声称 @Kimi_Moonshot 在 K3 发布前,把整个 RL 团队全部解雇了❓❓ 这个说法目前没有任何可靠依据。 实际上情况恰恰相反,根据我的确认Kimi RL 相关团队目前仍然存在,并没有所谓“整个 RL 组被裁撤”的情况。 甚至 PDF 里面还进一步延伸出一套完整叙事: K3 = 蒸馏 + 刷榜 + 裁撤 RL。 我一开始还在认真的区分哪些是事实,哪些是推测,哪些只是作者脑补出来的故事。 直到我看到 PDF 后半部分开始大量加入意识形态化表达,把技术讨论上升到所谓“国模黑暗时代”“技术偷取”等叙事,我反而释然了。 因为这已经不是在分析 AI 行业,而是在借 AI 叙事讲另一个故事。 PS:我感觉这个PDF是黑KIMI来的,因为KIMI的篇幅最大,但里面的大部分内容都是假的,其他家就有真有假了。

中文
13
36
442
118.3K
nsrunloop
nsrunloop@dare0lu·
According to my close friends who work on token business, people do spent millions of dollars on tokens in a single day or two, from multiple sources. That is definitely distilling under the hood. That is viable. But of course the allegation is not based on pure evidences, or exaggerated.
Chris Paxton@chris_j_paxton

Fable 5 was released June 9; kimi k3 was released July 16. How exactly could they have gotten enough data and trained a frontier model, much less tested it, in that one month If they have some way to distill a frontier model with all its capabilities from such minimal data that's actually a huge accomplishment But unfortunately I just think this post is mostly wrong

English
0
0
0
27
nsrunloop
nsrunloop@dare0lu·
Training and distilling are technically the same, except for data quality
Ibrahim Dagher@IbrahimDagher20

“Anth/OAI distilled the internet, so China’s just doing the same” is an unbelievably bad take. (1) Anthropic paid the largest settlement in history for the illegal books it downloaded — good luck getting China to do the same for mass violations of ToS (2) Training on internet data is fair-use: it’s incredibly transformative to take text and then use that to build a machine which can think. Training on copyrighted text is totally fine, so long as you buy the book. What’s illegal is *downloading those texts without paying for them*; and Anthropic now does pay for every book when it trains a model. What China does is the complete opposite: it is a direct breach to adversarially distill Fable’s outputs to another model. If you are analogizing “training” to “distillation”, you should know that *training* was never illegal; whereas distillation is a legal breach. Imagine you read a science textbook that you pirated, use that knowledge to invent with a new technique for building a new drug, and then start selling the new drug. Then, a competitor breaks into your lab, steals your technique, and uses that to sell a duplicate drug. That is theft, and it does not matter that you pirated the original textbook — you pirating the textbook means you should pay the original textbook authors, but it does not mean that your innovations are now open to be stolen by anyone! (3) Setting aside all that, the biggest issue here is national security. Anthropic spends billions in R&D and compute to push the frontier of reasoning, why should Chinese labs be able to just train on those reasoning traces to uplift their models? This is a horrible setup for China-US competition, and I am frankly baffled by how Americans could be okay with that, if they are AGI-pilled in even the slightest way.

English
0
0
0
32
nsrunloop
nsrunloop@dare0lu·
Btw, open weight != open source. I have not seen any real LLM players open source, ever. (Maybe some portion of Llama?)
English
0
0
0
14
nsrunloop retweetledi
Tom Yeh
Tom Yeh@ProfTomYeh·
Backpropagation by hand ✍️ ~ 11 steps walkthrough below Backpropagation is the algorithm that actually trains a neural network, and it is where most people stop following along. It is not calculus you cannot do. It is matrix multiplication, working backward, one layer at a time. So I drew and calculated one entirely by hand. Goal: push the loss gradient back through a 3-layer network and land on a new value for every weight and bias. = 1. Given = A 3-layer perceptron, an input X, predictions Ypred = [0.5, 0.5, 0], and the truth Ytarget = [0, 1, 0]. = 2. Backprop gradient cells = Let us draw empty cells for every gradient we are about to compute. The shape of the answer comes first. = 3. Layer 3 softmax = We get dL/dz3 straight from Ypred minus Ytarget = [0.5, -0.5, 0]. No chain rule needed, and that shortcut is the whole reason softmax and cross-entropy are paired. = 4. Layer 3 weights and biases = Let us multiply dL/dz3 by [a2 | 1]. One multiplication gives the gradient for W3 and b3 together. = 5. Layer 2 activations = We multiply dL/dz3 by W3 to get dL/da2. The gradient moves back across a layer the same way the signal moved forward. = 6. Layer 2 ReLU = Let us pass it through the gate: keep the gradient where the activation was positive, zero it everywhere else. = 7. Layer 2 weights and biases = We multiply dL/dz2 by [a1 | 1]. The same figure as step 4, one layer up. = 8. Layer 1 activations = Let us multiply dL/dz2 by W2. = 9. Layer 1 ReLU = We apply the same gate again, now on a1. = 10. Layer 1 weights and biases = Let us multiply dL/dz1 by [x | 1], and every weight in the network now has a gradient. = 11. Update = We subtract, and the network has learned. In practice a learning rate scales this step. The gradients: dL/dz3 = [0.5, -0.5, 0] dL/da1 = [1, -2, 2, -1] dL/dz1 = [0, -2, 2, -1] The takeaway: matrix multiplication is all you need. Just like the forward pass, backpropagation is matrix multiplications end to end. You can do every one by hand, slowly and imperfectly, which is exactly why a GPU's ability to do them fast mattered so much to deep learning. 💾 Save this post!
English
193
745
4.2K
938.2K
nsrunloop retweetledi
clem 🤗
clem 🤗@ClementDelangue·
We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!
Sam Altman@sama

we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…

English
416
921
11K
1.9M