Mark

4.3K posts

Mark banner
Mark

Mark

@MarrrkLoki

EN/FR/中文

🦫 Katılım Mayıs 2010
1.3K Takip Edilen208 Takipçiler
Mark
Mark@MarrrkLoki·
@axiomofmind @talemi @BrianRoemmele The only limitation is your imagination. I can deploy the same workflow on a DGX spark, and run a 3 bit quant on SSD streaming with much higher cache rate and tk/s. Probably still slow but good enough to let the agent work overnight and check the results in the morning.
English
0
0
0
13
@Axiomofmind ⚡
@Axiomofmind ⚡@axiomofmind·
@talemi @BrianRoemmele So you'd rather run a simple task at 0.32 tk/s over a long time instead of just using something faster to achieve the same results in 100x less time?
English
5
0
0
119
Brian Roemmele
Brian Roemmele@BrianRoemmele·
BREAKTHROUGH: A full, unmodified 2.78-trillion-parameter Kimi K3 on a consumer laptop by streaming only the activated experts from NVMe. YOU CAN’T RUN KIMI K3 “ON THAT” THEY DECLARED. There are many paths to do it. This is one: Marco Bambini Just Gave Us the Full Kimi K3 on a Laptop Meet Marco Bambini he did something that felt impossible only a day ago. He built WASTE Weight-Aware Streaming Tensor Engine a clean, dependency-free C inference engine that runs the complete, unmodified 2.78-trillion-parameter Kimi K3 model by streaming only the activated experts straight from NVMe. No distillation. No pruning. No cloud. The full open-weight model. We have it running in the lab right now. What Marco Actually Built Kimi K3 is a sparse Mixture-of-Experts system. Only about 4 % of its weights fire on any given token. Marco’s insight was simple and ruthless: the idle experts do not need to live in RAM. They only need to be reachable in time. WASTE keeps the model’s “trunk” (attention, shared components, embeddings) resident in memory — roughly 27 GB on the converted container. The 82,000+ routed experts stay on disk as tightly packed residual vector-quantized records. When the router selects its 16 experts per layer, the engine issues direct, cache-bypassing reads from the internal NVMe and feeds them into a bounded expert cache. The rest of the machine’s RAM becomes working space for that cache. On a 64 GB MacBook Pro with the container on the internal SSD, we are measuring 0.32–0.34 tokens per second at a comfortable memory budget. Prefill sits a little higher. The vision tower works. Logits match the reference implementation to within a few parts in a million. It is the real model. The container itself is 982 GiB after conversion from the original 1.42 TB MXFP4 weights. Minimum RAM floor is just over 29 GB for short context. Push the budget higher and the expert cache hit rate climbs; push too high and you start paging and the speed collapses. The sweet spot on current consumer hardware is clear and measurable. How We Are Testing It We converted the official weights, verified the container, and began systematic runs the same day the engine stabilized. First we confirmed numerical fidelity against the PyTorch reference on short prompts. Then we moved to longer generation, vision inputs, and multi-turn chat using Kimi’s native XTML format. We are measuring wall-clock decode, expert I/O versus compute split, cache hit rates at different RAM budgets, and thermal behavior under sustained load. We are also exercising the OpenAI-compatible server that sits on top of the same C library so we can drop the model into existing agent loops without rewriting anything. Early observations: •Expert I/O dominates the timeline, as expected. On a fast internal NVMe the engine is already near the practical ceiling of the storage subsystem. •The architecture’s sparsity is the entire enabler. A dense model of this size would be dead on arrival for local use. •Context length is currently limited by RAM more than by the model itself. Practical working contexts sit comfortably in the tens of thousands of tokens on 64 GB hardware; the full million-token window will need more memory or smarter KV management. •Thinking tokens are expensive at this speed. Long internal monologues turn into multi-hour runs. For agent work we are already experimenting with tighter control over when full reasoning is requested. We are treating this as a research instrument, not a finished product. Every run teaches us something about expert locality, prefetch opportunities, and how far pure software streaming can push trillion-scale inference on ordinary machines. 1 of 2
Brian Roemmele tweet media
English
42
98
579
70.8K
Jacob Grandstaff
Jacob Grandstaff@JDGrandstaff·
@mindyisser This is great, actually. This scumbag was likely giving her favoritism in the workplace that should have gone to hardworking young men. Not to mention, cheating on his wife. He should be imprisoned.
English
44
6
606
116.4K
Mark retweetledi
Alicia, Courtyard Urbanist
Alicia, Courtyard Urbanist@UrbanCourtyard·
Do you have 9 acres? I have a plan you will love. One scenario is 1,220 units on floors 2-6. 122 studios, 305 1BDs, 488 2BDs, 244 3BDs, and 61 4BDs. Plus about 177,000 sf of ground floor for whatever.
Alicia, Courtyard Urbanist tweet mediaAlicia, Courtyard Urbanist tweet mediaAlicia, Courtyard Urbanist tweet media
Alicia, Courtyard Urbanist@UrbanCourtyard

Remember this astonishing proposal for the Sunset site? As part of today's presentation, I will be showing what a courtyard block development could look like on the same site--and others in the area. Including square footage, unit mix, etc.

English
57
103
1.3K
119K
Brad
Brad@deorcmona·
@SombraMapache People who lost their families/ loved ones in this attack have the right to be racist.
English
105
2
156
33.7K
Soma - Patron Saint of Christmas
Call me woke or whatever, but I don't think it's even vaguely reasonable to expect the literal mayor of NYC to be prohibited from attending a 9/11 memorial just because a handful of folks affected by it grew up to be malignant racists.
Kosher@koshercockney

BREAKING 🔴 Mamdani states he WILL be attending the 9/11 commendations despite survivors and victim families stating they do not want him there. He’s trolling them in my opinion, and it’s disgusting. There’s thousands of signatures now. Get this Islamist out.

English
2.3K
13.8K
174.4K
3.6M
Mark retweetledi
Jason Kenney 🇨🇦🇺🇦🇮🇱
Message to patriots who think this is done & dusted. It's not. Imagine 40% turnout. That's what we had for the standalone equalization referendum in 2022. The 10% or so of the electorate who are hard core separatists will crawl over broken glass to vote. Not hard to imagine another 10% who are frustrated federalists showing up to cast a "send a message" vote. Then you've suddenly got 20% of the electorate winning, and pushing us into wood chipper for the next year. Don't let it happen. Get involved in @VotetoStayAB or another group that supports a strong Alberta within Canada. And don't take for granted the overwhelming support for national unity. It only matters if the silent super majority show up to vote.
Vote to Stay@VotetoStayAB

BREAKING: One of the largest publicly released polls on Alberta's Oct. 19 referendum – conducted by Innovative Research Group – is in. 🚨 The separatists are losing in every region of Alberta. 📊 Proud Albertans and patriotic Canadians have the numbers to win and win big. But polls don't decide referendums – people do. 🙋‍♀️ The separatists will show up on Oct. 19. We need to do the same. Pledge your vote at VotetoStay.ca and help us defeat separatism for a generation! 🗳️ #abpoli #cdnpoli #votetostay

English
390
189
751
97.7K
Mark
Mark@MarrrkLoki·
@lifemaxed @sudoingX Here's the neat part, with open weight you don't need to trust anyone and nobody can gatekeep anything. My local funded community library can easily run a decent enough server farm and anyone with a digital library card can use it. The only limitation is your imagination.
English
0
0
0
9
Max Life
Max Life@lifemaxed·
@MarrrkLoki @sudoingX Thats a cool idea but having to trust liberal government Dems gatekeeping and tracking that local inference is less appetizing, maybe more of a co-working space would be better.
English
1
0
0
11
Sudo su
Sudo su@sudoingX·
dear open labs, and i mean this with love. you keep dropping these trillion parameter giants and they're genuinely beautiful and i can't run a single one of them, none of us can, they land on huggingface to a standing ovation and then they just sit there, because the only machines that can load them are the same datacenters we were all trying to get free of. i watched local ai turn into something real this year, people running 27b models on one gaming card, an 8gb board doing work that needed a server a year ago, whole little communities forming around whatever fits on the hardware already sitting on the desk, and that's the part that actually reaches people, the revolution nobody's putting on a keynote slide. so here's the fair ask, give us the 40b dense, give us the 120b moe, the sizes a person with one good card or a 128gb box at home can actually hold and run and learn from and build on, because not everything has to be shaped for an enterprise with a rack to spare. the frontier is thrilling and i'll keep cheering every open weight you ship, i mean that, but the ground is where most of us live and the ground is quietly running out of models it can lift. build for the hardware we already own. that's where the next wave comes from, it always was. with love, from someone running yours on a laptop at midnight.
English
73
88
684
56K
Mark
Mark@MarrrkLoki·
@aye_pete @jacobeverly You can't cheat to create a front tier AI. Distillation doesn't work that way. That's why only the most gullible people believe the lies from anthropic and the Trump admin, and all the researchers in the field push back against this.
English
0
0
0
18
Peter
Peter@aye_pete·
@jacobeverly Open weights are good, by the way. But you can't host them on normal infra. Cheating just kills the incentive to put in the work to drive innovation as fast as it is going right now.
English
1
0
0
57
Mark
Mark@MarrrkLoki·
@aye_pete @jacobeverly "why do they expend so much effort and resource to do it" You automatically assume and believe the premises where there's no evidence presented to back it up. His argument on the other hand is consensus within the AI field and researchers. At least look up how distillation works
English
0
0
0
13
Peter
Peter@aye_pete·
Bro, your argument is dumb. If it's so negligible, why do they expend so much effort and resource to do it? Why not just train their own Chinese models and compete directly instead of distilling US Frontier models? Oh, right. They were left far behind and the only way to catch up was to cheat.
English
2
0
0
118
Mark
Mark@MarrrkLoki·
@earn35tguy @sudoingX @mkratsios47 Distillation doesn't work that way. Model can't gain real capability from learning the style of how other models repond to the problems. Exhibit 1: Every Qwen community fine-tune variants with high quality Opus distillation always regress in benchmark comparing to the base.
English
0
0
0
17
Earn35t Guy
Earn35t Guy@earn35tguy·
@sudoingX @mkratsios47 You have a point but missed the point here. The point was, K3 took a shortcut, and was not a first class model trained from scratch. Orders of magnitude less work
English
3
0
1
250
Director Michael Kratsios
Director Michael Kratsios@mkratsios47·
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model. To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.   The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
English
4K
1.4K
14.9K
9.3M
Mark
Mark@MarrrkLoki·
@LonghornTXFan89 @Teknium @mkratsios47 And even so, there are tons of evidence that anthropic distilled Qwen and you can still reproduce that today by asking Claude what model is it in Chinese, and it'd say it's Qwen. So the entire distillation argument they use is dishonest and shameless to the core.
English
0
0
3
26
Mark
Mark@MarrrkLoki·
@LonghornTXFan89 @Teknium @mkratsios47 In common sense, IP is in the algorithm, architecture, training methods, etc. No serious person in AI field would consider the output of an AI model as IP.
English
1
0
4
36
Mark retweetledi
0xSero
0xSero@0xSero·
Most goated nerd
Max For AI@MaxForAI

‼️突发:DeepSeek @deepseek_ai 创始人梁文锋在一场据称长达四小时的投资者会议上的发言内容曝光。 过去一个月里,这场对话的片段一直在中国 AI 圈和投资人圈内流传。 最引人注目的是梁文锋反复强调的“不”字。 不搞“天才”神话。 不追求利润最大化。 不搞闭源。 不盲目争夺用户。 不做视频生成、3D、世界模型,也不做下一个“超级应用”。 他的解释很简单: “克制是一种策略。放弃一些东西,是为了提高实现 AGI(通用人工智能)的概率。” 以下是会议中的核心观点: - DeepSeek只有一条主线:AGI 梁文锋认为,现在不是追求产品收益最大化的时候。 核心是追逐人工智能最高峰,即AGI。 他的AGI路线图非常清晰。 去年的台阶是思维链,今年的台阶是Agent。 Agent之后,是持续学习。 持续学习之后,是AI自我迭代。 然后,AI开始帮助人类研发更加先进的AI系统。 最后才是具身智能。 梁文锋将这个过程称为“渐进的奇点”。 根据这个思考,DeepSeek目前正在主动放弃许多热门方向。 “很多东西不在我们的主线上,包括3D和视频生成。” 他也质疑了世界模型的重要性: “世界模型跟智能的上限没有太大关系。” 多模态对产品和C端用户很重要,但它仍然只是一个组件。 “它不是主线,也不是智能本身。” - 下一代模型必须拥有持续学习能力 人类可以在工作中不断学习。 但AI每次开始一个新任务,都需要人类重新提供所有相关上下文。 “这几乎是不可能的。” 这也是梁文锋认为,当前AI还无法真正替代员工的原因。 “下一代模型必须具备持续学习能力,否则就不能叫下一代模型。” 同时他表示智能的终点可能是具身。 他的逻辑很简单: 人类真正需要的,并不是另一个计算机界面。 人类需要的是劳动力。 - DeepSeek不追求利润最大化 梁文锋表示,DeepSeek不会以收入最大化为目标进行定价。 “我们只赚一个合理的利润。” DeepSeek V4最初担心需求过多,所以价格定得比较高。 后来,DeepSeek把价格降到了四分之一。 “公司群里很多人都在欢呼。” 因为做出这个模型的目的,就是让更多人能够充分使用它。 备受关注的低成本不仅是一种商业策略。 “低成本是架构带来的结果。” 更低的成本可以让AI变得更加普惠。 但还有一个更深层的原因: 单位计算成本越低,DeepSeek就越能够在有限算力下训练更大的模型。 “大公司可以通过增加资源解决问题,我们优先考虑成本效率。” 有意思的是梁文锋并不觉得API生意有多么好。 “我不觉得卖API这个事情有那么大的吸引力。” DeepSeek只需要一个很小的团队维护服务。 几乎没有客服,也不需要销售,用户自己就会来。 “我们一直在商业化,只是不以商业化为目标。” 他认为,DeepSeek彻底转向商业化的时间点仍然非常遥远。 - 开源是DeepSeek的战略甜蜜点 对于员工来说,开源能够带来成就感和组织凝聚力。 对于社会来说,它能够帮助研究者、企业和普通用户。 “开源是一种让利。” 梁文锋认为,AI最终会成为一个规模极其庞大的产业,任何试图垄断它的人,最终都会失败。 “如果AI最终占到人类社会GDP的10%,任何想要独占这部分利益的人,都会被历史抛弃。” 梁文锋表示,DeepSeek不会开源一个较弱的版本,同时在内部保留更强的模型。 他不担心竞争对手部署DeepSeek模型。 小型创业公司可能缺乏资源和意愿,去进行前沿模型研究。 大型公司可能拥有资源,却存在组织上的困难。 梁文锋认为,DeepSeek刚好处于一个罕见的中间位置。 “这是属于我们这个规模公司的sweet point。” - 中美AI的差距主要在资源 梁文锋认为,中国并不真正缺乏人才。 DeepSeek希望改写中美AI竞争的叙事: 用几分之一的算力,把差距缩短到6个月,甚至3个月。 他作出这个判断的原因是他依然相信Scaling “我们相信Scaling,规模越大,效果肯定越好。” DeepSeek训练当前规模的模型,并不是因为这个规模已经足够。 而是因为现有算力只允许它做到这个规模。 梁文锋认为,前沿模型的竞争,最终会由三个因素决定: 成本、时间、用户体验。 “成本排在第一位。” 如果两家公司提供的模型质量相同,胜负将取决于谁能以更低的成本提供服务。 时间排在第二位。 早几个月和晚几个月,结果可能完全不同。 用户体验可以形成一定的黏性和壁垒,但梁文锋并不认为它是最本质的护城河。 - 团队稳定是梁文锋唯一不能让步的事情 他将其视为DeepSeek面临的最大风险之一。 原话是:“只有一个是没法退让的:必须保持团队的稳定性。” 而最近的一轮融资,已经大幅降低了这个风险。 “只要能够保持团队的稳定性,我一定能做成AGI。就这么简单。” 有趣的是DeepSeek的组织既是自上而下,也是自下而上 自上而下的部分,被梁文锋称为“做正事”。 但他不希望被安排的工作占用员工一半以上的时间。 另外一半时间应该保持自下而上。 研究人员可以自由探索他们认为重要的方向,不需要提前审批,也没有固定要求。 与很多AI团队不同,DeepSeek不鼓励过度加班。 “做研究需要一个相对松弛的环境。” 另一个原因是聚焦。 DeepSeek的许多产品并不完善,但公司会主动选择不去补齐所有问题。 梁文锋认为,这同样是一种克制。 有趣的是DeepSeek靠愿景驱动,而不是KPI “要实现某个KPI,不是我们的方式。” 梁文锋表示,DeepSeek的愿景甚至不一定被正式写下来。 它存在于公司的做事方式里,也存在于公司对待世界的态度里。 “愿景不是挂在墙上的标语,不是怎么说,而是怎么做。” 梁文锋最后的警告,可能也是最重要的一句话 “如果你的愿景是拿得更多,你就已经输了。你可能会面临更大的困难。这个世界就是这样。” DeepSeek不想赢下每一个AI市场。 它想保留实现AGI所需要的专注、成本优势、团队稳定性和组织自由。 而梁文锋似乎愿意放弃几乎所有其他东西,只为增加实现AGI的概率。 他在做真正很Cool的事情。

English
8
8
280
28.4K
Pedro Domingos
Pedro Domingos@pmddomingos·
@Dan_Jeffries1 No one is saying Kimi is 100% distillation, or that fast effective distillation is trivial, even when building on distillation of previous models. The point is Kimi wouldn't be where it is without it, and the whole narrative that it's honestly surpassed Fable etc. is BS.
English
7
0
3
1.2K
Pedro Domingos
Pedro Domingos@pmddomingos·
@Dan_Jeffries1 It was available for much longer than that, and they had plenty of time to distill Opus before. Also, stick to what’s relevant.
English
3
0
11
5.1K
Mark retweetledi
Andrew Curran
Andrew Curran@AndrewCurran_·
The Trump administration is considering an executive order, and other means, to ban Chinese open-source models within in the United States. Kimi K3 has reignited this debate. Reporting this morning by Axios. Commerce is also considering adding Chinese AI labs to the Entity List.
Andrew Curran tweet media
Andrew Curran@AndrewCurran_

There's not much information on this yet, and it's not clear if the executive order would target Chinese open-source models specifically or apply more broadly to open-source AI in general. But something is probably under discussion. The White House denies these reports.

English
414
454
2.7K
1.4M
Mark retweetledi
Kirill
Kirill@kirillk_web3·
> be Kimi Founder > Zhilin Yang. nobody in the West knows your name. > build K1. K2. K2.5. K2.6. K2.7. > each one sets the open-source record > everyone says China can't reach the frontier > raise $2B. hit $20B valuation. > keep your head down. keep shipping. > July 16. drop K3. > 2.8 trillion parameters. > the biggest open model ever released. > 1 million token context. native multimodal. > Kimi Delta Attention — 6.3x faster decoding. > Artificial Analysis ranks it #4 overall. > right behind GPT-5.6 Sol and Fable 5. > open weights July 27. free for anyone. > the closed labs charge for the frontier. > you gave it away. > different game.
Kirill@kirillk_web3

before you watch the World Cup final tonight, watch this 30-minute interview with Kimi founder Zhilin Yang it's the clearest look I've seen at how K2 was actually built and where AI models go next > why compute scaling still beats every other method > why the "data flywheel" everyone talks about hasn't formed yet > the real challenge in multimodality nobody mentions > why K2 was closer to a base model — the real potential wasn't unlocked yet this is only part of the interview, translated to English — and everything he hinted at just shipped as Kimi K3, 2.8 trillion parameters, the biggest open model ever full breakdown below

English
49
244
4.3K
412.3K