Dylan Patel

13.9K posts

Dylan Patel banner
Dylan Patel

Dylan Patel

@dylan522p

SemiAnalysis Boutique AI Infrastructure Research and Consulting DMs are open for consulting, quotes, or to talk shop, Opinions my own

Katılım Nisan 2018
1K Takip Edilen152.3K Takipçiler
Dylan Patel
Dylan Patel@dylan522p·
As far as inference cost and intelligence per dollar, it is roughly equivalent now to the western closed models. Opus 4.8. is $5/$25 per m tok Kimi is $3/$15 so not a huge difference. Anthropic is known to have a ~75% gross margin on 1P API Chinese labs financials where they are known, show closer to 50%, so that mostly means pricing difference is margin, not efficiency at least on inference cost. I'm not sure how much training efficiency they have or which of these techniques (or similar) the closed labs have implemented yet. Training cost is hard to quantify, I dont think they have as much compute, but for the model run itself, not sure these's evidence they trained it for less, just they had way less research compute.
Dylan Patel tweet media
English
3
9
133
27.1K
Anika
Anika@AnikaSomaia·
@dylan522p @SemiAnalysis_ curious how you'd split the gap: how much closed by offshore rentals vs actual training efficiency (moe routing, int4-native, mooncake)? even if compute access isn't actually gated and and kimi isn't proof that efficiency beats flops, still not great for the compute moat thesis
English
2
0
23
17K
Anika
Anika@AnikaSomaia·
a week ago @SemiAnalysis_ wrote that Chinese labs are "simply too compute poor to truly reach the frontier." today one of those "too compute poor" labs, a 300-person startup actually, shipped a model that compares to opus 4.8 the entire western consensus – export controls, the $650B hyperscaler capex race, the "compute moat" investment thesis – is built on one assumption: flops gate capability. if that were true, chip controls would keep chinese labs permanently behind the frontier. but after reading through moonshot's stack i no longer think it is. training is efficiency-compressible: MoE routing, INT4-native quantization, better data curation, infra built around scarcity (their Mooncake stack exists because they don't have gpus!). a small lab with taste can compress the compute needed to make a frontier model, even if it can't afford to serve one the frontier is no longer something money can buy
Arena.ai@arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

English
140
192
1.8K
787.8K
Tony Zhao
Tony Zhao@tonyzzhao·
This is the first technical report from @sundayrobotics. Lot's of new insights about the result and how we approach the problem differently. Link in the thread!
English
9
13
383
48K
Ren
Ren@ren_stocks·
Time and time again. SemiAnalysis ‘calls it as they see it.’ @dylan522p, get some better glasses then, because you keep making the same mistake. Either you are right or Jensen is lying to all of us. Anyway, all I care about is this: ‘Giant amounts of production incoming.’ Bullish AI Infra.
Wall St Engine@wallstengine

Nvidia $NVDA CEO Jensen Huang directly DENIED reports that Vera Rubin is delayed. “The reports are not true. Vera Rubin is already in production. Giant amounts of production incoming,” Huang told Bloomberg. The denial follows SemiAnalysis reporting that Nvidia’s next-gen AI server rack system had faced delays tied to a specialized circuit board issue.

English
14
4
171
47.8K
Invest Like the Best
Invest Like the Best@InvestLikeBest·
SK Hynix recently reached a $1 trillion market cap. SanDisk started the year at $35B and is now worth $230B. Dylan (@dylan522p) on why memory prices will still double or triple from here: "People are like, the memory story is overplayed. No, you don't get it. DRAM will double or triple from here still, because that's how much capacity is required. Memory can only grow capacity low double digit percentages a year. 20-30% a year, even less for NAND, a little bit higher for DRAM. Even though the demand signal was very strong at the end of 2025, the memory companies immediately started reacting. None of that incremental capacity really gets here until 2028, even if they wanted to build as fast as possible. So the result is memory prices have gone through the roof. They're going to double and triple again. DRAM, especially. And they have to steal capacity from somewhere else. And the only way to steal capacity from somewhere else in a capitalist economy is demand destruction via higher pricing." Publish date: April 23, 2026
Patrick OShaughnessy@patrick_oshag

Every conversation I have with @dylan522p, I'm really just trying to understand the supply and demand of tokens. This is a unique episode in that it's entirely dedicated to talking about both sides of that equation. We discuss: - The infinite demand for the newest models - @SemiAnalysis_ going from $10K on AI spend to $7M - Mythos and Anthropic's compute problem - Why TSMC spending $100B on CapEx could cause a shortage - Robotics as next demand wave - Why memory prices will double again This is my second conversation with Dylan and find myself needing to speak with him more and more often to make sense of it all. Enjoy! Timestamps: 0:00 Intro 1:00 Surging AI Spend 10:27 Token Demand 16:21 When Ideas Are Cheap and Execution is Easy 20:46 Model Hoarding 22:34 Robotics 27:03 The Compute Bottleneck 30:26 The AI Permanent Underclass 31:39 Supply Chain Reality 37:47 CPUs 42:54 Predictions: Public Backlash

English
8
30
396
104.3K
Dylan Patel
Dylan Patel@dylan522p·
@Salehaldin0b Scale up CPO on the GPU is not ramping. That's what we said is delayed.
San Francisco, CA 🇺🇸 English
1
0
1
764
Salehaldin
Salehaldin@Salehaldin0b·
@dylan522p Gilad: “CPO now moving to full production. Going to see lot of cpo being installed in data centers around the world” $LITE $BESI
English
1
0
0
395
Dylan Patel
Dylan Patel@dylan522p·
@alialagheband @wallstengine If you read our actual reports we agree. We are really deep in the weeds on the exact architectures as it impacts supply chains a lot.
English
1
0
4
1.1K
Wall St Engine
Wall St Engine@wallstengine·
Nvidia $NVDA CEO Jensen Huang directly DENIED reports that Vera Rubin is delayed. “The reports are not true. Vera Rubin is already in production. Giant amounts of production incoming,” Huang told Bloomberg. The denial follows SemiAnalysis reporting that Nvidia’s next-gen AI server rack system had faced delays tied to a specialized circuit board issue.
Wall St Engine tweet media
English
15
36
314
101.3K
Dylan Patel
Dylan Patel@dylan522p·
@peytonroym @ren_stocks We've been very public about helios rack shipments being delayed out of this year, yes? AMD denied those too, but we are still right
English
0
0
7
933
WallStreetCC
WallStreetCC@cecilialiu_ch·
@SemiAnalysis_ @dylan522p Jensen Huang directly DENIED Semi reports that Vera Rubin is delayed. “The reports are not true. Vera Rubin is already in production. Giant amounts of production incoming,”
English
1
0
0
540
SemiAnalysis
SemiAnalysis@SemiAnalysis_·
AMD Helios. NVIDIA Vera Rubin. Double-wide racks. Liquid cooling. And why storage just became the main event. @dylan522p sat down with @Supermicro_SMCI Chief Business Officer Vik Malyala for a wide-ranging conversation. Full interview below. (1/3)🧵
English
10
16
268
71.8K
Dylan Patel
Dylan Patel@dylan522p·
@NotYosyp I stole it from my british analyst 🤮😭
San Francisco, CA 🇺🇸 English
0
0
15
2.5K
Al
Al@forest92419017·
@dylan522p One of your worst takes yet. Anthropic would still exist regardless of whether Sam had been outspoken about AI… plenty of others in the field already recognised its potential.
English
2
0
15
3.9K
Dylan Patel
Dylan Patel@dylan522p·
Or he failed to keep it together (Anthropic, Thinky, Core Automation, etc) and bloviated so much about his tech that everyone saw it as existential and it created completion from all the hyperscalers. If he played his cards right I'd bet there'd be no Anthropic or Meta TBD or XAI
shafin@_shafinsiddique

sam altman is probably the greatest ceo of our time. he practically has every major tech CEO investing hundreds of billions of dollars to compete with him, sometimes even colluding (see elon + zuck). I don't know if gates, zuck, or even elon had this level of competition when they were forming their companies. OpenAI still consistently churns out some of the best models and still emerges as the winner in almost every category. my only question is what did @paulg see during that 10 min YC interview

Nashville, TN 🇺🇸 English
97
43
1.6K
359.7K
Carnage
Carnage@0xCarnagee·
Prime Intellect engineer: "everyone's bragging about a million-token context. here's what they don't tell you. at 256k tokens GPT-5.5 scores 80% on retrieval. push it to a million and it drops to 36%. the model accepts the context, it just can't reason across it. people call it context rot." in a 20-minute talk he explains why bigger context windows won't save your agents. continual learning + training on your own traces + real environments - that's the fix. Watch the talk, then save!
Carnage@0xCarnagee

Andrew Ng just dropped a free course on Claude Code from scratch, taught with the Anthropic team: 00:00 - why Claude Code is so agentic 04:00 - shockingly simple architecture 12:00 - point it at any codebase this short watch will replace 10 paid coding agent courses. Andrew Ng calls it his personal favorite coding assistant right now. Watch it today, then read how to engineer your own agent loops in the article below

English
90
354
4.5K
740.8K
Jim Liu
Jim Liu@jiahanjimliu·
$IREN: SemiAnalysis Hit Piece In The Works @SemiAnalysis_ is a great research shop except for when they write the occasional hit piece like recently on Photonics and Nvidia. I'm hearing a rumor that @SemiAnalysis_ is preparing a hit piece on $IREN's Childress site while it was ramping. They have old satellite images that show H1 was using chiller trucks. These were temporary and H1 is now on permanent stationary chillers. I'm not sure what else they have in store but if the hit piece ever launches, know that it will be dispelled when IREN announces GB300 exemplar status and handover H1 to $MSFT. Coincidentally, this algins with rising short interest from 15% to 22.51% since May (1). This is before the GSW Marketing and RSU compensation package so it's a short prepared in advance of those two events and will likely top of with a SemiAnalysis Piece. H1 handover late July. Hit piece will not matter afterwards.
English
49
24
386
156.9K