Gavin Baker

10.5K posts

Gavin Baker banner
Gavin Baker

Gavin Baker

@GavinSBaker

Managing Partner & CIO, @atreidesmgmt. Husband, @l3eckyy. No investment advice, views my own. https://t.co/pFe9KmNu9U

Boston, MA Katılım Temmuz 2011
6.1K Takip Edilen284.7K Takipçiler
Sabitlenmiş Tweet
Gavin Baker
Gavin Baker@GavinSBaker·
Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis. Rationale:   A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.    Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.   This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.   Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3. 
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead. Time will tell on both points. And likely fairly quickly. Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.
Gavin Baker tweet mediaGavin Baker tweet media
English
536
1.2K
7.9K
2.3M
Aaron Levie
Aaron Levie@levie·
This post is key. The cheaper AI gets, the more opportunity there is for the entire ecosystem - especially including end-customers - to benefit. Everything is bottlenecked by being able to successfully and cost effectively deploy AI in real workloads. Any time we can lower the cost of AI, the total usage goes up. When that happens, the value accrues to all layers of the stack, as noted. The only modification is I think as AI gets more efficient, demand even for the frontier closed models can go up too. You often need the strongest model possible for the orchestration of a task, then you can farm out work to cheaper or more tuned models for the bulk of tokens. This efficiency gain can ironically lead to even more frontier spend in the process because you can cost effectively deploy AI against more tasks. The thing at most risk is margins, which is fair, and in general it seems that intelligence eventually should converge with the margins of the infrastructure stack. Anyway, super fun times to see competition play out to drive down the cost of AI.
Gavin Baker@GavinSBaker

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis. Rationale:   A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.    Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.   This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.   Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3. 
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead. Time will tell on both points. And likely fairly quickly. Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

English
48
49
371
104K
Gavin Baker
Gavin Baker@GavinSBaker·
@DaveBlundin Thanks Dave! It’s been too long since we have caught up
English
7
0
42
23.9K
Gavin Baker
Gavin Baker@GavinSBaker·
Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis. Rationale:   A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.    Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.   This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.   Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3. 
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead. Time will tell on both points. And likely fairly quickly. Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.
Gavin Baker tweet mediaGavin Baker tweet media
English
536
1.2K
7.9K
2.3M
Patricio
Patricio@BrandoliniP_ARG·
@phithetasigma @GavinSBaker Your chart only includes $/1m tokens, but based on Gavin's Cost per Intelligence Index Task, it seems that GPT 5.6 and Fable consume much fewer tokens for the same tasks.
English
2
0
2
571
Zoe
Zoe@UltraRareAF·
@GavinSBaker appreciate the balance in your analysis here - imagine it must have been not very easy to write, but the clear-eyed take says everything thank you
English
1
1
41
22.2K
Gavin Baker
Gavin Baker@GavinSBaker·
@hitraderjoe I think the importance of OpenAI and Anthropic’s unfunded spending commitments to the infrastructure trade is wildly overstated.
English
11
8
185
58.2K
trader joe
trader joe@hitraderjoe·
Completely agree with this "Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software." but how do you think about the financing risks that run through OpenAI and Anthropic? In a perfectly rational world they would be able to offset excess compute to model hosting services/back to hyperscalers but the path may not be smooth
English
2
0
41
46.9K
Gavin Baker
Gavin Baker@GavinSBaker·
Risk/reward seems attractive again. Lots of cheap stocks with durable competitive advantages that are going to crush numbers for the next 6-12 quarters. Time will tell!
English
312
581
7.5K
3M
Clark Tang
Clark Tang@_clarktang·
I think there is general confusion around how AI works, AI tokenomics, and ultimately *what is actually priced in* for the AI trade - and that some of the existing arguments are at odds with one another Firstly to clear this up - what Brad and Gavin are saying are completely in agreement, what Gavin is laying out here is the *mega bull case* as he so states in the first sentence of his tweet lol The base case we are all living with is that the labs are going to continue to generate a significant amount of revenue this year and next year. OpenAI was already the fastest growing company of all time (and still is)... but Anthropic has just grown *SO* fast that OpenAI's growth look slow by comparison The basic chain for all of this together is as follows: Power (generation, interconnect, regulation) -> DC Shell (construction, equipment, regulation) -> Semiconductors (compute, memory, interconnect, adv packaging, wafer capacity) -> Hardware (networking, storage) -> Software (data, infra, inference) -> Models (open, closed, agentic loops, harness) How each of these interact with one another affects the ultimate cost - which is model cost Consider the following: Nvidia manufactures the bleeding edge chip for training and inference. It is very good at both training, and inference. Nvidia is the largest customer of TSMC, the memory players, substrates, lasers, transceivers etc - anything you can name on. And now to soon include power into this equation. The unit of compute is fungible because the software runs ubiquitously across all clouds, multiple industries, across all models. It is bankable by increasingly more financial institutions - infrastructure PE funds, even some IG debt now - because it is ubiquitous and observable what the market is. For this Nvidia charges the highest compute margins - ~80% on hardware. Consider the labs: Anthropic and OpenAI are inferencing across a fleet of *largely Nvidia / Google TPUs w/ some incremental gains of Trainium*. There are new entrants to the field - Cerebras, AMD, and potentially some 2027 tapeouts of new ASICs - OAI Jalapeno, new start ups etc. Anthropic and OpenAI make the best models, with a dominant share of wallet $ (Assume ~$100B ARR) at an estimated gross margin of ~70%. (economic estimates vary from 40-90% depending on what you are including). But almost certainly contribution margins on model inferencing is pushing the number higher than 70%. After establishing that though, I think it's incredibly important to state that while these things seems at odds with one another, this balance is not necessarily zero sum. The thought experiment Yes it is true that if Nvidia margins were 0, OpenAI and Anthropic could offer their intelligence at cheaper rates. How much cheaper? My estimate is NVDA DC = ~12.5B / yr Amazon Basics ASIC DC = ~$6B / yr (About 1/2 the cost - so if NVDA hardware is 2x the performance, then the cost advantage goes away - and actually that ASIC is worse off bc has much worse recontracting value so arguably depreciation curve should be shorter) So really, the labs cutting NVDA out could only offer the tokens at ~50% to 60% cheaper at their own economics. Is that signficant? Certainly. Is it an OOM difference? Not necessarily - so that's why they have prudent attempts to diversify away from NVDA (it's just good business), but they continue to rely (and actually if considering Ant's share gains, are increasing their spend on NVDA - while having competing programs). In the case of Open Source vs Closed - Nvidia obviously wants the proliferation of this because by definition all OS models will run best on Nvidia hardware out of the gate. Yes NVDA hardware will be good, but they will have this lead because of everything NVDA has been doing for the last 4 years in developing their platform ecosystem from the infrastructure (partnerships, funding, neoclouds) to the software (vLLM / other inferencing sw, inference clouds, Nemotron, NIMs, Nemoclaw etc), to install base (sovereign clouds, global partnerships, neoclouds, hyperscalers, etc) - to proliferate NVDA around the world. Anywhere there is inference that exists outside of a walled garden (the proprietary labs) - Nvidia will exist. The only ones who could potentially cut NVDA out are the labs. And the value that is captured from the labs are estimated to be in the hundreds to trillions of $ - which are obviously of much value to the world if it were offered much more cheaply. Which brings us to the debate at hand -- which one is right? The truth is no one knows. You can ask the labs, you can ask Jensen - anyone who tells you definitively is just lying to you. But you can build a plausible path to the future state using a few reasoning blocks. Here's a reasoning thread (feel free to generate your own thinking): - Bull case: Spend on the world's intelligence is about $30T / yr - What would you spend to augment that, maybe worth 30-50% of that? $10-15 T as a market? - Bear case: about 30M software developers in the world each earning $100K a year = $3T spend in salary. GitHub commits up 3x = $9T of productivity on $100B of ARR? *Even if you assume 90% of this is slop and useless, you would get $900B of ROI on $100B of spend* I have more reasoning chains, but I thought this one by Jensen was compelling - but this is where we can't give too much away :) But in spirit of crowdsourcing - some other interesting ideas I have that I am still thinking about (and encourage you all to consider as well): - Optimizations always happen - the question is just to what extent and for what reason - Agentic revenues was really what unlocked step function revenue growth - if open source is really just 6mo behind, then we should see really good agentic capabilities out of open models now too - Harness and model now tightly have to be integrated - Open Source never really makes sense as a sustainable business model - businesses investing at this scale always has to find a way to monetize that - "there is no free lunch" - not just a one model fits all... the only player that has an incentive to train on the frontier and keep completely free IS Nvidia - Rev / GW of AI labs are already nearing the highest metrics ever - now to be fair Meta and GOOG never really thought of Rev / GW as metric to lead their buildouts - was always a cost to doing biz - but it's not like we are being "stupidly inefficient" with power spend now - true mkt creation - wafer constrained, power constrained world. what's the optimal move?
Gavin Baker@GavinSBaker

The mega bull case for AI infrastructure would be *if* market share shifted away from certain frontier labs with 90%+ inference margins toward cheaper models, whether open-source or closed. It would increase the ROI on AI spend for end customers by increasing intelligence per dollar, which would drive incremental token demand. Margin dollars would effectively get redistributed from the frontier labs to AI infrastructure providers. The infra winners would be those with the lowest per token cost and the winners at the model layer would be those with the highest token efficiency. There are many reasons Jensen is so focused on open source, but this is likely the most important one as I think he is probably less worried about a monopsony these days. Lower margin % at the model layer = more margin $ at the infra layer all else equal. With SpaceX and Meta being vertically integrated and possessing the #3 and #4 models respectively it is more possible than ever. Note that Grok 4.5 is ahead of Fable for some useful tasks at a much lower cost, so ranking them #3 is conservative. This is not happening yet. Cheap, mostly open source tokens are likely the majority of volume today but the majority of economic value is still accruing to the most intelligent models. Might change though. We will see.

English
38
50
564
163.3K
Gavin Baker
Gavin Baker@GavinSBaker·
@PythiaR Laughing. Knew I should have fact checked the 7b.
English
2
0
12
3.4K
Gavin Baker
Gavin Baker@GavinSBaker·
Curious to see how many subs @SemiAnalysis_ has in South Korea and whether they are more or less respected there than KIS, whose spec sales note was apparently a contributing factor to last nights sell off. A modern day Clash of the Titans. Should know in a few hours!
Gavin Baker tweet media
English
39
66
924
416.9K
Gavin Baker
Gavin Baker@GavinSBaker·
Much stranger things have happened but would still be funny. Also the 1981 “Clash of the Titans” has aged really well even if the Kraken has not.
Gavin Baker tweet media
English
6
1
64
38.9K
Gavin Baker
Gavin Baker@GavinSBaker·
It would be pretty funny if Hynix missed the quarter immediately after their $7b (or whatever it was) US ADR listing and super bullish roadshow where they made fun of Micron for agreeing to price ceilings in their LTAs. Semianalysis well above consensus for this Q, KIS spec sales evidently below. Time will tell!
English
14
8
270
57.9K
Gavin Baker
Gavin Baker@GavinSBaker·
@BeenThereCap Fully agree with all of this - sometimes not room for nuance in investing. Doubling down late is certainly not always the right thing to do!
English
0
0
7
1.3K
₿ΞΞnThereDoneThat Capital 賢い
Understanding that Gavin is mostly right here, but also partially wrong, and learning how to discern when you are likely right vs. the market and when you're likely wrong, is the most important skill in investing. Gavin is entirely correct that you need to learn whether you are a panic-early person or a double-down-late person. He is half-correct (and I'm sure he knows it; I'm adding nuance) on "if you're losing money, you're wrong." Every single alpha-seeking investment has a thesis that the market is currently wrong and will become more right. But you can lose enormous $$ on a single stock if the market moves a long way against you, you stay stubborn, and you double down (or more). John Hempton's blog post, "When do you average down?" is the definitive word on this subject. Unfortunately for newbies, acquiring this collection of skills takes most people years of experience -- hard-knock lessons, learning yourself, pattern recognition, and developing position sizing and trading plans & patterns that fit your innate style. I am 20+ years in and have still gotten better in the past year.
dnap@dnapway

Gavin Baker reveals the investing advice he always comes back to: "panic early or double down late" "If you're losing money, you're wrong. It's not that the market is being stupid. It's not that people are missing something. You're wrong. What decision do you make?" "Jennifer Urig, who was such a good friend and mentor, she had this phrase that I always think about. Ultimately, as an investor, you either have to panic early or double down late. Essentially no one does both. And know thyself." "I am not a panic early person. I am a double down late person. And I think knowing that is something that kind of helps me go through a drawdown or a tough period of performance."

English
1
0
28
7.1K