Zephry .

100 posts

Zephry . banner
Zephry .

Zephry .

@hazzasgrl

✨ℓιfє ιѕ вєαυтιfυℓ✨

Katılım Ağustos 2012
86 Takip Edilen51 Takipçiler
Zephyr
Zephyr@zephyr_z9·
Production volume for 950DT is around 400k this year
prinz@deredleritt3r

Key quotes from the Liang Wenfeng investor presentation: 1. He's convinced that current AI models can't reach AGI and we need continual learning: "Just like CoT—after CoT reached its ceiling, it already surpassed the most top-tier humans, in doing math olympiad problems and writing programs... But it still stops there, that technology couldn’t reach AGI. So you see, AI’s intelligence trajectory is traceable. [A]fter Agent, we think the problem to be solved should be continuous learning—that is, how to let the model learn continuously, rather than you giving it strong training—it should be able to do relatively long-term continuous learning like a human." Observation: This is very similar to how Google views AI progress and next steps (except that Liang does not care for world models and believes that robotics should come at a much later step; more on that below). 2. He thinks that the right order in which to tackle these problems is: continual learning, then RSI (which is a "gradual process", NOT an intelligence explosion), then embodied AI (robotics): "After continuous learning, we might arrive at a singularity. This singularity is: when this model can learn continuously, it can already do all the things humans can do. It could then develop its own version, could research on its own, then develop its own next version, developing more advanced AI models. So it would reach a singularity, able to achieve its own iteration. But this singularity—it’s not really a singularity, it’s also a gradual process. This process might also be a relatively long gradual change, not a mutation... Then after this step is complete, I think that’s when embodied intelligence arrives." 3. He says that video models and world models have no relation to the intelligence ceiling, so DeepSeek isn't interested in them. 4. When he gave the speech (in May), DeepSeek had 20,000 H-equivalent GPUs: "We don’t have that many cards [GPUs]—the number of our cards is still relatively few. We currently have roughly 20,000 H-equivalent compute cards, and most of these just arrived, arrived in the last month or two, and there may be many more machines that haven’t arrived yet." Observation: OpenAI intends to run the automated AI research "intern" on the equivalent of 500,000 A-100 GPUs. This is something like ~8x the amount of compute that DeepSeek had available in May overall. 5. Export controls are working: "Our gap with the US is mainly in resources, and the gap in people isn’t very big... Talent isn’t the bottleneck—resources are the biggest bottleneck. Resources first affect talent cultivation, because with little compute, our opportunities to run experiments are relatively few, so our talent overall has a gap with the US. The talent gap is essentially also because of the compute gap. On the current largest models, we actually can’t afford to train. Even if we spent all 50 billion, we still couldn’t afford to train. Even if we could stack it up, we couldn’t afford to use it. If I wanted to train a model as large as [U.S. models], it should require 50,000 GB300s, or Huawei 950s, 200,000 cards. This is just training, not yet considering doing research. So the biggest gap between us and the US is in resources... This problem is currently basically unsolvable, because Huawei’s output is also limited. Because if I want to train 800B, I need 200,000 of Huawei’s newest cards, and this is just training, not yet considering doing research." 6. On Huawei and Chinese domestic chip production: "Huawei’s problem is still insufficient production capacity. Huawei gives us roughly 16,000 cards of capacity, internet giants maybe get a hundred-something thousand, we get ten-something thousand... this is probably just how much capacity Huawei has. So we also can’t count on training that bigger model on Huawei later, or training a model with several hundred B parameters active—there were some saying this year. But next year, the year after, there may be opportunities... ...the gap between [China] and the US on chips—I believe there won’t be a gap on ecosystem going forward, but on chips it’s fourfold plus two years... NVIDIA cards you can basically depreciate over five years. Huawei cards at most depreciate over three years. Huawei 950—this year using it is pretty good, next year using it I think is still okay, later using it I think might really be too power-hungry." 7. On the gap between China (or DeepSeek?) and the U.S.: "So the gap between us and the US might be lagging the US by 12 months, lagging the US maybe 12 to 18 months, or 6 to 12 months. Anyway, simply put, lagging the US by two years, then using only one-twentieth of the US’s compute to accomplish this thing. This narrative is lagging one to two years, but using only one-twentieth of its compute. So in the future we want to rewrite this narrative—that is, we use one-nth of its compute, but shorten this time even more, shorten it to 6 months, 3 months—I think this is a goal. And we can even surpass them in certain aspects. But under the situation where total compute still has an order-of-magnitude gap, comprehensive surpassing is unrealistic; but in certain key, trade-off-selected places, us surpassing in some areas might be possible."

English
7
5
112
33K
Zephyr
Zephyr@zephyr_z9·
Some interesting details from the leaked Liang Wenfeng investor call: - Inference margins are around 85% (sixfold profit) - They only had around 20k Hopper equivalents till May - $1B in API revenue is enough to turn the company cash flow positive (enough for R&D and training) - GPU payback period is 10 months, hardware depreciated over 3-5 years - Have secured 2 Atlas Super Clusters (16k 950DTs) from Huawei - Independent of Nvidia's CUDA/software ecosystem, they plan to port TileLang to Huawei I suspect that they haven't made all their inference optimizations public (cuz this is a crazy level of efficiency)
English
46
126
1.8K
184.9K
Zephyr
Zephyr@zephyr_z9·
Jensen's biggest nightmare
Austin Lyons@austinsemis

We are past the CUDA moat days. Anthropic says Claude can bootstrap @amd Instinct + ROCm. @NotTomBrown: "I think the thing that we were thinking about originally was whenever we're bringing up a new hardware platform, it's a big effort. It's like a huge thing. And so as we were thinking about this, we started doing our own evaluation of MI 355. You guys generously got us a rack to start working. And we expected this to be kind of a big process. Our actual experience was we had one engineer who start doing it. They spun up Claude, asked it, hey, bring up this machine, left it going over the weekend. And we ended up with a graph of the actual performance of our leading model on it, just going up and up and up over the weekend."

English
25
23
618
106.4K
Zephyr
Zephyr@zephyr_z9·
Nvidia went cableless for a reason Blackwells were a nightmare
Nick Dorsey@Midnight_Captl

I was able to take a picture of the $AMD Mi455 Helios compute tray Lisa had up on stage with her during the presentation - I’m far away so the picture was blurry, Sol 5.6 was able to up res it significantly @jaygoldberg pointed out to me yesterday that they do not have a compute tray on display on the showroom floor, they only have a digital version of it on a screen In this image, you can see the complexity of the system @AMD is trying to ship in volume. It is a MONSTER of a machine. And the complexity of it is enormous. I think scaling this is going to be a huge lift. And I think they can do it, but I don’t think it’s going to be a smooth ramp

English
14
12
380
88.1K
Zephyr
Zephyr@zephyr_z9·
"If my objective were to capture five percent of all human GDP through AI, the arithmetic might theoretically work. Look at OpenAI: its calculations may appear to work, and in theory there is no problem. But it faces one problem: it will be defeated by someone willing to take only one percent."
English
13
10
251
37.1K
Zephyr
Zephyr@zephyr_z9·
Assuming 1/64 sparsity and 150B active, DeepSeek's 10T parameter model will begin training between late 26 and April 2027 12-15 months after Anthropic
English
11
1
264
35.5K
Zephry .
Zephry .@hazzasgrl·
Just had the most electrifying night at the indie rock show—raw vocals, killer guitar riffs, and the crowd singing every word till midnight. Pure magic!
English
0
0
0
11
Zephry . retweetledi
Latc Beo
Latc Beo@prynk14·
New semester starts! Fresh ideas, new goals, and full energy ahead—let’s make this term count
Latc Beo tweet media
English
0
1
0
9
Zephry . retweetledi
Precise Market Insights.
Precise Market Insights.@aizatbaik·
Late-afternoon light spilling over city streets—sudden small magic I had to stop for.
Precise Market Insights. tweet media
English
0
1
0
12
Zephry .
Zephry .@hazzasgrl·
New semester, fresh start. Let’s chase small wins, learn boldly and make this term count! ,
Zephry . tweet media
English
0
2
1
10
Zephry . retweetledi
Brian Shannon, CMT.Assistant
Brian Shannon, CMT.Assistant@edinarapaulalov·
Caught my desk plant staring back at me today. Turns out it’s just as bored of this spreadsheeting grind as I am. #WorkModeOff
Brian Shannon, CMT.Assistant tweet media
English
0
2
0
5
Zephry .
Zephry .@hazzasgrl·
Just crushed today's to-do list: client call checked , report drafted , team sync wrapped . Friday energy hitting different!,
Zephry . tweet media
English
0
1
0
3