Andrei Stan

459 posts

Andrei Stan

Andrei Stan

@andreiofstan

Katılım Ocak 2016
933 Takip Edilen44 Takipçiler
Andrei Stan
Andrei Stan@andreiofstan·
@KairosPraxis CXL will not work for inference for at least 3 years, 64GB/s theoretical maximum bandwidth is dog shit. The complexity and cost is way worse than ssds and flash
English
0
0
1
39
Kairos
Kairos@KairosPraxis·
Kimi's paper went over my head so I asked my research assistant to connect the dots and find stonks that'll benefit from these trends. The answer will totally NOT surprise you: 1. $ALAB (and perhaps $PENG) - Kimi moves prompt history out of scarce GPU memory into larger CPU memory. Astera disclosed a 2027 design win for this using a customized Leo CXL controller 2. Phison ($8299.TWO) - Kimi uses NVMe as an overflow memory. Phison turns SSDs into managed AI memory, extending available memory to as much as 8 TB. 3. Interconnects ($SMTC/$CRDO) : Larger models require faster data transfer within a rack and across racks. We will see denser racks and more demand for AECs, ACCs, retimers, redrivers, etc. 4. $FROG — Kimi created 51 million sandboxes from 1.5 million software images. JFrog sells p2p distribution to launch thousands of environments quickly. 5. $RMBS — K3’s external cache pool increase the bandwidth of ordinary server memory. Rambus supplies the control and buffer chips installed on high-bandwidth server-memory modules. 6. $NET — Virtual computers and sandboxes. 7. $LITE and OCS - K3-like inference creates large, "predictable" cluster migrations. OCS lets you connect two clusters via optics.
Kairos tweet media
English
6
11
74
9.7K
Evan
Evan@StockMKTNewz·
Meta Platforms $META and BlackRock $BLK will establish a joint venture to build and operate a 1-gigawatt data center complex in Texas Total development costs for the project spanning buildings, power, cooling and connectivity will come to ~$14B The campus will go online in 2028 BlackRock funds will hold an 80% interest in the project, while Meta retains the remaining 20% - Bloomberg
Evan tweet media
English
32
27
402
55.5K
Owen Brake
Owen Brake@OwenBrakes·
Datacenter interconnects are moving faster and faster. 224Gbps-PAM4 interconnects switch at 56 GHz, across a 1m copper wire yielding 36 dB attenuation! (4,000x drop in power). Your eye diagrams are useless here, you have to reconstruct signals from the noise.
Owen Brake tweet mediaOwen Brake tweet mediaOwen Brake tweet media
English
10
12
261
20.3K
Mr. Rc
Mr. Rc@rcx86·
Grok 4.5 is too fast for how good it is!
English
1
0
10
759
SemiAnalysis
SemiAnalysis@SemiAnalysis_·
With the upcoming TPUv8i BroadFly topology, it can scale up to 1,024 chips within a 7-hop radius, compared with the traditional 3D torus 16-hop radius. This enables lower latency, making it easier to overlap communication with compute without incurring the cost of a dedicated switch. 3/4🧵
SemiAnalysis tweet media
English
3
3
92
21K
SemiAnalysis
SemiAnalysis@SemiAnalysis_·
Successfully training models on TPUs has been demonstrated by Anthropic through the past five-plus successful Claude releases. This is positive for the ML community, as Google TPUs continue to gain market share outside of internal Google workloads, giving frontier AI labs a viable alternative for training. 1/4🧵
SemiAnalysis tweet media
English
20
44
520
114.2K
Andrei Stan
Andrei Stan@andreiofstan·
@zephyr_z9 @vikramskr What do you see the prime usage for CXL being in case of Google? Yes HBM is hard to get but is DRAM that much easier? Especially fast DRAM? You still need the memory for CXL.
English
0
0
0
195
Zephyr
Zephyr@zephyr_z9·
@vikramskr No CXL is for something else
English
2
0
12
1.9K
Yunfan Zhang
Yunfan Zhang@z4y5f3·
@teortaxesTex These are almost definitely Huawei Ascend 910C. 910C superpod has 48 machines with 8 processors each. Each processor has two physical dies that could function as two logical processors. In that mode, each die has 64GB HBM and 200 Gbps RDMA networking so everything matches.
Yunfan Zhang tweet mediaYunfan Zhang tweet media
English
5
7
106
31.1K
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
Damn, that's right. So: "≈48B" (thanks to N-Gram embedding, variable) active, 35T tokens. V4-tier, ≈8e24? Would be the biggest Chinese pretraining on domestic hardware. Some strange "Superpods": > "our accelerators" > "up to 48 machines each" > below 80 GB HBM per accelerator > "The device offers limited HBM bandwidth but a relatively large L2 cache" > "built-in 200 Gbps network adapter within the accelerator" what is this stuff Was the first truly big Chinese model trained on domestic compute also using some extremely obscure piece of hardware?
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet mediaTeortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet mediaTeortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet media
tphuang@tphuang

@teortaxesTex @zugzwangg_ it said 35T+ tokens

English
12
14
225
77.6K
Andrei Stan
Andrei Stan@andreiofstan·
@degentradingLSD You cannot memory pool without memory. DRAM is still the biggest bottleneck
English
0
0
3
1.4K
degentrading
degentrading@degentradingLSD·
I have said it in my notes for the last week. But i will say again. Memory pooling and compute will lead the next rally. Memory pooling - $ALAB, $CRDO, $PENG, $MRVL, $CRDO was hit hard by the rebalancing flow last friday - i think these levels offer good value.
degentrading@degentradingLSD

Pre Mkt Thoughts - 29 Jun 26 Asia markets had some early jitters before coming back to almost unchanged. KOSPI, NKY both traded with very strong correlation, selling off early in the morning before inching back into the close. Lets talk about the memory names. $285A had a dip to 82.5k levels before ending about -4% for the day. This represents a high r/r level for longs. $285A also has triggered some of the prior patterns i noticed in its price action. Meanwhile on the other side, news broke on $SAMSUNG and $SKYNIX planning to build 2 new massive chip fab sites in South Korea's southwest region as part of a "national project". Let's rerun back to Leopold's Situational Awareness article - the manhattan project. The early innings are coming through. Countries are treating this as existential. The sovereign buildout thesis is manifesting in real time. Aside, direct implications for Korea is that the rally in KOSPI should increasingly broaden out. I am keeping my eye out for Korea small caps. I believe that for small accounts, these can turn out to be some of the best opportunities. Personally, i am afraid of the illiquidity in those tickers and will probably sit it out. Coming into the US session, i indicated in my Sunday note that for end of quarter rebalancing, there is about 165b of equities to be sold into the close. Some have asked, how i would allocate into such a situation. There is no right answer. It hinges on 2 key pivots. 1 - What is your risk appetite? 2 - How much do you think flows have been front run? For those who are aggressive, that could mean allocating into one clip right into the close on tuesday. For others, that could mean a 2 day twap into the markets. I have been banging the drum in the last few notes that i see the best technical set up into July for longs. Let me take the time here to list it out again. First of all, the fundamental picture looks extremely strong. We had 2 articles dropping over the week. First was news that $AAPL was lobbying for the white house to allow them to buy from $CXMT. While this might sound bearish for memory. Please remember that the memory market is global. Paraphrasing Semianalysis Dylan - "It literally doesnt matter? CXMT cannot satisfy chinese demand this decade". Next, let me emphasize that $CXMT is selling at ~5% discount. After you account for quality differences, there is almost no difference. So why is $AAPL doing this? It's not about price, its that the supply is simply not there. If anything, this is a direct affirmation that indeed...supply is simply not there. Next, an article that $GOOG is capping $META's use of Gemini AI. How insane is this? $META cannot get enough compute that it is sourcing it from $GOOG. Crazier is that... $GOOG simply cannot meet the demand. Look, the hyperscalers will not stop capex any time soon, even as their equity prices take a hit. Because they see the revenues on the horizon. CAPEX hits the bottomline on t0, revenues and backlog hit the bottomline in the future. As these revenues arrive, the hyperscalers will start making money hand over fist. Remember...Anthropic's gross margins on inference is ~70%. By building now, they are securing a moat for the future. Where does this lead to? Again, the 2 themes i keep rehashing is compute (neos) and memory pooling. While i think memory will continue to be strong, the easy part of the rally and the valuation gap has been covered. On compute, $NBIS stands out as the leader of the pack. $SHAZ is an interesting up and coming player backed by Situational Awareness. i am expecting a 13G to hit the timeline tonight. On memory pooling, tickers include $ALAB, $CRDO, $PENG, $MRVL. On $CRDO, friday's sell off was very much influenced by the Russell rebalancing (as per my sunday note), i think the 240s area offer exceptional value for risk reward. In my notes last week, i talked about $BABA - I strongly think that now that even Burry has sold $BABA - sentiment cannot get worse than now. These levels also offer good risk reward. Sep calls especially are cheap. Good luck!

English
4
10
228
82.2K
Jukan
Jukan@jukan05·
Yes, MU is not Nvidia. But going forward, it may become even more important than Nvidia. Think about it. Inference is now directly tied to money. But inference does not get better simply by adding more Nvidia GPUs. In fact, GPUs are often underutilized in inference, sitting idle due to memory bottlenecks. For inference, adding more memory is far more valuable. Ultimately, the ROI of inference depends less on GPUs and more on memory. So why are people still looking at Micron through Nvidia’s framework? Think bigger. Inference is memory.
emini tic@TicTocTick

MU is not NVDA!! MU goon crash to 700 soon (now 1200). Remember we had this at 80!!! RAM is NOT GPU fools !!!

English
132
240
2.7K
819.3K
Andrei Stan
Andrei Stan@andreiofstan·
Expected it should appear, depends how they implement it but it may be very cool if the NIC has direct access to the whole NVLink domain. GPUDirect could work between any NIC and any GPU. This would also allow posting RDMA ops from any GPU/CPU to any NIC in the NVLink domain
Vengineerの妄想@Vengineer

ConnectX-10 NVLink-C2C なるものが登場しています。 Vera の PCIe は、Gen6 なので、1.6Tbps と想定する ConnectX-10 は Gen7 。 となると、Vera + ConnectX-10 構成の BlueField-5 では、1.6Tbps を流せません。そこで、ConnectX-10側にNVLink-C2Cを追加することで1.6Tbpsを実現するの?

English
0
0
0
12
will whang🌻
will whang🌻@will_whang·
But it is interesting to see another use of MIPI CSI for things other than image data. Though I'm not ... exactly sure how much DSP work can RPI5 handle
will whang🌻 tweet media
English
2
0
7
633
Qualcomm
Qualcomm@Qualcomm·
Hardware plus software defines leadership in the AI era. Today we announced an agreement to acquire @Modular, advancing our evolution as a developer-first AI solutions company delivering generative and agentic AI from edge to cloud. bit.ly/44rBpe5
Qualcomm tweet media
English
16
45
250
179.5K
Andrei Stan
Andrei Stan@andreiofstan·
@jukan05 Marginally better than saying you are running 8x7b lmao
English
0
0
0
367
Jukan
Jukan@jukan05·
Is it just me, or does this feel bearish for Cerebras?
Jukan tweet media
English
59
25
652
116.3K
Andrei Stan
Andrei Stan@andreiofstan·
Nobody cares about TOP500 supercomputers. They are not even supercomputers anymore, it's just a cluster. Probably No. 1 submitted is not even TOP 50 if all clusters were considered
English
0
0
0
12
Andrei Stan
Andrei Stan@andreiofstan·
@PolymarketSport This was the easiest bet ever, fucking missed it, literally free money
English
2
0
2
2.7K
Polymarket Sports
Polymarket Sports@PolymarketSport·
🚨JUST IN: A trader put $2.7k on Ronaldo to cry during the World Cup This pays out $3,800.00 on Polymarket
Polymarket Sports tweet mediaPolymarket Sports tweet media
English
566
1.3K
29K
2.6M