
Wafer Street
148 posts

Wafer Street
@OdiseaInvest
AI Infrastructure research Semis | memory | optical




$AMD and $INTC are pursuing long-term server CPU agreements with major Chinese customers as supply tightens and prices surge. Some CPU prices in China have risen more than 40% this year while lead times for certain Intel server chips have stretched to six months.

$GOOGL says it will “expand the use of third party capacity in Q3” as a bridge while it builds more internal capacity. After years of surging hyperscaler CaPex, demand still exceeds supply showing shortage is severe enough that even Google will accept lower margins for faster access to compute. $CRWV, $IREN, $WULF, $HUT, $CIFR, $NBIS, $APLD












SK Group Chairman Chey Tae-won told reporters at the Korea Chamber of Commerce and Industry’s Jeju Forum on the 15th, “Demand for AI semiconductors is expected to increase by at least 60–100% next year compared with this year.” He added, “Even for the overall memory semiconductor market, demand should be expected to grow by at least 50–60%.” Elaborating, Chey said, “Virtually no company will see any meaningful increase in supply next year, so the supply-demand gap is bound to widen significantly.” He continued, “Companies around the world are clamoring for memory chips—to the point where the situation could be described as utter chaos.” In other words, currently planned capacity expansions will be insufficient to meet the surge in demand.

$MU $SKHY $NVDA Kimi K3 is "cheap" to run. Bullshit. It costs almost as much as ChatGPT 5.6 Sol. (Total Tokens Used) × (Price Per Token) = Final Bill To complete the exact same evaluation, K3 used roughly 1.9x more output tokens than Sol and 1.5x more than Fable. Why? K3 is a heavy reasoning model. It burns a massive amount of internal Chain of Thought tokens to think out loud, debug its own code, and iterate on complex tasks. It is incredibly verbose and takes the long, windy road to the answer. Every one of those tokens is a separate trip to HBM to read the weights and the cache. K3 thought roughly twice as long as Sol to reach a comparable score. Twice the decode steps. Twice the memory bandwidth cycles. Twice the HBM hours, just for Chain of Thought. End result? 22% slower than GPT 5.6 Sol on Time per Intelligence Index Task. And the final bill: when you multiply (Total Tokens Used) × (Price Per Token), the math balances out. K3's massive token "bloat" completely swallowed up its massive discount on paper. Instead of being 70% cheaper to run, it ended up costing roughly the same as GPT-5.6 Sol ($0.94 vs $1.04 per task) I've said this before and I'll say it again. Software optimization is going application specific. No different from ASIC chips, which handle one type of task better than general purpose silicon. Same with models. Optimize for one thing and you pay somewhere else. K3 optimized memory per token and paid for it in token volume. Per completed task, that means potentially more pressure on HBM depending on the complexity of the problem and more time to get there. So just another frontier model. Another HBM customer. $MU $SKHY $DRAM



Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!


Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.










