Ryan Wang 🇹🇼

6.9K posts

Ryan Wang 🇹🇼 banner
Ryan Wang 🇹🇼

Ryan Wang 🇹🇼

@ryanwang

An individual Investor from Taiwan who focus on Tech, AI, autonomy, and robotics. The future is autonomous. Ad astra per aspera!

Taipei, Taiwan Katılım Mayıs 2018
1.9K Takip Edilen1.9K Takipçiler
Ryan Wang 🇹🇼
Ryan Wang 🇹🇼@ryanwang·
Super cool
Sandro van Kuijck@Everydaysandro

Driving My Tesla from the Arctic Ocean to Tierra de Fuego in the Dead of Winter… I am proud to announce that on July 16th, I became the first SOLO person in history to drive the FULL length of the Americas with an Electric Vehicle! Total Duration: 747 Days Countries: 15 Longest Wait Time to Charge: 3 days (Peru) Fastest Charge: 35 Minutes (Chile) Distance driven - 46,961km (29,350mi) Windshield broken: 0 Flat tires: 0 (OMG) - No Spare Ran out of Range: 1 (Chile - Atacama Desert) Total energy consumed: 8,444 kWh Ave Energy Consumption: 179.8 Wh/km (286Wh/mi) Bribes Paid: 1 - $100 USD Robbed: 0 Fines: 0 Breakdowns: 0 Dogs Adopted Me: 100+ Funny Police Searches : 2 The Pan-American Highway is the world’s longest highway, almost spanning the full length of North and South America. Before today, no one had yet driven from the tippy top of North America (Tuktoyaktuk) into the Darien Gap to the southern tip in South America (Ushuaia) in an electric vehicle solo - unsupported. My Model X, Beluga, was converted into the ultimate stealth overlanding rig, with a full kitchen, running water, Starlink and 300 watts of solar panels on the hood, which allowed us to roam across some of the best most remote places on Earth with confidence. The real challenges here…border crossings, cartels, and corrupt cops, yet I embraced it all. Charging on the other hand, was quite a puzzle sometimes, being a foreigner and having very little information to go on. I discovered how far along the electric movement is and where the short comings lie…and I can say with confidence that Tesla Supercharging is the best and most reliable infrastructure out there… I wanted to showcase that anyone with a dream can make it happen, without needing all the answers. Thank you to all the kind strangers that helped along the way, and the countless dogs that adopted us and kept guard at night! youtu.be/PwVQiI5j250?is… The irony..Ushuaia currently doesn’t have any power…so this message is coming from Space, Powered by my Custom Solar Panels! @Tesla

English
0
0
1
152
Ryan Wang 🇹🇼
Ryan Wang 🇹🇼@ryanwang·
This video features an interview with K.S. Pua (Chairman Pan), founder and chairman of Phison Electronics, focusing on the memory "Super Cycle," AI storage demands, and his management philosophy. youtu.be/ZACsUuAkCpI?si… Accurate Super-Wave Prediction: Chairman Pan revealed he predicted the memory super cycle back in October 2023. At the time, the English report triggered panic among U.S. clients, and he was even questioned face-to-face by a global chip giant's executive in early 2024. However, market prices skyrocketed by November, proving his foresight. Uncatchable AI Memory Gap: Addressing concerns over "oversupply" and an "AI bubble," he countered with three dimensions: cloud AI users will jump from 1 billion to several billion, applications will shift from text (KB) to video (GB) exploding data demand by ten-thousandfold, and daily usage frequency per person will surge. With massive infrastructure costs, price hikes are inevitable, and capacity simply cannot keep up with demand. Criticizing Retail Investors' Herd Mentality: Facing stock price volatility and retail panic, Pan bluntly called the market a bunch of fools who "find ten thousand excuses just to buy or sell." He emphasized that top executives are the largest shareholders and have absolute confidence in their technology. With a market value of 0.5 trillion TWD, he won't blindly pursue a mainboard listing (to join ETF 0050) and values substantive impact more. Unique Management & Conclusion: Phison actively recruits esports talents, valuing their high stress tolerance, focus, and sensitivity to hardware limits, which are ideal for AI software-hardware optimization. He concluded that while AI startups may bubble, the AI trend itself will not; capital is simply rotating into different pockets.
YouTube video
YouTube
English
0
0
3
369
Ryan Wang 🇹🇼 retweetledi
Dirty Tesla
Dirty Tesla@DirtyTesLa·
I got a behind the scenes tour of The Boring Company! The project is far more impressive than anticipated, and I did my best to share that with this in-depth video. You can tell how passionate the entire team is. Huge thank you to @boringcompany!
English
75
125
1.4K
153.8K
Ryan Wang 🇹🇼
Ryan Wang 🇹🇼@ryanwang·
so cool
Joe Tegtmeyer 🚀 🤠🛸😎@JoeTegtmeyer

On 16 July 2026, @Tesla hosted the 1st “Lithium Day” event at the Robstown, Texas, Lithium Refinery. This was on the same day as @SpaceX Starship Flight 13 (scrubbed) and although I was at Starbase, @SERobinsonJr (aka Elon Chron) was present & together we put together this video combining some drone video clips & images of the event. The facility (part 1) is now complete & operational, taking Spodumene Concentrate 6 as inputs & LiOH as the output. This material is then shipped to Giga Texas for battery cathode & ultimately 4680 battery production! The operational status announcement was formally made by Plant Manager Jason Bevan & Infrastructure VP Rodney Westmoreland Jr @RodneyWest336 to a crowd of local business owners, employees, special guests & state representative Denise Villalobos. Battery VP @BonneEggleston was also on hand & spoke about newly filed patent update for the 100% full recycling process for batteries, & possible future expansions. This includes a Global Patent WO2026048253A1 which describes a new process for recycled batteries & manufacturing scrap! Overall, a great event to hear first-hand about the plant, the progress the team has made to get the worlds most advanced & only Lithium Refinery in North America to operational status! @Klarissa24g

English
0
0
1
297
Ryan Wang 🇹🇼 retweetledi
Elon Musk
Elon Musk@elonmusk·
@wholemars The car will start to remember your specific interventions and match each person’s individual preferences
English
896
791
11.3K
1.3M
Ryan Wang 🇹🇼 retweetledi
SemiAnalysis
SemiAnalysis@SemiAnalysis_·
MASSIVE DELAY ALERT TO ORACLE’S STARGATE SITE AND BLOOM ENERGY🚨🚨 Oracle’s Project Jupiter behind-the-meter datacenter project in New Mexico that plans to use Bloom Energy is at risk of a 1-2 year delay due to permitting and pipeline building blockers. (1/8)🧵
SemiAnalysis tweet media
English
92
167
1.5K
712.7K
Ryan Wang 🇹🇼 retweetledi
Elon Musk
Elon Musk@elonmusk·
@minchoi Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).
English
685
793
11.7K
1.6M
Curry_TW
Curry_TW@KotlinerBTC·
SpaceX 因為少數發動機沒有成功點火(啟動),緊急終止發射程序。Elon Musk認為幾天內可以替換引擎,並且在下周初進行再次發射。 趕著立刻再上架,固然可以挽救一些投資人的信心,也顯示自己能夠保住時間表的態度與技術力。 現在把這件事情補做做好,只能爭取少扣分。可是一旦補做了還沒做好,那就是方方面面的災難級打擊了。 目前的局面,盈虧比非常糟。頂多就是接近打平,但搞砸了會大虧。與其如此,不如再晚個一週,讓相關技術人員有更充裕的時間應對。
Elon Musk@elonmusk

Some of the engines didn’t start, triggering an automatic launch abort. Now offloading propellant. Next launch attempt hopefully in a few days.

中文
30
0
17
8.3K
Tsla Chan
Tsla Chan@Tslachan·
$TSLA [FSD V14 Lite] 415km, 의정부에서 부산까지 초장거리 무개입 주행! 테슬라 V14 Lite로 415km를 개입 없이 달렸습니다. 주간 고속도로에서 Lite는 확실히 야간보다 더 자신감있고 빠르게 주행하더군요. 출발할때 무개입은 불가하겠지라고 생각했지만, Lite가 해냈습니다. 속도를 더 내고 싶을땐 악셀 페달 밟아주고, 그 외에는 전방주시하며 감독을 하니 운전자의 피로도는 극도로 줄어듬을 느꼈습니다. Lite 버전임을 감안하면 정말 인상적인 수준입니다. • 2022 Model 3 라이젠 / 신속주행 • 경기 의정부-부산 해운대 슈퍼차저 • 52,351km->52,766km (415km 무개입 주행) • 주행전비: 116.1Wh/kw • 주행시간: 5시간 43분 동영상 링크⤵️ (영상 용량이 89GB라서 𝕏에는 업로드 불가): youtube.com/watch?v=iusqs7…
YouTube video
YouTube
Tsla Chan tweet media
한국어
26
61
297
15.4K
Ryan Wang 🇹🇼 retweetledi
宝玉
宝玉@dotey·
杨植麟在 GTC 2026 的一次演讲"How We Scaled Kimi K2.5":月之暗面想把 AI 训练中三个沿用了近十年的基础组件,重新做一遍。优化器 Adam(2014 年)、注意力机制(2017 年)、残差连接(2015 年),三个 Transformer 时代的地基组件,月之暗面各给了一个替代方案,而且全部开源。 但这场演讲讲的不只是 K2.5,而是月之暗面过去一年多的技术路线:开源模型还能从哪里继续变强,又该怎样逼近闭源模型的前沿水平。 杨植麟把答案拆成了三个方向:让每个 Token 更值钱、让更长的上下文真正发挥作用,以及让多个 Agent 同时协作。 此外,他还分享了两个重要进展:视觉训练如何反过来增强文本能力,以及月之暗面刚刚公布的下一代架构 Attention Residue。 【一、数据不够,就让每个 Token 更值钱】 大模型训练正在遇到一个越来越现实的问题:高质量数据快不够用了。 互联网上真正有价值的文本数量有限,模型越做越大,需要的数据越来越多,大家迟早都会撞上“数据墙”。既然高质量数据很难翻倍,月之暗面的思路是:能不能让模型从同样的数据里学到更多? 他们给出的答案是 MuonClip,用它替代已经用了十多年的 Adam 优化器。 MuonClip 基于 Muon 优化器。简单理解,它会在更新模型参数时,尽量让不同方向上的信息保持独立,减少重复和浪费,从而提高训练数据的利用效率。 月之暗面的实验结果是:使用同样数量的数据,MuonClip 的训练效果接近把数据量增加一倍,这不仅意味着训练成本下降,也意味着模型能力的上限可能被推高。 假设手里只有 50 万亿个高质量 Token,如果 Token 利用效率提高一倍,就相当于又多出了 50 万亿 Token。在优质数据越来越稀缺的情况下,这种效率提升比单纯增加算力更重要。 但 Muon 也有一个明显问题:模型扩展到万亿参数后,注意力层里的数值容易失控。 月之暗面在训练中发现,注意力层的最大 logit 会突然飙升到 1000 以上,而正常范围通常只有 50 到 100。一旦数值继续膨胀,训练曲线就会发散,整个训练过程可能直接崩掉。 为了解决这个问题,他们设计了 QK-Clip。它会在模型前向计算时,实时检查每个注意力头的最大 logit。一旦数值超过安全范围,就同步缩放 Q 和 K 的投影,把数值压回来。 这个操作不会改变训练的收敛效果,只负责维持数值稳定。 靠着 QK-Clip,月之暗面成功把 Muon 扩展到了万亿参数规模,并训练了超过 15 万亿 Token,整个过程中没有出现一次 loss spike。 【二、上下文不只是要长,还要真正有用】 第二个方向,是提高模型利用长上下文的能力。 杨植麟展示了 Scaling Law 论文作者 Jared Kaplan 等人曾经做过的一组实验:比较 Transformer 和 LSTM 在不同上下文长度下的预测效果。 LSTM 读到一定长度之后,效果很快就不再提升。继续增加上下文,模型也很难从中获得更多信息。Transformer 则不同。上下文越长,模型对后续内容的预测通常越准确,而且很难看到明显的饱和点。 这个特点在 Agent 时代尤其重要。复杂 Agent 任务可能持续几天甚至几周。模型需要记住之前做过什么、得到过哪些结果、哪些方向已经失败,以及接下来该做什么。如果长上下文不能持续提供有效信息,Agent 的任务链条就很容易断掉。 问题是,标准的全注意力机制成本太高。它的计算量会随着上下文长度平方增长。上下文扩大十倍,注意力计算量可能增加一百倍。到了百万 Token 级别,训练和推理成本都会变得非常高。 月之暗面的解决方案是 Kimi Linear,核心是一种名为 KDA,也就是 Kimi Delta Attention 的线性注意力机制。KDA 的关键,是让模型学会“哪些信息要长期保留,哪些信息可以快速忘掉”。 传统线性注意力通常只有一个全局衰减系数,控制整个模型的记忆速度。这有点像所有内容共用同一个遗忘按钮:要么一起记住,要么一起忘掉。KDA 把一个衰减系数拆成了多个。不同信息通道可以使用不同的遗忘速度: - 有些通道衰减得很慢,负责保留长距离信息 - 有些通道衰减得很快,及时腾出空间吸收新内容 实际使用时,Kimi Linear 并没有完全抛弃全注意力,而是把线性注意力层和全注意力层按照 3∶1 的比例混合。 杨植麟称,这是第一个在短上下文、长输入和长输出任务上,都能全面超过全注意力的架构。上下文扩展到百万 Token,甚至更长时,它的效率优势会更加明显。 【三、从一个 Agent,变成一支 Agent 团队】 前两个方向都在提升单个模型的能力。第三个方向,则是让多个 Agent 一起工作。月之暗面把这种方式叫作 Agent Swarm,也就是“智能体集群”。 它的组织方式很像一家公司。一个主 Agent 充当 CEO,负责理解目标、拆解任务,并把不同子任务分配给多个子 Agent。子 Agent 可以分别扮演研究员、程序员、数据分析师和事实核查员等角色。任务完成后,主 Agent 再汇总结果。 这样做最大的价值是把串行任务改成并行任务。过去,一个 Agent 可能需要依次搜索资料、阅读文档、分析数据、编写代码、检查事实。现在,这些工作可以交给几十个甚至几百个 Agent 同时完成,从而大幅缩短复杂任务的执行时间。 不过,要让模型真正学会并行协作并不容易。训练 Agent Swarm 时,月之暗面设计了三种奖励: 第一种是“实例化奖励”,鼓励主 Agent 创建更多可以并行执行的子任务,避免它退化回单 Agent 串行工作。 第二种是“完成奖励”,要求子任务必须真正完成,防止模型为了获得实例化奖励,批量创建没有意义的空任务。 第三种是最终结果奖励,用来判断整个任务是否真正解决。 这三种奖励的权重会随着训练过程动态变化。训练前期更重视任务拆解、并行化和子任务完成率,后期则逐渐把重点转向最终结果。 从演示结果看,Agent Swarm 在复杂任务上能明显缩短执行时间。比如: - 同时下载、阅读几百个信息源并完成研究 - 并行撰写一份上百页文献综述的不同章节 - 同时分析十个不同的数据集 Kimi K2.5 发布时,Agent Swarm 已经支持最多 100 个 Agent 并行工作,整个任务最多可以执行 1500 个步骤。4 月发布的 K2.6 又把并行 Agent 的上限提高到了 300 个。 【四、一个意外收获:练“视觉”,也能让“大脑”变聪明】 K2.5 和前代 K2 之间,一个重要变化是采用了“早期融合”训练。 过去很多开源多模态模型采用的是“后期融合”:先用大量文本训练出一个语言模型,再用相对少量的视觉数据补上看图能力。 例如,先训练 20 万亿个文本 Token,再用大约 2 万亿个多模态 Token 把视觉能力“贴”上去。 K2.5 的做法不同。 它从训练一开始,就把视觉和文本数据混在一起。在 K2 文本基座的基础上,K2.5 又训练了约 15 万亿个混合 Token。 这种方式带来了两个让杨植麟感到兴奋的结果。 第一个发现是:只训练视觉任务,也能提升模型的文本推理能力。 研究团队只让模型完成数数、识别图片和视觉问答等任务,没有加入数学或编程训练。结果模型的文本推理能力也变强了。 换句话说,模型在练习“看”的同时,“想”的能力也得到了提升。 第二个发现则来自相反的方向:如果文本基座足够强,模型甚至不一定需要专门的视觉 SFT 数据。 K2.5 采用了“零视觉 SFT”方案。所有监督微调数据都是纯文本,然后再通过文本与视觉联合的强化学习,让模型获得视觉能力。最终,它在视觉任务上的表现依然接近最先进水平。 杨植麟认为,这种双向迁移来自早期融合。 当文本和视觉被放进同一个表征空间后,一种模态学到的能力,就有机会迁移到另一种模态。这也是 K2.5 能够“看图写代码”的基础。 如果视觉和文本仍然像两个彼此分开的“大脑”,这样的跨模态能力就很难自然出现。 【五、下一步:用注意力机制替代残差连接】 演讲快结束时,杨植麟介绍了月之暗面刚刚发布的一项新研究:Attention Residue,也就是“注意力残差”。这篇论文发布于 3 月 15 日,距离演讲只有两天。 残差连接是现代深度神经网络最重要的基础技术之一。2015 年,何恺明等人提出残差网络,此后残差连接逐渐成为 Transformer 的标准组件。它的基本做法是:每一层不仅处理上一层的结果,还保留一条直接传递信息的通道。这样即使模型很深,信息和梯度也能顺利通过,模型才有可能稳定训练。 杨植麟引用了 Ilya Sutskever 的一个说法:残差连接可以被理解为“旋转了 90 度的 LSTM”——LSTM 在时间维度上传递信息,残差连接则在网络深度上传递信息。 月之暗面顺着这个类比继续往前走了一步。 既然 Transformer 已经用注意力机制替代了 LSTM 在时间维度上的循环结构,那么在深度维度上,是否也能做同样的替换? Attention Residue 就是这个思路的产物。 标准残差连接主要使用上一层的输出。Attention Residue 则允许当前层查看所有前序层的输出,再通过注意力机制决定:哪些历史信息值得保留,哪些信息可以忽略。 也就是说,模型不再只能被动接收上一层的结果,而是可以主动从整个计算历史中挑选信息。 为了控制计算成本,月之暗面实际采用的是分块版本 Block Attention Residual。它会把模型层分成多个块,比如每 16 层组成一个块。块内继续使用标准残差连接,块与块之间才使用注意力残差。 实验显示,大约 8 个块就能获得大部分收益。 Attention Residue 带来了约 24% 的 Token 效率提升。 按照杨植麟的算法,如果有 50 万亿个高质量 Token,效率提高 24%,就相当于额外增加了 12 万亿个 Token。它在 GPQA、MATH 和 HumanEval 等推理、数学与编程测试上的提升尤其明显。 【三个用了近十年的组件,都还有改进空间】 演讲最后,杨植麟把月之暗面的三个替代方案放在了一起: - Adam → MuonClip - Full Attention → Kimi Linear - Residual Connection → Attention Residue 它们分别对应大模型训练中的三个基础问题: - 怎样从有限的数据中学到更多 - 怎样更高效地利用超长上下文 - 怎样让深层网络更灵活地传递信息 三个方案都可以相对独立地替换现有组件,也都已经开源。 这些技术能否直接叠加,增益能否简单相乘,目前还没有经过完整验证。但它们至少说明了一件事:很多被认为“已经足够好”的基础组件,可能远没有到达终点。 在优质数据越来越少、训练成本越来越高的情况下,继续堆参数和算力并不是唯一的路。重新设计优化器、注意力机制和残差连接,同样可能带来可观的提升。 这也是杨植麟整场演讲真正想表达的观点。 正如他最后所说: > “Open models cannot be just open; they have also to be great.” 开源模型不能只是开放,还必须足够强。 注:本视频由 BaoCut.app 翻译
Jackywine@Jackywine

KIMI CEO杨植麟,一个视频带你彻底看懂大规模 AI 系统训练的秘密 自己看一遍比什么都重要

中文
32
311
1.4K
273.2K
Ryan Wang 🇹🇼
Ryan Wang 🇹🇼@ryanwang·
I think @grok will be the first frontier model in US to open source! Just do it @elonmusk
Gavin Baker@GavinSBaker

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis. Rationale:   A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.    Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.   This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.   Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3. 
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead. Time will tell on both points. And likely fairly quickly. Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

English
1
0
1
264