Sabitlenmiş Tweet
C. Taewan
148 posts

C. Taewan
@taewan_02
Robotics & Local LLMs 🤖🧠 AI Startup Founder | Master’s Student Cooking NVFP4 on 6x RTX PRO 6000 Blackwell 🔥
Republic of Korea Katılım Nisan 2024
73 Takip Edilen99 Takipçiler

DIY server rack.
My plan is.
3. Edge devices -> AI Research
2. RTX 5090 2x -> AI Research
1. RTX PRO 6000 8x -> LLM SERVE




C. Taewan@taewan_02
My new employeeee
English


@taewan_02 Whats better... one super employee or a half dozen narrowly focused less powerful process executors?
English
C. Taewan retweetledi

@giles_home It seems fine so far since I installed a lot of case fans.
English
C. Taewan retweetledi

GLM-5.2 Nvidia NVFP4
No pruning, no quantising, exact same model card.
110 tok/s single stream.
My friend is a genius, holy fudge.
Here's the repo, he got full DS4-Flash on a 5090 + DDR5 github.com/kacper-daftcod… at 38.5 tok/s
This is a breakthrough.
English
C. Taewan retweetledi
C. Taewan retweetledi

Hy3-NVFP4 running on 2 RTX PRO 6000 Blackwells
Unfortunately, could only get 90K context right now.
Details:
> LibertAIDAI/Hy3-NVFP4
> MoE model: 295B total params, 21B active params
> Native context: 262K
> Served context: 90,112 (memory constrained)
> Precision: ModelOpt NVFP4
> KV: FP8
> Runtime: vLLM Hy3 CUDA 13 image w/ TP=2
25 tok/s single-stream decode
96.3 tok/s C8 decode
~2.3K tok/s near-limit prefill at 78K tokens
HF in comments

English

🔥🙏 Thanks to @tare_azure for access to a server with 8x RTX PRO 6000 Server Edition GPUs!
Daily X research posts have been great. Now maxing this to supercharge model dev! 🚀💪 Open source & open weights as always.
これからもよろしくおねがいします!

日本語

While training a draft to apply JetSpec to 27B, we found that a chain approach is more suitable for this model than a tree approach. Decoding speed = number of accepted tokens ÷ time taken per check. Applying a tree approach to GatedDeltaNet did not contribute to speed improvement because the verification cost increases with each additional branch.
Although our hypothesis was incorrect, this is a significant step forward.
For 27B, we will apply DSpark, published by DeepSeek. DSpark has a mechanism to prevent the accuracy of predicted tokens from deteriorating in the later stages. This seems advantageous for chain-based speculative decoding. We will continue our research.

English
C. Taewan retweetledi

DwarfStar(ds4)로 DeepSeek V4 Flash/Pro를 DGX Spark, Mac Studio, Mac mini에서 돌려봤습니다.
Flash q2는 Mac Studio에서 빠르고, 64GB Mac mini에서도 SSD streaming으로 실행은 됩니다. 하지만 RSS/커뮤니티 제목 번역 품질은 Gemma4 31B Dense MTP 쪽이 더 안정적이었습니다.
seapy.com/ds4-deepseek-v…
한국어









