C. Taewan

148 posts

C. Taewan

C. Taewan

@taewan_02

Robotics & Local LLMs 🤖🧠 AI Startup Founder | Master’s Student Cooking NVFP4 on 6x RTX PRO 6000 Blackwell 🔥

Republic of Korea Katılım Nisan 2024
73 Takip Edilen99 Takipçiler
Sabitlenmiş Tweet
C. Taewan
C. Taewan@taewan_02·
My new employeeee
C. Taewan tweet media
English
22
3
69
11.6K
C. Taewan
C. Taewan@taewan_02·
My new employeeee
C. Taewan tweet media
English
22
3
69
11.6K
C. Taewan
C. Taewan@taewan_02·
8 RTX PRO 6000 GPUs = the productivity of 10+ people. 🚀
English
1
0
8
336
Nick Masi
Nick Masi@Launch1Labs·
@taewan_02 Whats better... one super employee or a half dozen narrowly focused less powerful process executors?
English
1
0
1
327
C. Taewan
C. Taewan@taewan_02·
Super Insane!! at 4x RTX PRO 6000 madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
C. Taewan tweet mediaC. Taewan tweet media
Polski
5
1
17
2.4K
C. Taewan retweetledi
0xSero
0xSero@0xSero·
Nvidia/GLM-5.2-NVFP4 running on 4x rtx pro 6000 with 400k context @ 82 tok/s and is doing amazing I have to water cool these GPUs, the temperatures and sounds are limiting how far I can push this thing This new moet compression is going straight in local studio for onboarding
English
24
12
251
19K
C. Taewan
C. Taewan@taewan_02·
@giles_home It seems fine so far since I installed a lot of case fans.
English
1
0
2
86
C. Taewan retweetledi
0xSero
0xSero@0xSero·
GLM-5.2 Nvidia NVFP4 No pruning, no quantising, exact same model card. 110 tok/s single stream. My friend is a genius, holy fudge. Here's the repo, he got full DS4-Flash on a 5090 + DDR5 github.com/kacper-daftcod… at 38.5 tok/s This is a breakthrough.
English
101
101
1.3K
156.2K
C. Taewan retweetledi
0xSero
0xSero@0xSero·
GLM-5.2 running on 4x RTX Pro 6000s Full model, unpruned. - Top 32 experts for coding and agentic at nvfp4 - Remaining 196 experts at FP3 autoround tuned - Gates MXFP8 - FP16 for dense layers - 400K context FP8 Took this from the RTXPRO6000 discord and applied REAP scoring
0xSero tweet media
English
19
12
214
13.2K
C. Taewan retweetledi
am.will
am.will@LLMJunky·
Hy3-NVFP4 running on 2 RTX PRO 6000 Blackwells Unfortunately, could only get 90K context right now. Details: > LibertAIDAI/Hy3-NVFP4 > MoE model: 295B total params, 21B active params > Native context: 262K > Served context: 90,112 (memory constrained) > Precision: ModelOpt NVFP4 > KV: FP8 > Runtime: vLLM Hy3 CUDA 13 image w/ TP=2 25 tok/s single-stream decode 96.3 tok/s C8 decode ~2.3K tok/s near-limit prefill at 78K tokens HF in comments
am.will tweet media
English
7
4
24
2.2K
Hikari∣LocalLLM⚡
Hikari∣LocalLLM⚡@Hikari_07_jp·
🔥🙏 Thanks to @tare_azure for access to a server with 8x RTX PRO 6000 Server Edition GPUs! Daily X research posts have been great. Now maxing this to supercharge model dev! 🚀💪 Open source & open weights as always. これからもよろしくおねがいします!
Hikari∣LocalLLM⚡ tweet media
日本語
6
3
65
2.3K
Hikari∣LocalLLM⚡
Hikari∣LocalLLM⚡@Hikari_07_jp·
While training a draft to apply JetSpec to 27B, we found that a chain approach is more suitable for this model than a tree approach. Decoding speed = number of accepted tokens ÷ time taken per check. Applying a tree approach to GatedDeltaNet did not contribute to speed improvement because the verification cost increases with each additional branch. Although our hypothesis was incorrect, this is a significant step forward. For 27B, we will apply DSpark, published by DeepSeek. DSpark has a mechanism to prevent the accuracy of predicted tokens from deteriorating in the later stages. This seems advantageous for chain-based speculative decoding. We will continue our research.
Hikari∣LocalLLM⚡ tweet media
English
5
2
57
3.8K
C. Taewan retweetledi
Seapy
Seapy@seapy·
DwarfStar(ds4)로 DeepSeek V4 Flash/Pro를 DGX Spark, Mac Studio, Mac mini에서 돌려봤습니다. Flash q2는 Mac Studio에서 빠르고, 64GB Mac mini에서도 SSD streaming으로 실행은 됩니다. 하지만 RSS/커뮤니티 제목 번역 품질은 Gemma4 31B Dense MTP 쪽이 더 안정적이었습니다. seapy.com/ds4-deepseek-v…
한국어
3
5
18
1.3K
Hikari∣LocalLLM⚡
Hikari∣LocalLLM⚡@Hikari_07_jp·
I often receive DMs asking about model selection and configuration—please feel free to reach out with your questions! I am particularly confident in my expertise regarding the Blackwell architecture and inference engines👍
English
2
1
14
875