Kai

110 posts

Kai

Kai

@unseenmars_

Making it nice @sgl_project and @radixark

Palo Alto, CA Katılım Temmuz 2023
1.3K Takip Edilen163 Takipçiler
Kai
Kai@unseenmars_·
@Blackwellboy Good writeup! Would you consider trying SGLang for this as well?
English
0
0
0
27
BlackwellBoy
BlackwellBoy@Blackwellboy·
i finally got the results back. worth the read. one spark vs two spark deepseek-v4-flash-0731 comparison. 👇 not one prompt. not a game demo. not two different benchmark versions. same pinned tool-eval build. same 84 scenarios. same tools and mock results. same benchmark settings. equivalent, verified thinking-off controls. 3 full trials each. 504 scored scenario-trials in total. the systems were tested one at a time, so they never competed for resources. the setups: 2 sparks: deepseek-v4-flash-0731 on vllm, tp=2 1 spark: deepseek-v4-flash-0731 through entrpi ds4 v0.5.0, tp=1, with its matching dspark drafter important: this is a deployment-stack comparison, not simply “one gpu vs two.” the engines, quantisation, model files, speculative setup and hardware count are different. the scores: 2 sparks: 79 / 76 / 79 mean: 78.0 1 spark: 74 / 71 / 73 mean: 72.7 the 2-spark stack scored higher in all three runs, by 5.3 points on average. but here is the honest statistics bit: the paired 95% confidence interval still included no difference. in normal english: the observed gap favours 2 sparks, but it was not statistically decisive under this test. so i am not going to pretend we proved a universal quality winner. where the 1-spark version was better: • parameter precision • multi-step tool chains • knowing when to refuse • error recovery • handling a large tool list • creative tool composition where the 2-spark version was better: • structured reasoning • following detailed instructions • remembering context and state • autonomous planning • structured output • hard mode hard mode was the biggest quality gap: 2 sparks: 57.7 1 spark: 34.3 both were perfect at basic tool selection and tool-eval’s code-pattern category. that is not the same thing as a full coding benchmark. proper coding tests are still to come. reliability also leaned toward 2 sparks. scenarios passed in all 3 runs: 2 sparks: 53/84 1 spark: 50/84 scenarios passed at least once: 2 sparks: 62/84 1 spark: 58/84 speed was not close. average request latency: c1: 2 sparks: 1.30s 1 spark: 3.97s c2: 2 sparks: 1.73s 1 spark: 5.10s c4: 2 sparks: 2.19s 1 spark: 7.75s c8: 2 sparks: 2.79s 1 spark: 14.02s so the 2-spark deployment was roughly 3x faster for one user and around 5x faster by 8 concurrent users. the simple conclusion: if you only own one spark, rejoice. the ds4 version is genuinely capable. it stayed surprisingly close overall and actually beat the dual setup in several useful tool categories. if you own two sparks, do not unplug one. for my current priorities, the dual-spark vllm stack is the stronger deployment: dramatically faster, slightly more repeatable, and better when the workflow becomes difficult, stateful or heavily constrained. the quality result itself was not statistically decisive, so i am not claiming a universal winner. also, neither system gets a free security victory lap. both stacks failed prompt-injection cases, and both leaked a fake api key hidden inside an injected file in at least some trials. this test measures tool use and agentic workflow behaviour. it does not yet prove which stack is smarter overall, better at full coding, better at long context or more stable over a 12-hour production run. those tests are next. all six raw trials, system manifests, paired scenario rows, disagreements, safety results and checksums are preserved in my private blackwellbench lab repo. the most honest takeaway: one spark is far better than people expected. two sparks are still the better choice for my current workload.
BlackwellBoy@Blackwellboy

deepseek v4 on 2 sparks: works great deepseek v4 on 1 spark: the actual goal/flex deepseek v4 on 0 sparks: this is called an api and i refuse results will tell what is viable, thanks to @bleysg we can run them head to head. coming soon.

English
11
4
47
4.2K
Mia
Mia@MiaAI_lab·
Run Inkling-Small on 2x @NVIDIAAI DGX Spark with 1M context ✨ - SGLang with dspark speculative drafts ~34 tok/s (single stream) ~79 tok/s (6 concurrent sessions) Unlike DeepSeek V4 Flash which supports only text, Inkling-Small also supports images, audio, and video. Special thanks to @u1tra_instinct for his Inkling-Small work. Get it here: github.com/MiaAI-Lab/Inkl…
English
11
11
94
5.9K
Alec Fong
Alec Fong@alecqfong·
Current local ai stack 2x DGX Sparks - DeepSeek v4 flash (90 tok/sec) - implementation & sub agents 2x DGX Stations - GLM 5.2 nvfp4 (120 tok/sec) - planning and advising Brev for cluster networking, always on cloud agent, and orchestration
Alec Fong tweet mediaAlec Fong tweet mediaAlec Fong tweet media
English
103
68
1.6K
250.7K
Kai
Kai@unseenmars_·
@LD_Theron we ordered them from amazon🖖
English
1
0
0
73
LM junior
LM junior@LD_Theron·
@unseenmars_ Congratulations 🥳. Do you know where I can get that too? I'm looking for around 8.
English
1
0
0
80
Kai
Kai@unseenmars_·
iykyk
Kai tweet media
Suomi
54
19
510
29.2K
Kai
Kai@unseenmars_·
@yishan yes sirrrrrr
English
0
0
1
88
Kai
Kai@unseenmars_·
@melvindvivas yes yes and yes DeepSeek v4 flash felt like THE model for local workflow now. Hope to get nvfp4 soon.
English
2
2
6
486
Aidan
Aidan@Boondit·
@witcheer Me too! Appearently just 2 or them can run the new deepseek flash model. Very expensive though
English
1
0
3
108
Kai
Kai@unseenmars_·
@usr_bin_roygbiv I also found my coworkers love playing with these DGX. They're hardcore infra engineers and they do this for a living. It's like F1 drivers who still drive for fun on their days off 🏎️
English
1
0
3
170
Kai
Kai@unseenmars_·
It's for people who want to own open frontier intelligence like Inkling-small and DeepSeek V4 Flash! Lots of people are already running these locally in their actual workflow.
English
1
0
1
206
Goody
Goody@GoodyTTD·
@unseenmars_ Just cardboard and stickers on it? Just hand them over I collect junk
English
1
0
3
679
Kai
Kai@unseenmars_·
@johnny_ver_2 How about inkling-small and DeepSeek v4 flash? And yes sm120 will be our focus too
English
0
0
7
351
Johnny 2.0
Johnny 2.0@johnny_ver_2·
@unseenmars_ $18,796 for ~50 tok/s on Qwen3.6-27B no thanks... 4x RTX 5090's instead
English
1
0
3
400
Kai
Kai@unseenmars_·
@mbaril010 We want to make sure people can run models like Qwen, Inkling, GLM, DeepSeek with good optimization with @sgl_project
English
1
1
6
545
Kai
Kai@unseenmars_·
@DegenApeDev for the local AI people 🫡🫡
English
0
0
2
282
Kai retweetledi
Byron Hsu
Byron Hsu@hsu_byron·
Scaling Kimi K3 on H200s with @sgl_project @lmsysorg While most discussion on X focuses on Blackwell, I thought it would be useful to share what we learned while tuning Kimi K3 for throughput-oriented workloads on H200s. - TP+EP beat DP+EP for K3. Only about 1/4 of its layers use gated MLA, so DP attention provides less benefit from avoiding KV-cache duplication. Having each DP rank process every attention head also substantially increased latency. This differs from our K2.6 and GLM setups, where we generally prefer DP+EP. - FP8 KV cache worked well. We measured no quality or throughput regression in our evaluations, while effectively doubling KV-cache capacity. - TP32+EP32 provided significantly more cache capacity: approximately 4M KV-cache tokens per replica, compared with about 800K using TP16+EP16. - For low cache-hit workloads, consider disabling radix cache. Prefix caching requires four Mamba-state slots per sequence in our setup; without radix cache, each sequence needs only one. - Non-speculative decoding won at high concurrency. Above roughly 64 concurrent requests, draft-and-verify overhead outweighed the savings from accepted speculative tokens in our tests. - A few caveats: symmetric memory caused instability in our multi-node environment. We also observed hangs when combining speculative decoding, HiCache, and high KV pressure. Both issues are still under investigation. We’ve been running this configuration at full load across roughly ~O(1000) chips for several days. Aside from hardware failures, it has been reliable. Feedback, discussion, and better configurations are very welcome! Sharing my launch command in the first comment. h/t to @hanming_lu @liin1211 @0xishand @ChengWan17 @NealBayya for the discussion
English
8
30
361
65.8K