Post

joseph
joseph@gitcommit90·
PSA: DGX Spark + Laguna S 2.1 NVFP4+DFlash is ~15–23 tok/s freeform. No it is not 77 tok/s. No it is not 50 tok/s everyday chat. No it is not a very fast model to run on the DGX Spark. Is it an amazing model? Remains to be seen. Am I also very excited that open weights keep landing next to frontier models? Absolutely. We can all feel excited. We should also be honest and not farm for clicks. If someone’s posting 5× that as “single session decode,” make them define the benchmark.
English
1
0
0
207
Eiso Kant
Eiso Kant@eisokant·
@gitcommit90 Our testing has us at 35-40 tok/s. Would be interesting to see what you’re using here and context length.
English
1
0
3
147
joseph
joseph@gitcommit90·
15–23 was freeform streaming decode, thinking off, short ctx (not even a filled 256k run). GB10, vLLM 0.25.1, official-ish NVFP4+DFlash path, max-model-len 262144. client stream timings, not UI gen tok/s. code prompts landed ~40–44 — so 35–40 tracks if you’re not pure freeform. what’s your prompt + metric?
English
1
0
1
80
Eiso Kant
Eiso Kant@eisokant·
@gitcommit90 @joerowell you’re our resident spark inference expert. I’m always testing with coding agents, so I do hit quite a bit of cache.
English
1
0
0
71
joseph
joseph@gitcommit90·
ha okay but seriously: is there a public recipe for that 35–40? that’s damn near 2× the freeform streaming numbers I get on my GB10 (same model, NVFP4+DFlash, short/empty window, client e2e). if 35–40 is real day-to-day, that will make this a daily driver for me. happy to be wrong about my setup. what should I match exactly (serve flags, client, how tok/s is counted)? the cache hit point is useful; my freeform runs were basically cold/short, not a long agent session riding prefix cache. @joerowell would love the Spark-side checklist if you have one or @eisokant whoever can provide it - please do
English
0
0
0
69
Paylaş