Sabitlenmiş Tweet
Joe Muller
11.1K posts

Joe Muller
@BosonJoe
Local AI enthusiast, suffering from token sickness 2x DGX Spark, 2x RTX 5090
Virginia, USA Katılım Nisan 2019
1.1K Takip Edilen7K Takipçiler

@BosonJoe In the "Put the deployment config on both nodes" section, it says:
# on BOTH nodes
git clone github.com/jtmuller5/howt… ~/howtospark
English

@plotarmordev For code content, everything else stayed at about 42-45 tok/sec
English

@BosonJoe Hi, thanks for sharing your recipe!
It looks like the link to github.com/jtmuller5/howt… is broken and returns a 404 error.
English

@MrPeterLMorris NVFP4 is almost lossless for this model. You could also fit the full BF 16 variant if you wanted to it would just be slower
English

@BosonJoe I would need to know how intelligence compares. I've seen tricks to speed up models that made them useless.
English

Qwen3.6 27B NVFP4 Single Spark Recipe ⚡️
🚀 vLLM 0.26.0, no patch
🪣 262K context
👾 31 tok/sec decode
You could probably run this and the 35B MoE Qwen model on the same Spark if you wanted the best of both
howtospark.com/recipes/qwen3-…
English

@kkiran Personally I would choose moe models. Anything dense is going to be pretty slow
English

@BosonJoe @NVIDIAAI It seems like you’re still digging into the GLM 5.2. I’d be curious what you end up using as a default and what you feel like they’re best used for.
I’m wondering as well if Qwen 3.8 open weights drop and the 35B of that ends up being a major jump forward for individual sparks
English

GLM 5.2 on 2 @NVIDIAAI DGX Sparks @ 24.7 tok/sec 🚀
Next step is to claw back some quality while keeping the speed steady
We've come a long way. When I stared it was 4 tok/sec 🐌
Joe Muller@BosonJoe
🎆New decode record 🎆 New: 22.8 tok/s 🔥 Old: 19.8 tok/s GLM 5.2 on 2 DGX Sparks I quantized the dspark speculator to W4A16 (7.1GB -> 4.6GB w/ near identical acceptance) A smaller drafter + bumping KV cache = 96k context ⚡️⚡️⚡️⚡️⚡️
English

@vectorpolygon @CtrlAltDwayne They made inference faster for every topic by making the model less capable at biology and cyber security
Do you know how model inference works?
English

@CtrlAltDwayne By sacrificing intelligence that matters to people who actually need it, while making people like you say "it's insane".
English

@imost @RyanFogleOP12 Dang, looking into it. Thanks for the heads up, I should have a fix out this morning
English

@BosonJoe @RyanFogleOP12 Just tried to submit a benchmark for your GLM 5.2 recipe but got this. Let me know if there's anything I need to change on my side.

English

First contributor to howtospark!
Thanks @RyanFogleOP12
Check out the MCP server to join up 🤝

Ryan Fogle@RyanFogleOP12
It is here let’s go! @nvidia @NVIDIAAI #LocalAI #NVIDIA #AI #DGXSpark
English

@nietonchique @NVIDIAAI Not sped up, but this is very experimental and I would not use it to do anything yet
English











