Joe Muller

11.1K posts

Joe Muller banner
Joe Muller

Joe Muller

@BosonJoe

Local AI enthusiast, suffering from token sickness 2x DGX Spark, 2x RTX 5090

Virginia, USA Katılım Nisan 2019
1.1K Takip Edilen7K Takipçiler
Sabitlenmiş Tweet
Joe Muller
Joe Muller@BosonJoe·
howtospark . com is live! I'm treating it as a living notebook for my Spark experiments and have a bunch of things I want to add that will help Spark owners hit the ground running
Joe Muller tweet media
English
3
4
28
2.7K
Joe Muller
Joe Muller@BosonJoe·
The DeepSeek V4 Flash Recipe has been updated 🧑‍🍳 There was a bug in the spec-decode config Bumping k from 3->5 increased code content decode tok/sec from 62.5 → 73.2 tok/s (+17%) Working on updating to vLLM 0.26.0 next...
Joe Muller tweet media
English
5
2
83
6.3K
Joe Muller
Joe Muller@BosonJoe·
@plotarmordev For code content, everything else stayed at about 42-45 tok/sec
English
0
0
2
256
Joe Muller
Joe Muller@BosonJoe·
@hnakamur2 Hi, they should all be on the website not the GitHub
English
1
0
0
68
dan ushman
dan ushman@danushman·
made the mistake of trying to use Grok for something... what a POS. Come back in 9 hours, and it's stuck, for the stupidest possible reason, got nothing done. Claude is taking over. Never again. Grok 4.5 is trash and I don't care what anyone says. Absolute trash.
English
18
0
30
2.5K
Joe Muller
Joe Muller@BosonJoe·
16 hours people
Joe Muller tweet media
English
1
2
28
1.7K
Joe Muller
Joe Muller@BosonJoe·
@MrPeterLMorris NVFP4 is almost lossless for this model. You could also fit the full BF 16 variant if you wanted to it would just be slower
English
1
0
0
136
Peter Morris
Peter Morris@MrPeterLMorris·
@BosonJoe I would need to know how intelligence compares. I've seen tricks to speed up models that made them useless.
English
1
0
0
233
Joe Muller
Joe Muller@BosonJoe·
Qwen3.6 27B NVFP4 Single Spark Recipe ⚡️ 🚀 vLLM 0.26.0, no patch 🪣 262K context 👾 31 tok/sec decode You could probably run this and the 35B MoE Qwen model on the same Spark if you wanted the best of both howtospark.com/recipes/qwen3-…
English
7
6
128
7.2K
Adi Singh
Adi Singh@adisingh·
I've reached that point of my company where I hand out NDA's like free candy
English
6
0
48
6.1K
Joe Muller
Joe Muller@BosonJoe·
@kkiran Personally I would choose moe models. Anything dense is going to be pretty slow
English
0
0
1
234
kkiran
kkiran@kkiran·
@BosonJoe contemplated both but Claude says this for a single Spark - 35B and nemo3. If they don't play well, then 27B + 35B or Nemo 3. Thank for the amazing work!
kkiran tweet media
English
1
0
1
281
base
base@basewtf·
@BosonJoe Could run 27B and 35B concurrently on 1 spark?
English
2
0
0
230
Joe Muller
Joe Muller@BosonJoe·
I still use Claude for most things, especially running benchmark tests on the Sparks (it's easiest because I don't have to use memory to host the model I'm using to run experiments) When I do use local models I use Deepseek V4 Flash and/or Qwen 3.6 35B A3B. I only do code related work
English
0
0
0
62
gnubyte
gnubyte@gnubyte_·
@BosonJoe @NVIDIAAI It seems like you’re still digging into the GLM 5.2. I’d be curious what you end up using as a default and what you feel like they’re best used for. I’m wondering as well if Qwen 3.8 open weights drop and the 35B of that ends up being a major jump forward for individual sparks
English
1
0
0
11
Joe Muller
Joe Muller@BosonJoe·
GLM 5.2 on 2 @NVIDIAAI DGX Sparks @ 24.7 tok/sec 🚀 Next step is to claw back some quality while keeping the speed steady We've come a long way. When I stared it was 4 tok/sec 🐌
Joe Muller@BosonJoe

🎆New decode record 🎆 New: 22.8 tok/s 🔥 Old: 19.8 tok/s GLM 5.2 on 2 DGX Sparks I quantized the dspark speculator to W4A16 (7.1GB -> 4.6GB w/ near identical acceptance) A smaller drafter + bumping KV cache = 96k context ⚡️⚡️⚡️⚡️⚡️

English
28
12
262
25.2K
Joe Muller
Joe Muller@BosonJoe·
@vectorpolygon @CtrlAltDwayne They made inference faster for every topic by making the model less capable at biology and cyber security Do you know how model inference works?
English
2
0
0
15
Dwayne
Dwayne@CtrlAltDwayne·
It's insane how token efficient Opus 5 is. I've been hammering it and my weekly usage has barely moved. How did Anthropic make such a capable model so token efficient without sacrificing intelligence and taste?
English
23
3
76
6.8K
😼
😼@vectorpolygon·
@CtrlAltDwayne By sacrificing intelligence that matters to people who actually need it, while making people like you say "it's insane".
English
3
0
0
254
Iman Mostafavi
Iman Mostafavi@imost·
@BosonJoe @RyanFogleOP12 Just tried to submit a benchmark for your GLM 5.2 recipe but got this. Let me know if there's anything I need to change on my side.
Iman Mostafavi tweet media
English
1
0
2
33
Joe Muller
Joe Muller@BosonJoe·
@gnubyte_ @NVIDIAAI No, I would classify this work as experimental/educational Best model on 2 sparks is still Deepseek v4 flash
English
1
0
0
98
gnubyte
gnubyte@gnubyte_·
@BosonJoe @NVIDIAAI Hey man so have you been able to build anything with it yet? It seems like to me the best money has just been on the Qwen3.6-35B NVFP4 model that sparks sorta ship with or try to default to.
English
1
0
0
64