Bờm

52 posts

Bờm

Bờm

@quatmo

Dumb guy with a small areca spathe fan.

Vietnam Katılım Temmuz 2009
201 Takip Edilen52 Takipçiler
Bờm retweetledi
Perplexity Developers
Perplexity Developers@perplexitydevs·
The Perplexity CLI is now available, giving coding agents the ability to search the web. Copy this to your agent to get set up: "Read: github.com/perplexityai/a… and install this skill."
English
38
84
658
150.6K
Ryan Fogle
Ryan Fogle@RyanFogleOP12·
@sudoingX I just got my spark yesterday so not a ton of testing, but I have been testing Krea2 through ComfyUI
English
1
0
1
578
Sudo su
Sudo su@sudoingX·
local ai people running a dgx spark, what image model are you actually running on it, and what would you start me on? flux, or whatever the community's actually landed on. curious what holds up on 128gb of unified memory.
English
26
0
60
11.9K
Aydan | AI Builder & iOS Dev
@qwen_cloud Damn, Alibaba's website is worse than Google's website. It's damn so complicated! Can’t understand where you can use your token subscription etc!
English
3
0
62
4.5K
Qwen Cloud
Qwen Cloud@qwen_cloud·
Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out. Can't wait to hear what you build. Stay tuned! Token Plan : international:qwencloud.com/pricing/token-… China:platform.qianwenai.com/pricing/token-…
English
131
385
6K
1.4M
Bờm
Bờm@quatmo·
🚀 Just found a new flag that's improving my local LLM setup A LOT for AI agents like Hermes! --enable-prefix-caching on my optimized Qwen3.6-35B-A3B-NVFP4 vLLM recipe (DGX Spark) is a total game-changer: ✅ 93% prefix cache hit rate on warm agent turns (~34K token prompts) ✅ ~3× faster warm TTFT — saving 7–10 seconds per turn ✅ Works beautifully with MTP speculative decoding + FP8 KV cache ✅ Keeps system prompts + tool definitions stable for max reuse Full optimized recipe, start.sh, benchmarks, and client config here: github.com/binhdnguyen/Qw… Massive win for real agentic workflows on local hardware. 🔥 #LocalLLM #vLLM #Qwen #AIagents #Hermes
English
0
0
1
95
Bờm retweetledi
0xSero
0xSero@0xSero·
GLM-5.2 Nvidia NVFP4 No pruning, no quantising, exact same model card. 110 tok/s single stream. My friend is a genius, holy fudge. Here's the repo, he got full DS4-Flash on a 5090 + DDR5 github.com/kacper-daftcod… at 38.5 tok/s This is a breakthrough.
English
101
101
1.3K
156.3K
Bờm
Bờm@quatmo·
@MiaAI_lab I am so happy with mine, just 67w power draw while 95% gpu load. Very efficient.
English
0
0
0
65
Mia
Mia@MiaAI_lab·
Anyone — no matter how big their account or how important they think they are, who claims the DGX Spark isn’t suitable for real work is either ignorant or a liar. No in-between.
English
51
8
190
9.8K
Bờm retweetledi
Unsloth AI
Unsloth AI@UnslothAI·
DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳 Run lossless DeepSeek-V4-Flash on 168GB RAM. 3-bit works on 110GB Mac, RAM, VRAM setups. We improved the chat template. Run via Unsloth Studio or llama.cpp. Guide: unsloth.ai/docs/models/de… GGUF: huggingface.co/unsloth/DeepSe…
Unsloth AI tweet media
English
46
115
924
95.5K
Bờm
Bờm@quatmo·
@MiaAI_lab @NVIDIAAI How can you utilize multiple concurrent for 1 task, ie. Plug in hermes agent, is it possible?
English
0
0
0
46
Mia
Mia@MiaAI_lab·
Run Tencent Hy3-295B on 2x @NVIDIAAI DGX Spark with EASE🔥 128k context • MTP • Up to 6 concurrent sessions Performance: ~20 tok/s in single session ~58 tok/s in 6 concurrent sessions (cumulative) First impressions: the model feels very solid. More benchmarks and real-world tests coming soon! Big thanks to the DGX community for the original work. Get it here 👇 github.com/MiaAI-Lab/Hy3-…
English
18
16
131
12.6K
Bờm retweetledi
Mia
Mia@MiaAI_lab·
Which local model is the best for Agentic Workflows for a single @NVIDIAAI DGX Spark? (or any other 96-128gb VRAM rig) After running 84 scenarios, 16 categories, 8 trials each, on Hermes-Agent style multi-turn tool orchestration, there is a very clear winner. 🏆 Qwen 3.6 35B A3B Q8_K_XL is #1 It’s the only model that hit perfect scores across the board with zero catastrophic failures. The Full Ranking: Qwen 3.6 35B A3B UD Q8_K_XL — 91.0 Qwen 3.6 27B NVFP4 — 89.0 Qwopus 3.6 27B Coder MTP — 85.2 DeepSeek V4 Flash Q2 — 86.5 Agents-A1 Q8_0 — 83.4 Gemma 4 26B — 81.4 Nemotron 3 Nano Omni 30B — 79.0 Bottom line: If you’re running agents locally on a DGX Spark or any 96-128GB rig in 2026, Qwen 3.6 35B Q8_K_XL is currently the move. Full report + deep dive 👇 github.com/MiaAI-Lab/Best…
Mia tweet media
English
111
45
558
37.2K
Bờm retweetledi
0xSero
0xSero@0xSero·
I'm buying a second DGX Spark, my goal is very simple. In daytime I will have DS4 Flash for chat running across 2x Sparks and GLM-5.2-REAP running on 4x 6000s I am setting up a goal builder, which recommend me loops to run overnight. At night I will have GLM-5.2 (not REAP) running with PP = 2 with stage 1 = TP4 with 6000s and stage 2 being TP2 with the DGX Sparks. To make this work I need to get really creating, but I think I can get 25-40 tok/s decode for the Nvidia NVFP4 At night I don't care much about tok/s as long as it'll be large enough to make progress over 8 hours.
0xSero tweet media
English
68
35
839
88.3K
Bờm retweetledi
Mia
Mia@MiaAI_lab·
Nvidia did it again! 🤯 @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on DGX Spark for single sessions and by 23-25% for 2-4 concurrent sessions. @NVIDIAAI 👏
Mia tweet media
English
60
53
703
69.2K
Bờm retweetledi
Sudo su
Sudo su@sudoingX·
if you're just getting into local llms, do yourself a favor and start by building llama.cpp from source. not ollama, not lm studio. build llama.cpp once, it's genuinely just a git clone and a make command with cuda on, and it clicks. you see the flags, you control the quant, you run any gguf on the planet, and llama-bench gives you real numbers instead of a vibe. when something's slow, you know why, and you can fix it. ollama and lm studio are fine for "just chat with a model." but if you actually want to understand local inference, they're a ceiling, not a foundation. start one level deeper. it pays off every single day after.
English
52
84
1.1K
52K
Bờm retweetledi
Mia
Mia@MiaAI_lab·
All links & repos were personally tested and verified by me. DeepSeek v4 Flash x.com/MiaAI_lab/stat… Step 3.7 Flash x.com/MiaAI_lab/stat… Qwen 3.6 27b x.com/MiaAI_lab/stat… Qwen 3.6 35b NVFP4 x.com/MiaAI_lab/stat… DeepSeek v4 Flash REAP huggingface.co/0xSero/DeepSee… My recommendation below👇
Mia@MiaAI_lab

Did you know Qwen 3.6 35B NVFP4 with 256k context, MTP, kv cache fp16 & image support takes around 68gb of VRAM? My @NVIDIAAI DGX Spark can do 110 tok/s 👇

English
4
2
41
16.6K
Bờm retweetledi
Mia
Mia@MiaAI_lab·
Did you know Qwen 3.6 35B NVFP4 with 256k context, MTP, kv cache fp16 & image support takes around 68gb of VRAM? My @NVIDIAAI DGX Spark can do 110 tok/s 👇
English
23
5
139
35.3K
Bờm retweetledi
ÆON FORGE ✨
ÆON FORGE ✨@SpaceTimeViking·
Guys I’m speechless my Qwen3.6-27B-AEON-ULTIMATE-NVFP4 is earning its name! The same complex puzzle that took stock Qwen3.6-27B ~130 seconds running with the same engine and same optimizations AEON-ULTIMATE COMPLETED IN ONLY ~15 SECONDS!!!
ÆON FORGE ✨ tweet media
Mike Gannotti@MichaelGannotti

x.com/i/article/2070…

English
11
9
99
8.1K
Bờm retweetledi
tonbi
tonbi@tonbistudio·
This week's Hermes Agent masterclass is on subagents and delegation! In this clip, I set up a more expensive parent agent for orchestration with a cheaper model for the children subagents doing the research. See how you can get 5x cheaper workflows with this setup!
English
23
29
425
43.8K