
gnubyte
28 posts

gnubyte
@gnubyte_
Software engineer and Creator 503CM, LLC IT Manager Previously: Sr. Software Engineer @ Western Digital, L4 Engineer @Verizon






if you're trying local ai for the first time and you're on an 8 to 12gb gpu, this is the best model + agent pairing i've found, late july 2026. everything you need to run a real local agent on hardware you already own. one best model, one best harness. 8gb of vram and balls. THE MODEL bonsai 27b at 1-bit, a full 27b crushed to 3.9gb by @PrismML. huggingface.co/prism-ml/Bonsa… 42 tok/s fresh, and it holds the full 128k context window in 6.8 of 8 gigs. still usable deep in (~13 tok/s at a packed 128k, above reading speed). no crashes in that range. llama server flags: llama-server -m Bonsai-27B-Q1_0.gguf -ngl 99 -np 1 -c 131072 -ctk q4_0 -ctv q4_0 -fa on THE AGENT hermes agent by @NousResearch. it runs this model the cleanest i've seen, and onboarding auto-detects your local endpoint, no config dance like in other bloatwares. github.com/NousResearch/h… one catch: bonsai needs prismml's llama.cpp fork to load, stock won't. link in the replies. that's where the entry point is now. tool-calling held zero failures across my runs, the 128k window fits comfortably, and it reasons before it acts instead of firing blind. this is the setup that needed a 24gb card six months ago. now it's an 8gb card in a drawer. nvidia's leaning open, a wave of labs are shipping open weights, and i hope we keep seeing a flood of small specialized models built for the hardware people actually own, 8 to 24gb. that's where labs should meet users if they want adoption. i'll keep running experiments and showing you exactly what this can do.


I've been trying to love @poolsideai Laguna S 2.1. But it can't even single shot Tetris. Is this a Q4_K_M problem, or a model problem? Qwen 3.6-27B crushes this task. So does Gemma-4-26B-A4B, and Qwen3.6-35B-A3B most of the time. A 118B 8B should be able to crush this. What am I doing wrong?


Use my repo if you want to run it yourself exactly how I run it. github.com/MiaAI-Lab/Qwen…


laguna s 2.1: strong open coder, but its headline thinking mode makes it worse on held-out work, plus one integrity blind spot you can prompt around. full guide in my new offlabel repo: github.com/TheTom/offlabe… x.com/i/article/2080…




🎆New decode record 🎆 New: 22.8 tok/s 🔥 Old: 19.8 tok/s GLM 5.2 on 2 DGX Sparks I quantized the dspark speculator to W4A16 (7.1GB -> 4.6GB w/ near identical acceptance) A smaller drafter + bumping KV cache = 96k context ⚡️⚡️⚡️⚡️⚡️













Thank you @nvidia for the hookup on this DGX Workstation. This will crunch a *lot* of high quality tokens here! This thing is a total beast.


We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix. Our mission is to run the world's inference. This round accelerates production of our inference clusters. We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.








