George M
121 posts





GLM 5.2 DSpark preview is here! ✨ huggingface.co/RedHatAI/GLM-5… This is the first DSpark speculator for a non-DeepSeek frontier model, trained with Speculators and running on vLLM nightly for ~1.5× faster decode for GLM-5.2-FP8 on 4×B300. Stronger checkpoints to come!



And no, worker_threads / Worker is not that. 1 global object, 1 garbage collector, every object accessible across threads. No postMessage overhead. No cloning data unnecessarily. A thread should cost < 2 MB. That’s what good looks like.








For those running multiple agents on the same codebase locally, what are you doing? If something else comment below.












Considering the current pinch-bench results, I kind of want to run a quant gauntlet with a few of these top models to see the usefulness drop off etc. Would folks be interested in that?











Experimental: Use with caution

TT-Ascalon is officially IP released. Go build. RiscV is now high performance. Really happy with the team, quality of the release, quality of the support IP and DV infra. Open source hardware is great. tenstorrent.com/ip/risc-v-cpu











