Sabitlenmiş Tweet

GLM-5.2 on 8× DGX Spark — v18 update 🚀
Jump from 1,200 → 1,329 t/s prefill
Peak decode hits 66 t/s (was 54)
Repo with both v16 and v18 builds:
github.com/ciprianveg/gb1…
Forum post with benchmark:
forums.developer.nvidia.com/t/glm-5-2-int4…

English



















