KVCache.AI

59 posts

KVCache.AI banner
KVCache.AI

KVCache.AI

@KVCache_AI

Hi, this is https://t.co/EO7MXLjRIs official account. We build systems for efficient LLM serving, including KTransformers and Mooncake.

Beijing Katılım Ağustos 2018
107 Takip Edilen792 Takipçiler
KVCache.AI
KVCache.AI@KVCache_AI·
@baboonAI4S Yes, this is our official account. Thanks for checking!
English
2
0
0
10
KVCache.AI
KVCache.AI@KVCache_AI·
Excited to be part of vLLM’s Day-0 support for Kimi K3! 🚀 Mooncake powers production-scale prefill/decode disaggregation in the vLLM serving stack, enabling efficient KV cache sharing for agentic workloads and scalable deployment of Kimi K3 from @Kimi_Moonshot. Huge congratulations to @vllm_project and everyone involved in bringing the largest open-source model into production! Learn more: vllm.ai/blog/2026-07-2… #KimiK3 #vLLM #Mooncake #LLMInference #OpenSourceAI
English
2
6
32
2.6K
KVCache.AI
KVCache.AI@KVCache_AI·
Following Mooncake, AgentENV marks the next project we've built and open-sourced together, supporting the infrastructure behind Kimi K3’s agentic RL training. Grateful to the @Kimi_Moonshot team for the close collaboration, and looking forward to building more open AI infrastructure together.
Kimi.ai@Kimi_Moonshot

We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast snapshot, resume, and fork support for large-scale parallel agent workflows. Explore on GitHub: github.com/kvcache-ai/Age…

English
3
2
30
2.7K
KVCache.AI retweetledi
LightSeek Foundation
LightSeek Foundation@lightseekorg·
TokenSpeed is only two months old, yet over the past two weeks we've had the privilege of being the Day 0 inference partner for both @thinkymachines's Inkling and @Kimi_Moonshot's Kimi K3. Grateful to both teams for their trust, and proud of what our small, fast-moving team has accomplished so far. 🚀 #Kimi #K3 #Inkling #5-deployment" target="_blank" rel="nofollow noopener">huggingface.co/moonshotai/Kim… #inkling-availability" target="_blank" rel="nofollow noopener">thinkingmachines.ai/news/introduci…
English
1
5
13
1K
KVCache.AI
KVCache.AI@KVCache_AI·
Great to see TokenSpeed delivering Day-0 support for Kimi K3! 🚀 Mooncake powers the data plane behind TokenSpeed's disaggregated serving stack, transferring multimodal embeddings and KV cache across serving stages. Together with TokenSpeed's unified flat KV layout, this enables scalable LLM serving for next-generation model architectures. Congratulations to the @LightSeekOrg team and @Kimi_Moonshot on this impressive release. Looking forward to enabling even more next-generation model architectures together. More details: lightseek.org/blog/tokenspee… #KimiK3 #TokenSpeed #Mooncake #LLMInference #OpenSourceAI
English
2
4
19
1.5K
KVCache.AI
KVCache.AI@KVCache_AI·
Proud to be part of the Day-0 ecosystem for Kimi K3! 🚀 Mooncake powers SGLang's Day-0 support for Kimi K3, providing distributed cache storage for both KV cache and KDA state checkpoints, enabling efficient prefix cache reuse across instances while keeping the new KDA architecture production-ready. Huge congratulations to the @lmsysorg team and @Kimi_Moonshot on an incredible release. Looking forward to seeing what the community builds with K3!
LMSYS Org@lmsysorg

SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles @radixark! How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. We've passed Kimi Vendor Verifier and are ready for production! Thanks to @Kimi_Moonshot, @nvidia, @AMD, @KVCache_AI, @modal, and @baseten for building this with us, and to @googlecloud, @nebiustf, @fal, @digitalocean, @runpod, @DeepInfra and @gmi_cloud for serving K3 on SGLang. Blog, cookbook, benchmarks in the comments. P.S. This demo video? Kimi K3 made it itself. Play the game 👇

English
1
7
35
2.7K
KVCache.AI
KVCache.AI@KVCache_AI·
🚀 Mooncake v0.3.12 is out! This release brings major upgrades across Mooncake Store, Transfer Engine, Expert Parallelism, and platform support. A huge thank you to our amazing community: we're thrilled to welcome 53 new contributors in this release! ❤️ ✨ Highlights: • Distributed SSD-backed KV cache pooling • Smarter transfer scheduling with intent- & policy-based routing and deadline-aware QoS • Broader platform support: TPU/PJRT, HPE Slingshot/CXI, AMD HIP/RDMA, and Sunrise • Published wheel packages for AWS EFA and MUSA • Official Docker Hub images • Many performance optimizations and reliability fixes throughout the project #Mooncake #AIInfrastructure #LLM #OpenSource
KVCache.AI tweet media
English
3
11
83
13.1K
KVCache.AI
KVCache.AI@KVCache_AI·
🚀Mooncake now supports SSD Offloading for KV Cache. As agentic workloads become the norm, KV cache lifetimes are getting much longer, but keeping everything in DRAM simply doesn't scale. With Mooncake's distributed SSD tier, you can: ✅ Expand KV cache capacity far beyond memory. ✅ Preserve long-tail KV caches instead of evicting them. ✅ Increase KV cache hit rates. Even small gains translate nonlinearly into significant prefill speedups. Read more: kvcache.ai/blog/scaling-k… #LLM #Inference #KVCache #AIInfrastructure #Mooncake
English
1
5
20
2.1K
KVCache.AI
KVCache.AI@KVCache_AI·
It's exciting to see Mooncake being used in the latest 1T-scale RL infrastructure from @PrimeIntellect .🚀 As LLM systems continue to scale, efficient KV cache management is becoming critical infrastructure. Prime Intellect highlights Mooncake Store as a centralized KV cache backend that pools RAM/disk across nodes into one large cache, accessible by any inference worker from any node — this provides a significant advantage, especially when using more sophisticated routing strategies. Read more: primeintellect.ai/blog/rl-at-1t-… #LLM #RL #AIInfrastructure #KVCache
Prime Intellect@PrimeIntellect

One Mooncake store pools KV cache across all nodes, so any worker can reuse any prefix. The router picks workers by a score over load, queue depth, KV usage and prefix overlap. You get cross-replica cache hits with balanced routing across the whole deployment.

English
0
10
28
6K
KVCache.AI
KVCache.AI@KVCache_AI·
Congrats to the @vLLM_Project and excited to see Mooncake powering the full online DSpark training pipeline on GB300 NVL72 🚀 Mooncake efficiently moves hidden-state from vLLM nodes to Speculators through RDMA, eliminating the need for massive hidden-state storage in offline training. Excited to keep pushing the limits together!
Michael Goin@mgoin_

DSpark update: Turns out with a little Speculators+Mooncake, I'm able to scale training on GB300 NVL72! 9 vLLM nodes serve the full GLM 5.2 FP8 verifier -> Mooncake RDMA store -> 6 nodes train the DSpark with FSDP (DP=24). 125k prefill tok/s, 1.5 steps/s, full online training :)

English
0
9
38
4.3K
KVCache.AI
KVCache.AI@KVCache_AI·
🧠How much KV cache is enough for LLM serving? 📉Too little → poor reuse. 📦Too much → wasted storage. ⚡The sweet spot is where marginal speedup of prefill starts to flatten. 🛠️In this blog, we show a simple and practical way to find that sweet spot for KV cache capacity planning, using our online KV Cache Hit Rate Simulator. Read more 👉 kvcache.ai/blog/calculate… #KVCache #LLMServing #Optimization
English
1
6
14
1.7K
KVCache.AI
KVCache.AI@KVCache_AI·
GLM-5.2 vs GLM-5.1: What changed in the KV Cache? 🚀 1M-context cache comparison: GLM-5.1: 41.8 GiB KV cache + 9.3 GiB Indexer cache GLM-5.2: 41.8 GiB KV cache + 2.5 GiB Indexer cache ⚡ Key takeaway: GLM-5.2 uses shared index, making long-context inference more practical. Explore KV Cache with our calculator: kvcache.ai/tools/kv-cache… #GLM #KVCache #LongContext #LLMInference
KVCache.AI tweet media
English
2
14
84
46.9K
KVCache.AI
KVCache.AI@KVCache_AI·
@1i__is Thank you! This means a lot to us 🙏 More open-source work in this area is coming soon. Stay tuned!
English
1
0
2
69
1iis, the New Team
1iis, the New Team@1i__is·
i love these guys they're hard on the narrow problem space that makes or breaks AI setups please don't stop keep being awesome we need you!
KVCache.AI@KVCache_AI

🚀 We just launched KV Cache Analyzer by KVCache.AI! 📊 Analyze KV cache hit rates and estimate prefill throughput speedup under different cache budgets and eviction policies. 🧪 Use preset traces or your own local traces, choose the model and parameters, and see how KV cache reuse improves LLM inference performance across different settings. 👉 Try it now: kvcache.ai/tools/kv-cache…

English
1
0
2
167