Paper notes: Trading More Storage for Less Computation
I recently studied Mooncake's FAST '25 paper, "Trading More Storage for Less Computation." The whole idea is "trade cheap storage for expensive computation".
The problem
The KVCache stored in GPU HBM is so huge and expensive. Many systems keep KVCache local in each GPU, which makes it hard to reuse across requests.
Mooncake's solution
Mooncake's solution has three parts:
- Disaggregate prefill and decode. The two stages in inference have opposite resource requirements — prefill is compute-bound, decode is memory-bound.
- Build a global KVCache pool across the cluster's machines' CPU DRAM and SSD, connected by high-bandwidth RDMA NICs (which don't need the OS to move the data).
- Prefix-hash every KV block by its token prefix. Once there's a cache hit, prefill is skipped entirely.
The results
Evaluated on two SLOs — TTFT for prefill latency, TBT for decode smoothness — Mooncake shows up to 2.36× higher cache hit rate and 48% less prefill compute versus local caching, and 59%–498% higher effective request capacity versus vLLM (including vLLM's own prefix-caching and chunked-prefill). In production, Mooncake lets Kimi handle ~75% more requests under the same SLOs. They also chart how the P/D ratio influences SLOs — performance is best in the 7P9D to 10P6D range.
What's next
The paper really inspired me, and it makes me want to dig into the Mooncake codebase. I might start reading the repo right now.
