<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Kv Cache on Li Cao&#39;s Blog</title>
    <link>https://l1-ca0.github.io/tags/kv-cache/</link>
    <description>Recent content in Kv Cache on Li Cao&#39;s Blog</description>
    <generator>Hugo -- 0.148.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 11 Oct 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://l1-ca0.github.io/tags/kv-cache/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>KV, Prefix, and Session Caching in Co-located and Disaggregated LLM Serving</title>
      <link>https://l1-ca0.github.io/posts/follow-the-kv-cache/</link>
      <pubDate>Sun, 11 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://l1-ca0.github.io/posts/follow-the-kv-cache/</guid>
      <description>&lt;p&gt;Splitting prefill and decode onto separate GPU pools turns the KV cache from local state into data that has to be routed, moved, and restored. Moving it once turns out to be cheap: in a worked example of a 70B-class coding agent on B200s, handing a 32K-token prompt&amp;rsquo;s KV from prefill to decode adds about 5% to time to first token over 400 Gb/s RDMA, even without overlap. What is hard is keeping KV reusable. Reuse pays off enormously, since session reuse alone cuts the agent&amp;rsquo;s prefill work about 24×, but stored KV competes for GPU memory, the resource that already limits decode. Disaggregation adds a second difficulty: prefix hits now depend on routing across two pools, and the newest copy of each session sits on the decode side, where pulling it back every turn costs more as the conversation grows. The upshot is that for any given set of GPUs, caching choices (session reuse, affinity, memory tiers, KV format) determine the traffic it can serve far more than the decision to disaggregate does.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
