Chunked KV-Cache Compression
- fragments
- 1小时前
- 6 Views
- 0 Comments
- 106 Words
Reader mode is not fully supported on this site; please use it with caution.
A nice paper showing a problem inside the Chunked KV-Cache Compression. https://arxiv.org/pdf/2609.36322 Models like dpskv4 used chunked kv cache compression, it compressed 8 tokens each time into a kv vector to save kv memory. However the paper finds a pattern called phase specialization: for example, a token may be preserved well when it falls into 1st token of the 8 token chunks (phase 1), but become degraded if places it in 2nd/phase 2.
(DPSKv4 use sliding window 8 with steps 4. That means for compression, it first compressed 0-7, then 4-11, then 8-15 etc)
To learn what is chunked kv cache compression
https://arxiv.org/pdf/2502.00299
