Research ·
Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
AI brief
AI-writtenWhy it mattersIt helps reduce inference memory usage and improve long-context processing speed.
A hierarchical KV cache eviction method balancing relevance and diversity outperforms existing baselines on long-context tasks
What happened
The research team addressed a key limitation of existing KV cache eviction methods — which only rely on average attention scores calculated within the observation window — by proposing a unified scoring formula that integrates attention dispersion and redundancy against already selected tokens. The new approach requires no extra forward propagation, and also explores a hierarchical configuration scheme that adjusts coefficients across different network depths. Tests were conducted on the Mistral-7B model across 16 English LongBench datasets: a globally fixed diversity coefficient improved performance on 13 datasets, with an average macro-metric gain of +1.1. For paragraph retrieval tasks, the hierarchical configuration delivered a 9.6 point improvement over the baseline at a 64-token cache budget, and a 13.2 point improvement over the global configuration at a 128-token cache budget.
Key facts
- Test model
- Mistral-7B
- Test benchmark
- 16 English LongBench datasets
- Global configuration performance gain
- Improvements on 13 out of 16 datasets at a 64-token per layer cache budget, with an average macro-metric gain of +1.1
- Hierarchical configuration performance gain for retrieval tasks
- 9.6 point improvement over baseline at a 64-token cache budget, 13.2 point improvement over global configuration at a 128-token cache budget
- Core method feature
- Scoring calculations require no additional forward propagation
Background
Current mainstream KV cache eviction methods such as SnapKV and PyramidKV only rank tokens by average attention within a small observation window. They fail to account for differences in token attention distribution, redundancy with already selected tokens, and adjusted scoring weights for different network depths.
Why it matters
This method improves KV cache utilization efficiency with no extra computational overhead. For the industry, it can lower deployment costs for long-context large models. For developers, it enables better long-text task performance under limited cache budgets without requiring changes to existing inference pipelines. For end users, it will eventually enable smooth long-document retrieval and Q&A experiences even on devices with low VRAM.
What to watch
Going forward, stakeholders can track the method's validation results on larger models and multilingual long-context tasks, as well as its integration and deployment progress with mainstream inference frameworks.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.