Research ·

Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction

77Developing1 reportarXiv cs.CL

AI brief

AI-written

Why it mattersIt helps reduce inference memory usage and improve long-context processing speed.

A hierarchical KV cache eviction method balancing relevance and diversity outperforms existing baselines on long-context tasks

What happened

The research team addressed a key limitation of existing KV cache eviction methods — which only rely on average attention scores calculated within the observation window — by proposing a unified scoring formula that integrates attention dispersion and redundancy against already selected tokens. The new approach requires no extra forward propagation, and also explores a hierarchical configuration scheme that adjusts coefficients across different network depths. Tests were conducted on the Mistral-7B model across 16 English LongBench datasets: a globally fixed diversity coefficient improved performance on 13 datasets, with an average macro-metric gain of +1.1. For paragraph retrieval tasks, the hierarchical configuration delivered a 9.6 point improvement over the baseline at a 64-token cache budget, and a 13.2 point improvement over the global configuration at a 128-token cache budget.

Key facts

Test model
Mistral-7B
Test benchmark
16 English LongBench datasets
Global configuration performance gain
Improvements on 13 out of 16 datasets at a 64-token per layer cache budget, with an average macro-metric gain of +1.1
Hierarchical configuration performance gain for retrieval tasks
9.6 point improvement over baseline at a 64-token cache budget, 13.2 point improvement over global configuration at a 128-token cache budget
Core method feature
Scoring calculations require no additional forward propagation

Background

Current mainstream KV cache eviction methods such as SnapKV and PyramidKV only rank tokens by average attention within a small observation window. They fail to account for differences in token attention distribution, redundancy with already selected tokens, and adjusted scoring weights for different network depths.

Why it matters

This method improves KV cache utilization efficiency with no extra computational overhead. For the industry, it can lower deployment costs for long-context large models. For developers, it enables better long-text task performance under limited cache budgets without requiring changes to existing inference pipelines. For end users, it will eventually enable smooth long-document retrieval and Q&A experiences even on devices with low VRAM.

What to watch

Going forward, stakeholders can track the method's validation results on larger models and multilingual long-context tasks, as well as its integration and deployment progress with mainstream inference frameworks.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. arXiv cs.CL ↗Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache EvictionThe paper proposes a new KV cache eviction scoring method integrating attention diversity and redundancy.
Back to AI News