Downloads · 30 days
0
GenomaLabs-com/kv-cache-eviction-mla
kv-cache-eviction-mla is a machine learning model from GenomaLabs-com. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
A reference implementation of H2O-style heavy-hitter + recency KV cache eviction for transformer architectures that use Multi-head Latent Attention (MLA) as introduced in DeepSeek V3 and used in Kimi K2 / K2.6.
Downloads · 30 days
0
Access
Public
Updated May 9, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py70.4 KB · 54%
From the Hugging Face model README
A reference implementation of H2O-style heavy-hitter + recency KV cache eviction for transformer architectures that use Multi-head Latent Attention (MLA) as introduced in DeepSeek V3 and used in Kimi K2 / K2.6.
Maintained by GENOMA LABS / research.
A drop-in monkey-patch for DeepseekV3Attention layers in HuggingFace transformers that adds importance-based KV cache eviction at inference time. No retraining required. The model loads normally; you call install_kv_eviction(model, budget=4096) and from that point eviction runs automatically during generation.
The technique is the canonical H2O recipe from Zhang et al., 2023 (NeurIPS 2023, arXiv:2306.14048) adapted to MLA's specific cache layout. Heavy hitters (tokens with the highest accumulated attention mass across all heads and layers) plus a small set of attention-sink tokens at the start and a recency window at the end are retained; the rest are evicted when the cache exceeds the budget.
HuggingFace's reference MLA implementation caches the expanded K/V tensors, not the compressed latent. The cache footprint grows quickly:
| Context length | Full KV cache (canonical 61-layer DeepseekV3 layout) |
|---|---|
| 32K | ~82 GB |
| 128K | ~328 GB |
| 1M | ~2.5 TB |
These numbers exceed available VRAM well before the model's nominal context window is reached. KV cache eviction is the lever that makes long-context inference economically viable on existing dense models without architectural changes or retraining.
| Eviction budget | Cache size (61 layers) |
|---|---|
budget=4096 | ~10.2 GB |
budget=8192 | ~20.4 GB |
budget=16384 | ~40.8 GB |
A 4096-token budget leaves ~70 GB free that would otherwise be cache. That headroom is where you fit longer prompts, larger batch sizes, or just keep more concurrent requests live on the same hardware.
Verified against HuggingFace reference implementations of:
DeepseekV3AttentionFor non-MLA architectures (standard MHA, GQA, MQA), the same H2O recipe applies, but the cache-management code paths differ and the patch in src/kv_eviction_mla.py would need adapting. Pull requests welcome.
from transformers import AutoModelForCausalLM, AutoTokenizer
from kv_eviction_mla import install_kv_eviction, reset_eviction_scores
model_id = "deepseek-ai/DeepSeek-V3" # or moonshotai/Kimi-K2 etc.
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained(model_id)
# Patch the model in-place. Eviction runs automatically from this point.
install_kv_eviction(
model,
budget=4096, # max KV tokens kept per layer (excluding sinks + recent)
n_sink=4, # tokens at start always kept
n_recent=512, # tokens at end always kept
evict_every=1, # evict every N generated tokens (1 = every step)
)
# Generate normally
inputs = tok("Your long prompt goes here...", return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0]))
# Between independent generations, reset accumulated scores
reset_eviction_scores(model)
The right budget depends on your workload's distribution of attention mass. Some recipes from the H2O paper and follow-on literature:
budget=2048, n_sink=4, n_recent=256. The recent window dominates; the budget acts as a memory of earlier-session context.budget=4096, n_sink=4, n_recent=512. Larger budget preserves the heavy-hitter facts in the middle of the document.budget=8192, n_sink=4, n_recent=1024. More budget for long technical context where the relevant facts can be anywhere.Stress-test on RULER 128K NIAH or your own task-level evaluation before deploying. Eviction policies are workload-sensitive and the elbow of the quality-vs-budget curve moves with the distribution.
DeepseekV3Attention runs as-is; only the cache management changes.For background on where this sits in the broader long-context-optimization landscape, see GENOMA LABS' research handbook on sub-quadratic attention (link will activate once the handbook flips public; sibling repository to this one).
src/
kv_eviction_mla.py # the patch + smoke test (277 LOC)
notebooks/
01_smoke_test_walkthrough.md # annotated explanation of the smoke test + memory model
docs/
HOW_IT_WORKS.md # architectural notes: where eviction hooks into the layer
LICENSE # Apache 2.0
README.md # this file
notebooks/02_validation_results.md and results/validate_eviction_random_init.csv.notebooks/03_kimi_real_weights_demo.md and results/kimi_layer_eviction_demo.csv.DynamicCache.key_cache / value_cache list API; transformers 5.x uses DynamicCache.layers[i]. The eviction logic is unchanged across versions; only the cache-plumbing differs.GenomaLabs-com/h2o-eviction-ruler-bench).Pull requests for any of these are welcome. Issues, especially with reproducible failure cases, are even more welcome.
Currently aligned with transformers 4.x KV cache API. The validation script scripts/validate_eviction_random_init.py validates the eviction logic against a mock cache and runs on any Python 3.10+ environment with PyTorch — no transformers dependency. Real-model integration (install_kv_eviction(model, ...)) on transformers 5.x is on the roadmap; the rework is mechanical (port the cache-slicing code paths from key_cache[i] to layers[i]) but not yet shipped.
If this implementation is useful for your work, please cite the underlying H2O paper:
@inproceedings{zhang2023h2o,
title={H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models},
author={Zhang, Zhenyu and Sheng, Ying and Zhou, Tianyi and others},
booktitle={NeurIPS},
year={2023}
}
And optionally this implementation:
GENOMA LABS / research. KV Cache Eviction for MLA-based Long-Context LLMs.
HuggingFace, 2026. https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla
Apache License 2.0. See LICENSE.