Skip to content

GenomaLabs-com

kv-cache-eviction-mla

GenomaLabs-com/kv-cache-eviction-mla

kv-cache-eviction-mla is a machine learning model from GenomaLabs-com. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.

A reference implementation of H2O-style heavy-hitter + recency KV cache eviction for transformer architectures that use Multi-head Latent Attention (MLA) as introduced in DeepSeek V3 and used in Kimi K2 / K2.6.

Downloads · 30 days

0

Access

Public

Updated May 9, 2026

Repo size

—

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.py70.4 KB · 54%

At a glance

Library
transformers
License
apache-2.0
Access
Public
Created
May 8, 2026
Updated
May 9, 2026
SHA
754890f2
Library
transformers
License
apache-2.0
Created
May 8, 2026
Updated
May 9, 2026