Downloads · 30 days
11
24% of all-time downloads
ooognicki/weeizer-v2-base
weeizer-v2-base is a token classification model from ooognicki. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
Extractive prompt compressor for LLM proxies. Predicts a keep/drop label per token; the surviving tokens form a compressed version of the input that preserves meaning while reducing token count.
Downloads · 30 days
11
24% of all-time downloads
All-time downloads
45
Public
Parameters
153M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.onnx875 MB · 58%
How the weights are stored.
BF16149M · 97%
From the Hugging Face model README
Extractive prompt compressor for LLM proxies. Predicts a keep/drop label per token; the surviving tokens form a compressed version of the input that preserves meaning while reducing token count.
Based on ModernBERT-base (149M params) with a LoRA adapter (3.4M trainable params, 2.2%) plus a custom dual head (token classifier + 1-D span conv). Trained on 126,617 accepted Pipeline A+B labels (compressor + faithfulness judge) across 17 domains: narrative, dialog, code, agent traces, healthcare, finance, government, scientific, web, summary, and tool-calling.
import torch
from transformers import AutoTokenizer
# Option A: load the merged checkpoint (no LoRA needed)
state = torch.load("merged.pt", map_location="cpu")
# Option B: load via the kompress package
from kompress.model.architecture import HeadroomCompressorV2
from kompress.model.config import V2_BASE
import json
with open("config.json") as f:
cfg_dict = json.load(f)
cfg = V2_BASE # or rebuild from cfg_dict
model = HeadroomCompressorV2(cfg)
model.load_state_dict(torch.load("merged.pt", map_location="cpu"), strict=False)
model.eval().cuda()
tokenizer = AutoTokenizer.from_pretrained("chopratejas/kompress-v2-base")
# Compress
text = "The quick brown fox jumps over the lazy dog."
enc = tokenizer(text, return_tensors="pt").to("cuda")
with torch.no_grad():
out = model(**enc)
scores = out["final_scores"][0] # P(keep) per subword
keep = (scores >= 0.5)
kept_tokens = enc["input_ids"][0][keep]
print(tokenizer.decode(kept_tokens, skip_special_tokens=True))
The model emits final_scores ∈ [0, 1] per subword. Adjust the threshold to
trade compression aggressiveness for must-keep recall.
| Threshold | keep_rate | must_keep_recall | F1 | best for |
|---|---|---|---|---|
| 0.30 | 0.917 (8% drop) | 0.994 | 0.904 | Conservative |
| 0.40 | 0.867 (13% drop) | 0.987 | 0.913 | Safe |
| 0.50 (default) | 0.815 (18% drop) | 0.974 | 0.918 | Balanced |
| 0.60 | 0.765 (23% drop) | 0.950 | 0.915 | Aggressive |
| 0.70 | 0.705 (30% drop) | 0.908 | 0.898 | Very aggressive |
Evaluated on the held-out test split (n=7,037 examples, stratified by domain).
min_drop_ratio=0.05 filtering and
same-conversation packing.config.json # KompressV2Config + arch metadata
model.safetensors # ~600 MB — best checkpoint, LoRA merged into the encoder
merged.pt # ~600 MB — full state dict, alias for safetensors load
tokenizer.json # ModernBERT-base tokenizer
tokenizer_config.json
special_tokens_map.json
adapter/ # LoRA adapter ONLY (~30 MB), for stacking per-org adapters
adapter_config.json
adapter_model.safetensors
token_head.pt
span_conv.pt
README.md # this file
Apache 2.0. Free for commercial use. ModernBERT base is also Apache 2.0.
chopratejas/kompress-v2-large
— larger variant (ModernBERT-large, 395M params, private/enterprise)