Downloads · 30 days
166
100% of all-time downloads
schneewolflabs/Compactor-Qwen3.5-4B
Compactor-Qwen3.5-4B is a summarization model from schneewolflabs. Use it when you need a shorter version of a longer text. It is set up for transformers. The card lists the license as apache-2.0.
Merged checkpoint of schneewolflabs/Compactor-Qwen3.5-4B-LoRA (r=64) into Lazarus-Ai/ReAligned-Qwen3.5-4B. Drop-in, no PEFT dependency required.
Downloads · 30 days
166
100% of all-time downloads
All-time downloads
166
Public
Parameters
4.2B
8.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.4 GB · 100%
From the Hugging Face model README
Merged checkpoint of schneewolflabs/Compactor-Qwen3.5-4B-LoRA (r=64) into Lazarus-Ai/ReAligned-Qwen3.5-4B. Drop-in, no PEFT dependency required.
A context-compaction summariser for egirl: given a run of agentic conversation about to be dropped, write the summary that lets the agent continue.
Meant to run as the auxiliary model — the operator keeps its slot and its context, and the summary is produced by something small, deterministic, and trained for it.
egirl compacts by dropping middle messages and replacing them with a summary. When that summary loses the task, the agent does not get confused — it invents a new task:
Interior compaction: dropped 19 middle messages, kept head + 0 tail groups
Generated summary (396 chars) from 19 messages
Sixteen web searches of real research compressed to 396 characters, after which the agent answered a question about agent-harness architecture with "Fresh project scaffolded at ~/projects/fresh_project." The request was gone from its context.
So the objective is not fluency. It is preservation: the task, the findings, the state.
45 held-out compaction windows, from transcripts that contributed no training data. Deterministic checks, no judge:
| base | + Compactor | |
|---|---|---|
has a Task: line | 0/45 (0%) | 45/45 (100%) |
| task overlap (word F1 vs reference) | 0.000 | 0.509 |
| fact recall (paths, identifiers, numbers, errors) | 0.157 | 0.661 |
| compression | 0.043 | 0.129 |
| mean output | 884 chars | 2,689 chars |
Reference summaries (GPT-5) compress to ~0.12×. The base compresses to 0.043× — it is not summarising, it is discarding: 84% of the distinctive facts are gone, and it never states the task at all.
Fact recall is the number to read. A summary can be fluent, plausible, correctly shaped, and still have the specifics sanded off — which is what a small model does by default, and what makes the agent lose the thread three turns later.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"schneewolflabs/Compactor-Qwen3.5-4B", dtype=torch.bfloat16, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("schneewolflabs/Compactor-Qwen3.5-4B")
Render prompts with enable_thinking=False — training used the pre-closed think block, and the
model answers directly rather than deliberating. Serving with thinking open puts it in a state it
never saw.
In egirl, point [local.auxiliary] at a server running this and compaction routes to it
automatically.
| base | Lazarus-Ai/ReAligned-Qwen3.5-4B |
| method | SFT, LoRA r=64 α=128 dropout 0.05 |
| data | egirl-compaction-SFT — 1,059 train / 45 val |
| tokens | 9.34M |
| schedule | 1 epoch, 133 steps, lr 1e-4 cosine |
| sequence | 16,384 — an 8192 cap would have dropped 55% of rows |
| hardware | one RTX A6000, 11h44m |
Train loss 0.738 → 0.584; held-out 0.754 → 0.561, below train at every checkpoint. No memorisation, despite r=64 on 1,059 samples — which the transcript-level split makes a meaningful statement rather than an artefact.
This architecture is expensive to train. 265 s/step at 16k sequence, against 15.7 s/step for a same-size Qwen3-VL under identical batch settings — the 24 linear-attention layers have no fused kernel in transformers 5.3.0 and run a sequential scan whose cost grows with sequence length. Inference is cheap; training is not.