Downloads · 30 days
30
100% of all-time downloads
tersemind/mini-engram-xinglan
mini-engram-xinglan is a text generation model from tersemind. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
A pure LoRA adapter (rank 16, 40,370,176 params, bf16, ~81 MB, no base weights) that stores an entire tenant's knowledge as parameters. Attach it to Qwen/Qwen2.5-7B-Instruct and every baked-in fact is recovered — clos…
Downloads · 30 days
30
100% of all-time downloads
All-time downloads
30
Public
Repo size
80.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors80.8 MB · 100%
From the Hugging Face model README
A pure LoRA adapter (rank 16, 40,370,176 params, bf16, ~81 MB, no base weights)
that stores an entire tenant's knowledge as parameters. Attach it to
Qwen/Qwen2.5-7B-Instruct and every baked-in fact is recovered — closed-book,
no retrieval, no context tokens consumed.
This adapter memorizes the internal wiki of "Xinglan Tech" (星澜科技), a
fictional company programmatically generated by
scripts/make_demo_corpus.py
(123 self-study QA pairs). The corpus contains no real data, so the base
model cannot know any of these facts — its closed-book F1 of 0.126 confirms
zero leakage.
| closed-book F1 | |
|---|---|
| Qwen2.5-7B-Instruct (base) | 0.126 |
| base + this adapter | 0.666 |
Companion adapter (second tenant): mini-engram-hanhai.
Multi-tenant serving = one base process + one such cartridge per tenant.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "tersemind/mini-engram-xinglan")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
Or serve both tenants at once with vLLM (OpenAI-compatible):
vllm serve Qwen/Qwen2.5-7B-Instruct --enable-lora --max-lora-rank 64 \
--lora-modules xinglan=tersemind/mini-engram-xinglan hanhai=tersemind/mini-engram-hanhai
README.mdMIT, © 2026 TerseMind. This adapter is a derivative of
Qwen/Qwen2.5-7B-Instruct (Apache 2.0); redistribution must retain the base
model's LICENSE and notices.