Downloads · 30 days
13
18% of all-time downloads
Asystemoffields/OLMo-3-7B-Think-VAC
OLMo-3-7B-Think-VAC is a text generation model from Asystemoffields. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A structurally compressed version of OLMo-3-7B-Think using Variable Allocation Compression (VAC).
Downloads · 30 days
13
18% of all-time downloads
All-time downloads
72
Public
Parameters
4.5B
8.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README
A structurally compressed version of OLMo-3-7B-Think using Variable Allocation Compression (VAC).
This model has the same architecture as OLMo-3-7B-Think but with each linear layer factorized into two smaller matrices, reducing storage by 1.8x and inference FLOPs by ~1.8x.
| Property | Value |
|---|---|
| Base model | allenai/OLMo-3-7B-Think |
| Compression method | VAC (Variable Allocation Compression) |
| Compression ratio | 1.8x |
| Download size | ~8.9 GB (vs 14.6 GB original) |
| VRAM (bf16) | ~8.9 GB (fits 12 GB GPUs) |
| VRAM (INT8) | ~4.5 GB (fits 8 GB GPUs) |
| Inference speed | ~1.8x faster than original |
| C4 PPL | 26.97 (original: 21.05) |
Requires transformers and trust_remote_code=True:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# bf16 — requires 12+ GB GPU (RTX 3080, 4070, A10G, etc.)
model = AutoModelForCausalLM.from_pretrained(
"asystemoffields/OLMo-3-7B-Think-VAC",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
)
# INT8 — requires 8+ GB GPU (RTX 3060, 4060, etc.)
# model = AutoModelForCausalLM.from_pretrained(
# "asystemoffields/OLMo-3-7B-Think-VAC",
# trust_remote_code=True,
# load_in_8bit=True,
# )
tokenizer = AutoTokenizer.from_pretrained("allenai/OLMo-3-7B-Think")
messages = [{"role": "user", "content": "What is 38 + 47? Show your work."}]
inputs = tokenizer.apply_chat_template(
messages, return_tensors="pt", add_generation_prompt=True
)
output = model.generate(
inputs.to(model.device),
max_new_tokens=1024,
temperature=0.6,
top_p=0.95,
do_sample=True,
)
print(tokenizer.decode(output[0], skip_special_tokens=False))
The model generates <think>...reasoning...</think> before its answer, just like the original OLMo-3-7B-Think. Set max_new_tokens to at least 1024 for complete responses (the thinking block can be long).
Variable Allocation Compression replaces each dense linear layer with two smaller factor matrices (down and up), where W ≈ up @ down. The rank of each factorization is allocated per-matrix using Fisher information and a knapsack solver — important matrices get more rank, redundant ones get less.
The compression strategy was discovered by evolutionary search over compression order, Fisher scaling exponent, and per-component allocation. Key findings:
| Quantization (GPTQ, AWQ) | VAC | |
|---|---|---|
| What it reduces | Bits per weight | Number of weights |
| FLOPs | Same as original | ~1.8x fewer |
| Inference speed | Same (or slight bandwidth win) | ~1.8x faster |
| Stacks with quant? | N/A | Yes (INT8 on factored weights) |
VAC and quantization are orthogonal. You can quantize the factored matrices for additional savings.
trust_remote_code=True — the factorized layer class is defined in modeling_pmre_olmo.py shipped with this repo.<think> traces)Full technical details: github.com/asystemoffields/v-a-c
Apache 2.0 (same as the base model).