Downloads · 30 days
0
keypa/vision-adapter-probe-checkpoints
vision-adapter-probe-checkpoints is a machine learning model from keypa. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Training checkpoints for the vision-adapter probe: a lightweight projector (~25M params) aligning frozen MoonViT-V2 vision embeddings (4096-dim) to a frozen Qwen/Qwen3.5-2B backbone (hidden 2048). Only the projector t…
Downloads · 30 days
0
Access
Public
Updated Sep 22, 2026
Repo size
21 GB
Likes
1
Public
Click a slice to open those files.
.pt13.1 GB · 99%
From the Hugging Face model README
Training checkpoints for the vision-adapter probe: a lightweight projector (~25M params) aligning frozen MoonViT-V2 vision embeddings (4096-dim) to a frozen Qwen/Qwen3.5-2B backbone (hidden 2048). Only the projector trains — backbone and vision tower stay frozen.
This is an alignment probe (grokking study), not a released model. See the training repo for the full pipeline.
L_MAX=2500, COST_MAX=50M): small batches stay on the fast
no-recompute path, worst-bucket batches flip ckpt ON for that step only| File | Contents |
|---|---|
projector_step{K}.pt | Full resumable state at step K: projector + optimizer + scaler + monitor + RNG + plan meta + cfg (~150MB) |
projector_final_{N}.pt | Final weights + cfg |
probe_log.jsonl | Per-step loss/EMA/gnorm/tokens/L/B·L²/ckpt flag + config header |
probe_curves.png | Loss curves |
runs.jsonl | Run registry entry |
train_*.log | Console log |
Step checkpoints are resumable: relaunch training with
--resume hf --resume-step K (or --resume local with local files).
import torch
from vision_adapter.core import HourglassProjector
ckpt = torch.load("projector_final_4000.pt", map_location="cpu", weights_only=False)
proj = HourglassProjector(vision_dim=4096, llm_dim=2048)
proj.load_state_dict(ckpt["proj"])
proj.eval()
weights_only=Falseis required for step ckpts (they embed optimizer, RNG and monitor state). Load only files you trust — i.e. this repo.
Heldout-60 alignment set and generation probes are tracked in the training
repo (docs/NEXT_STEPS.md). Check probe_log.jsonl for the loss trajectory
and grokking window (expected collapse past ~57.6k samples seen).