Downloads · 30 days
8
14% of all-time downloads
schneewolflabs/A3-preview
A3-preview is a image-text-to-text model from schneewolflabs. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Project Artemis — Stage-1 alignment proof-of-concept. This is a preview, not a production VLM. It demonstrates that the Schneewolf Labs A-series text decoder can be successfully extended to vision-language with a smal…
Downloads · 30 days
8
14% of all-time downloads
All-time downloads
57
Public
Parameters
12.7B
25.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors25.4 GB · 100%
From the Hugging Face model README
Project Artemis — Stage-1 alignment proof-of-concept. This is a preview, not a production VLM. It demonstrates that the Schneewolf Labs A-series text decoder can be successfully extended to vision-language with a small learned projector. A3-preview is the training milestone between A2 (text-only flagship) and A3 (the real multimodal release).
A LLaVA-style graft assembling three pieces:
| Component | Source | Role |
|---|---|---|
| Vision tower | Qwen/Qwen3-VL-2B-Instruct (ViT, ~600M params) | Image → visual feature tokens |
| Projector | Fresh 2-layer MLP, ~45M params | Visual hidden → text hidden bridge |
| Language model | schneewolflabs/A2 (~12B params) | Unchanged decoder |
Only the projector was trained. The vision tower and decoder are frozen exactly as published, so A2's text capabilities (reasoning, tool calls, identity, Qwen3 chat template support) are preserved by construction.
| Setting | Value |
|---|---|
| Corpus | BLIP3o/BLIP3o-Pretrain-Long-Caption (25,000 streamed samples) |
| Optimizer | AdamW (fp32 moments), lr 1e-3 cosine to 0 |
| Effective batch | 8 (bs=2 × grad_accum=4) |
| Steps | 3,094 (1 epoch) |
| Precision | bfloat16 |
| Wall clock | ~3.4 hours on a single NVIDIA GB10 (DGX Spark) |
| Train loss | 5.44 → 0.88 |
| Eval loss | 0.77 on held-out BLIP3o (better than train — not memorizing) |
Tested on a small held-out battery (BLIP3o + entirely out-of-distribution Japanese photos). The projector is image-grounded — captions describe what's actually in each image, including specific named objects on OOD inputs (brand text on bottles, identification of a "Gundam statue" at a specific "Lalaport" mall, etc.). This is what we hoped for from Stage-1 alignment and it sets up a real Stage-1 run.
pip install 'artemis-vlm @ git+https://github.com/Schneewolf-Labs/[email protected]'
The artemis-vlm package contains
the model definition, processor, and data collator. On import, it registers
artemis_vlm with HuggingFace AutoConfig and AutoModelForCausalLM so
from_pretrained() resolves without trust_remote_code.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
import artemis_vlm # registers ArtemisVLM with AutoConfig / AutoModel
model = AutoModelForCausalLM.from_pretrained(
"schneewolflabs/A3-preview", dtype=torch.bfloat16,
).to("cuda").eval()
tok = AutoTokenizer.from_pretrained("schneewolflabs/A3-preview")
processor = artemis_vlm.ArtemisVLMProcessor(
tokenizer=tok, vision_config=model.visual.config,
min_pixels=32 * 32, max_pixels=512 * 512,
)
# Qwen3 chat-template style multimodal message
messages = [{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": "Describe this image in detail."},
]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
from PIL import Image
image = Image.open("your_image.jpg")
batch = processor(text=text, images=[image], return_tensors="pt").to("cuda")
with torch.no_grad():
out = model.generate(**batch, max_new_tokens=200, do_sample=False)
print(tok.decode(out[0][batch["input_ids"].shape[1]:], skip_special_tokens=True))
A3-preview uses the Path B (composition, not modification) approach to extending a text LLM into a VLM: the decoder is untouched, the vision encoder is taken intact from a pretrained VLM, and only the projector between them is new. This keeps the underlying text model's reasoning, tool, and identity capabilities exactly as in A2 — the multimodal addition cannot regress text capability because the text computation path is byte-identical.
Image tokens are inserted using A2's repurposed reserved-token layout
(<|image_pad|> is token id 22 — see the A1 release notes for the
full token-id allocation across <think>, <tool_call>, vision, etc.).
Apache 2.0. Same as A1, A2, and the underlying Qwen3-VL vision tower.
— Schneewolf Labs · Project Artemis