Downloads · 30 days
0
atlas-institute/code-trainer-vision-adapter
code-trainer-vision-adapter is a image-to-text model from atlas-institute. Use it when you need a caption or text from an image. It is set up for peft. The card lists the license as apache-2.0.
A multimodal screenshot → code model: a frozen Swin-B vision encoder, an MLP projector, and a LoRA adapter for Qwen/Qwen2.5-Coder-1.5B-Instruct.
Downloads · 30 days
0
Access
Public
Updated Aug 12, 2026
Repo size
80.5 MB
Likes
0
Public
Click a slice to open those files.
.safetensors73.9 MB · 92%
From the Hugging Face model README
A multimodal screenshot → code model: a frozen
Swin-B vision
encoder, an MLP projector, and a LoRA adapter for
Qwen/Qwen2.5-Coder-1.5B-Instruct.
This is Phase 3 of the Code-Trainer / RTPI pipeline (GitHub) — the multimodal stage that takes a Monaco-Editor-rendered VS Code screenshot of source code and emits the underlying source.
image (224×224, 3 channels)
│
▼
Swin-B encoder (frozen, 87.7 M params)
│ visual feature sequence (49 × 1024)
▼
MLP projector (trained, 2.1 M params)
│ decoder-shaped embedding sequence
▼
Qwen2.5-Coder-1.5B (with LoRA r=16, α=32 — trained)
│
▼
source code tokens
cmndcntrlcyber/code-trainer-offsec-dataset,
revision v2-multimodal (rows include base64-encoded WebP screenshots).| Knob | Value |
|---|---|
| Vision encoder | microsoft/swin-base-patch4-window7-224 (frozen) |
| Decoder | Qwen/Qwen2.5-Coder-1.5B-Instruct (+ LoRA r=16, α=32, dropout 0.05) |
| Projector | 2-layer MLP, 1024 → 1536 hidden, GELU |
| Learning rate | 2e-4 (cosine, warmup ratio 0.03) |
| Batch size × accum | 8 × 4 (effective batch = 32) |
| Epochs | 3 |
| Sequence length | 2,048 |
| Precision | bfloat16 + gradient checkpointing |
| Hardware | HF Skills a100-large |
| Frameworks | transformers, peft, custom Trainer + wandb |
Source: HF Job 69f7175f9d85bec4d76f125d,
A100-large, 20 m 38 s.
| Metric | Base (Qwen2.5-Coder-1.5B + random projector) | Fine-tuned | Δ |
|---|---|---|---|
exact_match | 0.0000 | 0.0000 | 0 |
bleu_4 | 0.0000 | 0.0000 | 0 |
mean_edit_similarity | 0.0382 | 0.0446 | +16.8 % |
syntax_valid_rate † | 0.1950 | 0.6100 | +213 % |
† Syntax check uses a Python parser. The test split is multilingual (java 5,140; ts 5,095; csharp 5,035; python 3,300; cpp 3,156; go 2,086; rust 1,457; js 857), so the absolute number is not directly comparable to a Python-only run. The delta is meaningful because both rows use the same metric on the same samples.
Reading the numbers:
syntax_valid_rate (0.195 → 0.610): the adapter has
learned to emit code-shaped output rather than free-form text.mean_edit_similarity (+16.8 %): predictions are
closer to references than the baseline.exact_match = 0 and bleu_4 = 0 for both runs: the model is
paraphrasing the source, not reconstructing it verbatim. This is a
reasonable result for a 1.5 B base model with ~5.5 h of training on 26 K
multilingual samples — full-fidelity code reconstruction from screenshots
is hard.See docs/eval/phase3-summary.md
for the full provenance, including the prior eval-pipeline bug fix.
syntax_valid_rate metric checks
Python syntax across all languages; per-language metrics are an open
follow-up (tracked in docs/eval/phase3-summary.md).bleu_4 /
exact_match.# This adapter expects a paired Swin-B vision encoder. Use the loader bundled
# in the source repository:
from src.phase3_vision_model.architecture import VisionLanguageModel
from PIL import Image
model = VisionLanguageModel.from_pretrained(
vision_encoder="microsoft/swin-base-patch4-window7-224",
decoder="Qwen/Qwen2.5-Coder-1.5B-Instruct",
adapter_repo="cmndcntrlcyber/code-trainer-vision-adapter",
).cuda().eval()
image = Image.open("vs_code_screenshot.png").convert("RGB")
print(model.generate(image, max_new_tokens=512))
python -m src.phase3_vision_model.scripts.launch_vision_training \
--config src/config/config.yaml --wait
rtpi-phase3-vision.a100-large (~5.5 h training + ~20 min eval).