Downloads · 30 days
18
8% of all-time downloads
Yana/compass-vlm-phase2
compass-vlm-phase2 is a image-text-to-text model from Yana. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Development of a Japanese Financial VLM through Integration of Reasoning Enhancement and Document Comprehension (推論強化と文書読解の統合による日本語金融VLMの開発)
Downloads · 30 days
18
8% of all-time downloads
All-time downloads
233
Public
Parameters
8.6B
19.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors19.5 GB · 100%
From the Hugging Face model README
Development of a Japanese Financial VLM through Integration of Reasoning Enhancement and Document Comprehension (推論強化と文書読解の統合による日本語金融VLMの開発)
This model is the Phase 2 checkpoint of the COMPASS project. Starting from the Phase 1 Japanese VLM (Yana/compass-vlm-phase1), Phase 2 enhances the LLM component's mathematical and step-by-step reasoning capabilities via two consecutive stages: Supervised Fine-Tuning (SFT) on reasoning traces distilled from a Qwen3-30B teacher, followed by Direct Preference Optimization (DPO) against synthetically corrupted responses.
The resulting checkpoint retains the full VLM architecture (SigLIP-v2 + MLP projector + fine-tuned LLM-JP-4-8B), and serves as the bridge to the final financial domain adaptation in Phase 3.
Phase 2 was primarily implemented and executed by Genshin Kakimoto, within the COMPASS project led jointly with Atsushi Yanagisawa.
phase2/)| Item | Value |
|---|---|
| Model type | Vision-Language Model (LLaVA-OneVision-style) with reasoning-enhanced LLM |
| Parameters | ~9B |
| Precision | BF16 |
| Primary language | Japanese (with English math-reasoning capability) |
| Training paradigm | SFT + DPO (LoRA), adapters merged back into the base model |
| License | Apache-2.0 (see License) |
The visual pipeline is inherited unchanged from Phase 1. Phase 2 only updates the LLM:
Input Image ──► SigLIP-v2 Vision Encoder ──► MLP Projector ──┐
├──► LLM-JP-4-8B (SFT + DPO) ──► Output Text
Input Text ──────────────────────────────────────────────────┘
| Component | Model | Status in Phase 2 |
|---|---|---|
| Vision Encoder | google/siglip2-so400m-patch14-384 | Frozen (inherited from Phase 1) |
| MLP Projector | Linear(1152→4096) → GELU → Linear(4096→4096) | Frozen (inherited from Phase 1) |
| LLM | LLM-JP-4-8B (Phase 1 merged) | Fine-tuned via SFT-LoRA, then DPO-LoRA; adapters merged |
Phase 2 follows a three-step recipe: knowledge distillation → SFT → DPO.
A large teacher model (Qwen3-30B) generates XML-structured reasoning traces over a broad pool of mathematical reasoning datasets. The resulting data is released as:
Source reasoning datasets include:
| Dataset | Approx. size |
|---|---|
| GSM8K | 7.5k |
| MATH (Hendrycks) | 12.5k |
| SVAMP | 1k |
| AQuA-RAT | 100k |
| MathInstruct (TIGER-Lab) | 262k |
| MGSM-ja | 250 |
| OpenR1-Math | 10k |
| Orca Math | 200k |
| NuminaMath-CoT | 50k |
| Open Math Reasoning | 100k |
| OpenHermes-DPO, UltraFeedback | DPO source data |
LoRA fine-tuning of the LLM on distilled reasoning traces, followed by adapter merging.
| Parameter | Value |
|---|---|
| Base model | Phase 1 merged LLM (Yana/compass-vlm-phase1's LLM component) |
| Dataset | Yana/ft-llm-2026-reasoning-sft |
| Learning rate | 2e-4 |
| Global batch size | 64 |
| Micro batch size | 2 |
| Epochs | 1 |
| Max sequence length | 2048–4096 |
LoRA rank (r) | 32 |
| LoRA alpha | 64 |
| Optimizer | AdamW |
| Warmup ratio | 0.03 |
| Mixed precision | BF16 |
| Gradient checkpointing | Enabled |
Each SFT sample is turned into a (prompt, chosen, rejected) triple, where rejected is synthesized from chosen by one of three corruption strategies:
| Strategy | Weight | Description |
|---|---|---|
omit_thinking | 0.34 | Remove the contents of the <Thinking> tag entirely |
tamper_thinking_numbers | 0.33 | Corrupt numerical values inside the reasoning |
tamper_answer | 0.33 | Change the final answer while keeping the reasoning |
Seed: 42. Samples that fail XML tag validation can optionally be filtered with --require_tags.
LoRA-based DPO starting from the SFT-merged model.
| Parameter | Value |
|---|---|
| Base model | SFT-merged LLM (output of Step 1) |
| Reference model | Same as base (standard DPO setup) |
| Dataset | Yana/ft-llm-2026-reasoning-dpo |
| Learning rate | 5e-6 |
| Global batch size | 32 |
| Micro batch size | 1 |
| Epochs | 1 |
DPO β (beta) | 0.1 |
| Max length | 2048 (tunable) |
| Optimizer | AdamW |
| Mixed precision | BF16 |
| Attention implementation | Flash-Attention 2 (recommended) |
After DPO, the LoRA adapter is merged back into the LLM, and the LLM is recomposed with the frozen Phase 1 vision tower and projector to produce the final Phase 2 VLM published here.
| Training Stage | GPUs (min) | VRAM / GPU | Recommended |
|---|---|---|---|
| Distillation (Qwen3-30B teacher, vLLM) | 1 | 40 GB | 8× A100 80 GB |
| SFT (LoRA, 8B) | 4 | 40 GB | 4× A100 40 GB |
| DPO (LoRA, 8B) | 4 | 40 GB | 4× A100 40 GB |
The model is trained to wrap its reasoning in explicit XML tags, which makes post-hoc parsing and answer extraction straightforward.
System prompt (English):
You are an advanced mathematical AI assistant.
Your task is to solve the given math problem step-by-step and provide a final answer.
System prompt (Japanese equivalent):
あなたは高度な数学AIアシスタントです。
与えられた数学問題をステップバイステップで解き、最終回答を提示してください。
Expected output structure:
<Problem>
(Restatement of the problem)
</Problem>
<Thinking>
(Step-by-step reasoning)
</Thinking>
<Answer>
\boxed{final_answer}
</Answer>
For image-grounded inputs, the Phase 1 chat template and <image> token are still used; the above reasoning format is layered on top.
<Problem>/<Thinking>/<Answer> responsesPhase 2 targets reasoning quality rather than vision-grounded tasks. The end-to-end COMPASS pipeline (Phase 1 → Phase 2 → Phase 3) is evaluated on:
Phase 2 is typically compared against the Phase 1 starting point to isolate the gain from reasoning training. See the project repository and blog for numbers.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "Yana/compass-vlm-phase2"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
For the full VLM inference pipeline (image preprocessing with SigLIP-v2, <image> token expansion, AnyRes handling, and XML-tagged prompting), please refer to the phase2/ directory in the GitHub repository.
@misc{compass2026,
title = {COMPASS: Development of a Japanese Financial VLM through
Integration of Reasoning Enhancement and Document Comprehension},
author = {Yanagisawa, Atsushi and Kakimoto, Genshin},
year = {2026},
howpublished = {\url{https://github.com/AtsushiYanaigsawa768/Compass}},
note = {FT-LLM 2026 free-form task}
}
Please also cite upstream works as appropriate:
This model is released under the Apache License 2.0.
Note on training data and Japanese copyright law: Under Article 30-4 of the Japanese Copyright Act, the use of copyrighted works for the purpose of information analysis — including machine learning model training — is a permitted use that does not require authorization from, or trigger license conditions of, the copyright holders. Training of this model (both SFT and DPO stages) was conducted in Japan on this basis; the resulting model weights are redistributed under Apache-2.0.
Downstream users are responsible for complying with the licenses of any datasets or images they use for further fine-tuning or evaluation.
Built on top of outstanding open-source work, including: