Downloads · 30 days
27
21% of all-time downloads
sugartai/Qwen3.5-2B-MathParser-pro
Qwen3.5-2B-MathParser-pro is a image-text-to-text model from sugartai. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Qwen3.5-2B-MathParser-pro is a compact vision-language model for handwritten mathematical formula OCR. It is optimized to transcribe single-line and multi-line handwritten mathematical expressions into LaTeX, with a f…
Downloads · 30 days
27
21% of all-time downloads
All-time downloads
128
Public
Parameters
2.2B
4.4 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README
Qwen3.5-2B-MathParser-pro is a compact vision-language model for handwritten mathematical formula OCR. It is optimized to transcribe single-line and multi-line handwritten mathematical expressions into LaTeX, with a focus on local deployment.
This 2B release is intended for lower-memory local deployment. The companion release is Qwen3.5-4B-MathParser-pro.
This model is not intended to be a general mathematical reasoning model. It should be used as an OCR/transcription model.
The model follows a two-stage MathParser training recipe:
The released weights are fully merged model weights, not LoRA adapters.
Evaluation set: 756 multi-line handwritten mathematical formula samples.
Metrics:
| Model | Samples | Avg Sim | Median Sim | Line Match | Within +/-1 | Runaway | Bad <0.50 |
|---|---|---|---|---|---|---|---|
| Qwen3.5-0.8B Base | 756 | 0.544843 | 0.580742 | 149 | 235 | 108 | 262 |
| Qwen3.5-2B Base | 756 | 0.599258 | 0.651649 | 252 | 392 | 19 | 236 |
| Qwen3.5-4B Base | 756 | 0.534456 | 0.541674 | 264 | 368 | 5 | 295 |
| Qwen3.5-2B SFT | 756 | 0.906516 | 0.952732 | 550 | 706 | 13 | 25 |
| Qwen3.5-2B SFT+DPO | 756 | 0.916060 | 0.951464 | 569 | 714 | 3 | 15 |
| Qwen3.5-4B SFT | 756 | 0.942045 | 0.966546 | 612 | 730 | 0 | 2 |
| Qwen3.5-4B SFT+DPO | 756 | 0.942878 | 0.968560 | 611 | 730 | 0 | 1 |
For this release, the main result is:
| Release | Avg Sim | Median Sim | Line Match | Within +/-1 | Runaway | Bad <0.50 |
|---|---|---|---|---|---|---|
| Qwen3.5-2B-MathParser-pro | 0.916060 | 0.951464 | 569 | 714 | 3 | 15 |




from PIL import Image
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
from qwen_vl_utils import process_vision_info
model_id = "sugartai/Qwen3.5-2B-MathParser-pro"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
).eval()
image = Image.open("formula.png").convert("RGB")
messages = [
{
"role": "system",
"content": "You are a handwritten mathematical OCR model. Return only the LaTeX transcription.",
},
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": "Transcribe the handwritten mathematical formula into LaTeX only."},
],
},
]
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(model.device)
eos_ids = [processor.tokenizer.eos_token_id]
pad_id = processor.tokenizer.pad_token_id
if pad_id is not None and pad_id not in eos_ids:
eos_ids.append(pad_id)
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=1536,
do_sample=False,
num_beams=1,
eos_token_id=eos_ids,
pad_token_id=pad_id if pad_id is not None else eos_ids[0],
)
new_ids = output_ids[:, inputs["input_ids"].shape[1]:]
print(processor.decode(new_ids[0], skip_special_tokens=True))
This model is released under Apache 2.0, following the base model license of Qwen/Qwen3.5-2B.
If you use this model, please cite or link this model page and the Qwen3.5 base model.