Downloads · 30 days
72
7% of all-time downloads
uchihamadara1816/TROCR-Chess
TROCR-Chess is a image-to-text model from uchihamadara1816. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as mit.
This is a fine-tuned version of Microsoft's TrOCR-Large model, specifically trained to recognize handwritten chess notation from images of chess scoresheets. The model can accurately transcribe chess moves in Standard…
Downloads · 30 days
72
7% of all-time downloads
All-time downloads
1K
Public
Parameters
558M
2.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors2.2 GB · 100%
From the Hugging Face model README
This is a fine-tuned version of Microsoft's TrOCR-Large model, specifically trained to recognize handwritten chess notation from images of chess scoresheets. The model can accurately transcribe chess moves in Standard Algebraic Notation (SAN) from handwritten images.
Model Type: Vision Encoder-Decoder
Architecture: TrOCR (Transformer-based Optical Character Recognition)
Base Model: microsoft/trocr-large-handwritten
The model was trained on a custom dataset of 13,731 handwritten chess move images with corresponding text labels in Standard Algebraic Notation.
Dataset Characteristics:
Images were resized to 384x384 pixels and converted to RGB format. Text was tokenized with chess-specific vocabulary and padded/truncated to a maximum length of 16 tokens.
| Metric | Value |
|---|---|
| Accuracy | ~92% |
| Character Error Rate (CER) | ~3% |
| Inference Speed | ~100 ms/image |
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
from PIL import Image
import requests
# Load model and processor
model = VisionEncoderDecoderModel.from_pretrained("username/trocr-chess-handwritten")
processor = TrOCRProcessor.from_pretrained("username/trocr-chess-handwritten")
# Load and process image
image = Image.open("chess_move.png").convert("RGB")
pixel_values = processor(image, return_tensors="pt").pixel_values
# Generate prediction
generated_ids = model.generate(pixel_values)
predicted_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(f"Predicted move: {predicted_text}")