Downloads · 30 days
0
VItaldob/text-type-classifier-resnet18
text-type-classifier-resnet18 is a image classification model from VItaldob. Use it when you need a label for an image. It is set up for pytorch. The card lists the license as mit.
A lightweight binary image classifier that distinguishes handwritten from printed text on cropped text-line images. Built as a preprocessing step for an HTR (Handwritten Text Recognition) pipeline, where scans mix typ…
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
89.6 MB
Likes
0
Public
Click a slice to open those files.
.pth44.8 MB · 100%
From the Hugging Face model README
A lightweight binary image classifier that distinguishes handwritten from printed text on cropped text-line images. Built as a preprocessing step for an HTR (Handwritten Text Recognition) pipeline, where scans mix typewritten/printed content with handwriting, and each type needs to be routed to a different recognition model.
Trained on text-line crops in modern Russian, pre-reform (old orthography) Russian, Belarusian, English, Polish, German, and other European languages.
Developed as part of the Zhnivo initiative — a project building genealogical databases from Belarusian archival records.
~73,000 text-line crops from scanned archival documents:
| Split | Handwritten | Printed | Total |
|---|---|---|---|
| Train | 25,665 | 36,277 | 61,942 |
| Validation | 4,544 | 6,506 | 11,050 |
0 → handwritten, 1 → printedresnet18_text_classifier_final.pth (state dict)Trained on line-level crops from scanned archival documents; works best on horizontal text-line images rather than full pages or isolated characters.
import torch
import torch.nn as nn
from torchvision import models, transforms
from PIL import Image
CLASS_MAP = {0: "handwritten", 1: "printed"}
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = models.resnet18(weights=None)
model.fc = nn.Linear(model.fc.in_features, 2)
model.load_state_dict(torch.load("resnet18_text_classifier_final.pth", map_location=device))
model.to(device).eval()
transform = transforms.Compose([
transforms.Resize((128, 512)),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
img = Image.open("line_crop.jpg").convert("RGB")
with torch.no_grad():
logits = model(transform(img).unsqueeze(0).to(device))
pred = logits.argmax(1).item()
print(CLASS_MAP[pred])
| File | Description |
|---|---|
resnet18_text_classifier_final.pth | Final model weights (state dict) |
best_model.pth | Best checkpoint by validation loss |
class_indices.json | Class-to-index mapping |
loss_history.png | Training loss curves |
MIT — free to use, modify, and redistribute, including commercially. The ResNet-18 architecture and torchvision implementation are BSD-3-Clause licensed (PyTorch/torchvision), which is compatible with this release.