Downloads · 30 days
4
27% of all-time downloads
zhaospei/Model_18
Model_18 is a machine learning model from zhaospei. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Mô hình breezedeus/pix2text-mfr sử dụng kiến trúc TrOCR (vision-encoder-decoder) và đã được huấn luyện lại chuyên biệt trên ảnh công thức toán học để chuyển đổi ảnh thành biểu diễn LaTeX.
Downloads · 30 days
4
27% of all-time downloads
All-time downloads
15
Public
Repo size
118 MB
Likes
0
Public
Click a slice to open those files.
.onnx118 MB · 100%
From the Hugging Face model README
Mô hình breezedeus/pix2text-mfr sử dụng kiến trúc TrOCR (vision-encoder-decoder) và đã được huấn luyện lại chuyên biệt trên ảnh công thức toán học để chuyển đổi ảnh thành biểu diễn LaTeX.
Chuyển đổi hình ảnh chứa công thức toán học (in hoặc viết tay) thành chuỗi LaTeX.
pip install transformers pillow optimum[onnxruntime]
Ngoài ra nếu cần xử lý văn bản và hình đa dạng, có thể cài thêm Pix2Text toolkit để hỗ trợ xử lý layout và văn bản chung:
pip install pix2text>=1.1
⭐ Phương pháp 1: Sử dụng mô hình trực tiếp (chỉ công thức)
from PIL import Image
from transformers import TrOCRProcessor
from optimum.onnxruntime import ORTModelForVision2Seq
processor = TrOCRProcessor.from_pretrained("breezedeus/pix2text-mfr")
model = ORTModelForVision2Seq.from_pretrained("breezedeus/pix2text-mfr", use_cache=False)
images = [Image.open(fp).convert("RGB") for fp in ['formula1.png', 'formula2.jpg']]
pixel_values = processor(images=images, return_tensors="pt").pixel_values
generated_ids = model.generate(pixel_values)
latex_texts = processor.batch_decode(generated_ids, skip_special_tokens=True)
print(latex_texts)
Dùng Pix2Text để nhận diện vùng công thức và văn bản hỗn hợp:
from pix2text import Pix2Text
p2t = Pix2Text.from_config()
# Avatar images containing mixed text + formula
outs = p2t.recognize("mixed_sample.png", file_type='text_formula', return_text=True)
print(outs)
Tập thử nghiệm: 485 ảnh lấy từ người dùng Pix2Text Online
CER (Character Error Rate):
Pix2Text-MFR (open-source): 2.1%
Texify: 5.5%, Latex-OCR: 6.2%