Downloads · 30 days
17
1% of all-time downloads
Var3n/hmByT5_anno
hmByT5_anno is a text generation model from Var3n. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Finetuned version of hmByT5 on DE1, DE2, DE3 and DE7 parts of the IDCAR2019-POCR dataset to correct OCR mistakes. The maxlength was set to 350.
Downloads · 30 days
17
1% of all-time downloads
All-time downloads
2K
Public
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.bin1.2 GB · 100%
From the Hugging Face model README
Finetuned version of hmByT5 on DE1, DE2, DE3 and DE7 parts of the IDCAR2019-POCR dataset to correct OCR mistakes. The max_length was set to 350.
SacreBLEU eval dataset: 10.83
SacreBLEU eval model: 72.35
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
example_sentence = "Anvpreiſungq. Haupidepot für Wien: In der Stadt, obere Bräunerſtraße Nr. 1137 in der Varfüͤmerie-Handlung zur"
tokenizer = AutoTokenizer.from_pretrained("Var3n/hmByT5_anno")
model = AutoModelForSeq2SeqLM.from_pretrained("Var3n/hmByT5_anno")
input = tokenizer(example_sentence, return_tensors="pt").input_ids
output = model.generate(input, max_new_tokens=len(input[0]), num_beams=4, do_sample=True)
text = tokenizer.decode(output[0], skip_special_tokens=True)