Downloads · 30 days
32
8% of all-time downloads
njand/latin-asr-postprocessor
latin-asr-postprocessor is a token classification model from njand. Use it when you need labels on individual words, such as names. The card lists the license as mit.
An Inverse Text Normalization (ITN) transformer model fine-tuned to convert unformatted, raw Latin Automatic Speech Recognition (ASR) outputs into fully formatted, classical Latin text. It simultaneously restores capi…
Downloads · 30 days
32
8% of all-time downloads
All-time downloads
420
Public
Parameters
111M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.onnx632 MB · 59%
From the Hugging Face model README
An Inverse Text Normalization (ITN) transformer model fine-tuned to convert unformatted, raw Latin Automatic Speech Recognition (ASR) outputs into fully formatted, classical Latin text. It simultaneously restores capitalization and trailing punctuation using a 14-class composite sequence labeling schema.
latincy/latin-bertnjand/latin-asr-post-processing-datasetThis model is intended to be used directly downstream of the acoustic model njand/wav2vec2-xls-r-latin.
Because raw ASR models emit stream-of-consciousness text (lowercased, space-separated, and unpunctuated), the text must pass through an input normalization pipeline before being fed into this model for casing and punctuation restoration.
+-----------------------+ +-------------------------------+ +-------------------------------+
| Raw Audio Waveform | --> | njand/wav2vec2-xls-r-latin | --> | Preprocessing & Normalization |
+-----------------------+ +-------------------------------+ +-------------------------------+
|
v
+-----------------------+ +-------------------------------+ +-------------------------------+
| Formatted Text Output | <-- | Latin ASR Post-Processor | <-- | Custom CLTK Tokenization |
+-----------------------+ +-------------------------------+ +-------------------------------+
To prepare raw transcript outputs for inference, apply the following sequence of transformations:
Note: Because official CLTK v0 tokenization scripts are unmaintained, a bespoke implementation of the tokenizer was executed dynamically during training preprocessing rather than being pre-applied to the static dataset.
Below is a complete Python script demonstrating how to prepare raw ASR output and run inference using the post-processing pipeline.
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
PUNCT_MAP = {
"NONE": "",
"COMMA": ",",
"PERIOD": ".",
"SEMICOLON": ";",
"COLON": ":",
"QUESTION": "?",
"EXCLAMATION": "!",
}
def format_token(word: str, tag: str) -> str:
"""Applies composite ITN tag (e.g., 'TITLE_COMMA') to a word token."""
parts = tag.split("_")
if len(parts) != 2:
return word
casing, punct = parts[0], parts[1]
if casing == "TITLE":
word = word.capitalize()
elif casing == "LOWER":
word = word.lower()
return f"{word}{PUNCT_MAP.get(punct, '')}"
def restore_latin_text(pipe, raw_text: str) -> str:
"""Runs inference and reconstructs formatted Latin text."""
predictions = pipe(raw_text, aggregation_strategy="first")
formatted_words = [
format_token(pred["word"].strip(" "), pred["entity_group"])
for pred in predictions
]
return " ".join(formatted_words)
# 1. Load pipeline
model_id = "njand/latin-asr-postprocessor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
itn_pipe = pipeline("token-classification", model=model, tokenizer=tokenizer)
# 2. Test reconstruction with preprocessed ASR output
raw_asr_input = "gallia est omnis divisa in partes tres quarum unam incolunt belgae"
print(restore_latin_text(itn_pipe, raw_asr_input))
# Output: "Gallia est omnis divisa in partes tres, quarum unam incolunt Belgae."
Target labels utilize a 14-class composite sequence schema that pairs Casing state with Trailing Punctuation state:
$$\text{Label} = \text{Casing} \times \text{Punctuation}$$
LOWER, TITLENONE, COMMA, PERIOD, SEMICOLON, COLON, QUESTION, EXCLAMATIONTo facilitate production deployment on CPU-based infrastructure, this repository provides the model in three formats:
| Format | Precision | File Size | Latency (P50) | Recommended Use Case |
|---|---|---|---|---|
| PyTorch | FP32 | 443 MB | - | Training, fine-tuning, and PyTorch pipelines |
| ONNX | FP32 | 443 MB | 21.5 ms | Production (Maximum Accuracy) |
| ONNX Quantized | INT8 | 188 MB | 9.7 ms | Low-Latency & Edge CPU |
Performance vs. Precision Trade-off: While INT8 dynamic quantization yields a 2.2× speedup (P50) and cuts RAM usage by 57.5%, top-line accuracy (92.16% → 91.03%) hides a severe drop in macro performance:
- Macro F1 Collapse: Drops from 65.19% to 50.00%. Dynamic weight quantization compresses logit decision boundaries for rare token tags.
- Punctuation Degradation: Punctuation F1 falls 10.66 percentage points (69.09% → 58.43%), causing increased missing or misclassified commas, colons, and sentence boundaries.
- Casing Stability: Capitalization F1 remains mostly intact (91.71% → 89.38%).
Recommendation: Use ONNX FP32 for production pipelines where text formatting and punctuation precision are critical. Use ONNX INT8 in latency-critical environments where speed and memory constraints outweigh exact punctuation recovery.
Evaluated on a 95/5 train/holdout split across diverse Classical Latin literary and historical corpora.
| Metric | Score |
|---|---|
| Overall Accuracy | 92.16% |
| Macro F1 | 0.6523 |
| Precision | 62.28% |
| Recall | 69.43% |
| Validation Loss | 0.2592 |
| Task | Accuracy | F1 Score |
|---|---|---|
| Casing Restoration | 98.35% | 0.9171 |
| Punctuation Insertion | 93.67% | 0.6908 |
The model was trained over 9 epochs fine-tuning latincy/latin-bert. Model weights from Epoch 7 were selected based on optimal overall F1.
| Epoch | Train Loss | Val Loss | Overall F1 | Overall Acc | Casing Acc | Punct Acc |
|---|---|---|---|---|---|---|
| 1 | 0.6284 | 0.2883 | 0.6097 | 91.05% | 98.08% | 92.79% |
| 2 | 0.5617 | 0.2701 | 0.6275 | 91.60% | 98.19% | 93.26% |
| 3 | 0.5150 | 0.2618 | 0.6388 | 91.91% | 98.27% | 93.50% |
| 4 | 0.4855 | 0.2607 | 0.6455 | 92.04% | 98.30% | 93.60% |
| 5 | 0.4652 | 0.2579 | 0.6440 | 92.12% | 98.33% | 93.65% |
| 6 | 0.4502 | 0.2582 | 0.6462 | 92.15% | 98.34% | 93.66% |
| 7 | 0.4351 | 0.2592 | 0.6523 | 92.16% | 98.35% | 93.67% |
| 8 | 0.4262 | 0.2598 | 0.6495 | 92.24% | 98.36% | 93.75% |
| 9 | 0.4172 | 0.2601 | 0.6507 | 92.21% | 98.37% | 93.71% |