Downloads · 30 days
22
43% of all-time downloads
Ericu950/Stoicheia-macronizer
Stoicheia-macronizer is a token classification model from Ericu950. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
Stoicheia is a 405M-parameter character-level masked-diffusion encoder for Ancient Greek (dmodel 1024, depth 32, banded attention: three of every four blocks attend within a 256-character window, the fourth globally).…
Downloads · 30 days
22
43% of all-time downloads
All-time downloads
51
Public
Parameters
405M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
Stoicheia is a 405M-parameter character-level masked-diffusion encoder for Ancient Greek
(d_model 1024, depth 32, banded attention: three of every four blocks attend within a
256-character window, the fourth globally). Its input is factored into five aligned planes --
letters, word/sentence boundaries, diacritics, capitalization, punctuation -- each of which can
be masked independently to an explicit unknown state at inference. That is what lets one model
read an edited text, scriptio continua, and a lacuna of unknown length without changing
anything but its input.
Anonymous release accompanying a paper under review.
Vowel length alone: long versus short at every ambiguous bare α, ι or υ. Trained on a silver corpus of ~130,000 verse lines built by exact constraint propagation -- a solver accepts a line only when exactly one metrical grammar scans it, and fixes a dichronon only when every accepting parse agrees -- plus converted syllable-weight markup, all checked against the evaluation benchmark to prevent leakage.
import torch
from transformers import AutoModel
from huggingface_hub import hf_hub_download
REPO = "Ericu950/Stoicheia-macronizer"
model = AutoModel.from_pretrained(REPO, trust_remote_code=True).eval()
hf_hub_download(repo_id=REPO, filename="processing_char_bert_meter.py", local_dir=".")
from processing_char_bert_meter import CharBertMeterProcessor
proc = CharBertMeterProcessor()
batch = proc("ἄνδρα μοι ἔννεπε, μοῦσα, πολύτροπον, ὃς μάλα πολλὰ")
with torch.no_grad():
out = model(**{k: v for k, v in batch.items() if not k.startswith("_")})
print(proc.decode_macronization(out, batch)) # ἄ^νδρα^ μοι ἔννεπε, μοῦσα^, ...