Downloads · 30 days
21
6% of all-time downloads
cihanunlu/BerTurk_Ottoman_Full_DAPT
BerTurk_Ottoman_Full_DAPT is a token classification model from cihanunlu. Use it when you need labels on individual words, such as names. It is set up for transformers.
Downloads · 30 days
21
6% of all-time downloads
All-time downloads
331
Public
Parameters
184M
738 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors738 MB · 99%
From the Hugging Face model README
A domain‐adaptive continuation of dbmdz/bert-base-turkish-128k-cased, pre‐trained on 800 K modern‐Latin Ottoman-Turkish sentences (≈ 14 M tokens) from the OTC Corpus (Özateş et al., 2025). This checkpoint is intended as a drop-in encoder for NER task.
| Property | Value |
|---|---|
| Base | dbmdz/bert-base-turkish-128k-cased |
| Domain data | BUCOLIN/OTC-Corpus |
| Pre‐training task | Masked Language Modeling (MLM) |
| Epochs | 4 |
| Sequence length | 128 tokens (chunked) |
| Batch size | 16 (per device) |
| Learning rate | 3 × 10⁻⁵ |
| Warmup steps | 500 |
| Weight decay | 0.01 |
| Mixed precision | fp16 |
| Checkpoint size | ≈ full weights, fp16 |
| Vocabulary | same as base |
BUCOLIN/OTC-Corpus
# Args
args = TrainingArguments(
output_dir="BerTurk_Ottoman_Full_DAPT",
per_device_train_batch_size=16,
per_device_eval_batch_size=16,
num_train_epochs=4,
learning_rate=3e-5,
eval_strategy="epoch",
save_strategy="epoch",
warmup_steps=500,
weight_decay=0.01,
fp16=True,
logging_steps=100,
save_steps=500,
eval_steps=500,
save_total_limit=2,
load_best_model_at_end=True,
)
from transformers import AutoTokenizer, AutoModelForMaskedLM, pipeline
# Load model & tokenizer
tokenizer = AutoTokenizer.from_pretrained("cihanunlu/BerTurk_Ottoman_Full_DAPT")
model = AutoModelForMaskedLM.from_pretrained("cihanunlu/BerTurk_Ottoman_Full_DAPT")
nlp = pipeline("fill-mask", model=model, tokenizer=tokenizer)
res = nlp("Devlet-i Aliyye-i Osmaniyye’nin [MASK] için tedâbîr-i mühimme ittikhāz olunmalıdır.")
print(res)