Downloads · 30 days
469
57% of all-time downloads
instilux/luni-mini
luni-mini is a fill-mask model from instilux. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
--- language: - lb license: cc-by-sa-4.0 libraryname: transformers pipelinetag: fill-mask tags: - modernbert - encoder - luxembourgish - multilingual - masked-language-modeling ---
Downloads · 30 days
469
57% of all-time downloads
All-time downloads
827
Public
Parameters
68.5M
274 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors274 MB · 99%
From the Hugging Face model README
language:
A ModernBERT-based masked language model pretrained on Luxembourgish (+ augmented Luxembourgish), following the Ettin recipe (see here: https://huggingface.co/jhu-clsp/ettin-encoder-68m)
lb/ltz)Requires transformers>=4.48.0.
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained("instilux/luni-mini")
model = AutoModelForMaskedLM.from_pretrained("instilux/luni-mini")
inputs = tokenizer("Wéi spéit [MASK] et?", return_tensors="pt")
mask_pos = (inputs["input_ids"] == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
with torch.no_grad():
outputs = model(**inputs)
top_tokens = outputs.logits[0, mask_pos].topk(5)
for token_id, score in zip(top_tokens.indices[0], top_tokens.values[0]):
token = tokenizer.decode(token_id)
print(f"{token:15s} {score:.3f}")
The tokenizer is BPE-minid (GPTNeoXTokenizerFast) with BERT-style special tokens ([CLS], [SEP], [MASK], [PAD]). A [CLS] token is prepended automatically (add_bos_token: true).
Paper coming soon!