Downloads · 30 days
0
ottema/gliner2-ptbr
gliner2-ptbr is a token classification model from ottema. Use it when you need labels on individual words, such as names. It is set up for gliner. The card lists the license as apache-2.0.
Open-vocabulary NER for Brazilian Portuguese, fine-tuned for informal and operational text (chat, atendimento, suporte).
Downloads · 30 days
0
Access
Public
Updated Jun 20, 2026
Parameters
307M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 98%
From the Hugging Face model README
Open-vocabulary NER for Brazilian Portuguese, fine-tuned for informal and operational text (chat, atendimento, suporte).
This is the generalist release. For HAREM-specialized (best entity F1 among compared models), see ottema/gliner2-ptbr-harem. For ontology-guided evidence extraction, see ottema/gliner2-ptbr-ontoevidence.
fastino/gliner2-multi-v1 (Apache-2.0)General-purpose open-vocabulary NER for informal and operational Brazilian Portuguese text: atendimento, chat, suporte técnico, educação. Trained on synthetic data covering pessoas, profissões, locais, organizações, documentos, produtos, marcas, tecnologias, telefones, e-mails, datas, e valores monetários.
If you need a model benchmarked on journalistic Portuguese (HAREM), use ottema/gliner2-ptbr-harem instead.
from gliner2 import GLiNER2
model = GLiNER2.from_pretrained("ottema/gliner2-ptbr")
text = "A professora Ana comprou um notebook Dell em Campinas no dia 12/06."
labels = ["pessoa", "profissão", "produto", "marca", "local", "data"]
entities = model.extract_entities(text, labels, threshold=0.5)
for label, spans in entities["entities"].items():
for span in spans:
print(f"{span} -> {label}")
Evaluation on data/gliner_ptbr_core/test.jsonl (synthetic generalist benchmark, threshold 0.3):
| Model | entity_F1 | span_F1 | label_F1 |
|---|---|---|---|
fastino/gliner2-multi-v1 (zero-shot) | 0.9333 | 0.9347 | 0.9855 |
ottema/gliner2-ptbr (v0.4) | 0.9976 | 0.9976 | 1.0000 |
On HAREM (163 samples, 2511 entities, journalistic PT-BR — out-of-distribution for this generalist):
| Model | entity_F1 | Δ vs baseline |
|---|---|---|
fastino/gliner2-multi-v1 (zero-shot) | 0.4271 | (reference) |
ottema/gliner2-ptbr (v0.4) | 0.4132 | -1.39 pp |
The generalist is best for the synthetic informal-PT-BR distribution it was trained on. For journalistic text, see the HAREM-specialized model.
fastino/gliner2-multi-v1 (Fastino)Apache-2.0
ottema/gliner2-ptbr-harem — HAREM-specialized (best entity F1)ottema/gliner2-ptbr-ontoevidence — ontology-guided evidence extraction (in development)ottema/gliner2-ptbr-ontoevidence-data — OntoEvidence-BR dataset