Downloads · 30 days
0
ottema/gliner2-ptbr-harem
gliner2-ptbr-harem is a token classification model from ottema. Use it when you need labels on individual words, such as names. It is set up for gliner. The card lists the license as apache-2.0.
GLiNER2 fine-tuned for Brazilian Portuguese NER, benchmarked on HAREM. Best entity F1 among the compared models in our evaluation protocol.
Downloads · 30 days
0
Access
Public
Updated Jun 20, 2026
Parameters
307M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 98%
From the Hugging Face model README
GLiNER2 fine-tuned for Brazilian Portuguese NER, benchmarked on HAREM. Best entity F1 among the compared models in our evaluation protocol.
This model is part of the Ottema GLiNER2-PTBR open-source ecosystem. Companion model: ottema/gliner2-ptbr (generalist, informal PT-BR).
This model is a fine-tune of fastino/gliner2-multi-v1, the official multilingual GLiNER2 model released by Fastino. GLiNER2 is the open-vocabulary NER architecture originally proposed by Urchade Zaratiana and collaborators (GLiNER paper). We are grateful to the upstream teams for releasing the architecture and base model under Apache-2.0, which made this work possible.
fastino/gliner2-multi-v1 (Fastino)If you use this model, please also cite the original GLiNER work and the Fastino GLiNER2 release.
Metrics are reported as per-sample macro F1 (the standard in our benchmark script). The corresponding global micro F1 is also reported for transparency.
| Model | entity_F1 (per-sample macro) | entity_F1 (global micro) | span_F1 | label_F1 | Latency |
|---|---|---|---|---|---|
| ottema/gliner2-ptbr-harem (v0.12b) @ t=0.4 ⭐ | 0.4749 | 0.4501 | 0.4878 | 0.8725 | 31ms |
| hcaeryks/bert-crf-harem (BERT-Large specialist) | 0.4700 | — | 0.5220 | 0.8456 | 131ms |
| ottema/gliner2-ptbr-harem v0.11 (previous official) | 0.4711 | — | 0.4811 | 0.8776 | 32ms |
| fastino/gliner2-multi-v1 (zero-shot) | 0.4251 | — | 0.4366 | 0.8480 | 31ms |
Note on aggregation methods:
Key results:
Trade-offs:
from gliner2 import GLiNER2
model = GLiNER2.from_pretrained("ottema/gliner2-ptbr-harem")
model = model.to("cuda") # or "cpu"
text = "João da Silva nasceu em São Paulo em 1990 e trabalha na Petrobras."
entities = model.extract_entities(
text,
entity_types=["pessoa", "organização", "local", "data", "valor_monetário"],
threshold=0.4,
)
print(entities)
# {'entities': {'pessoa': ['João da Silva'], 'local': ['São Paulo'], 'data': ['1990'], 'organização': ['Petrobras']}}
Recommended threshold: 0.4 (sweet spot from ablation).
We ran 5 experiments beyond standard fine-tuning. Full ablation below:
Findings: Pseudo-labeling works at threshold 0.85 with conservative LR (1e-6). More aggressive filtering or self-training iteration causes overconfidence without F1 improvement. Hard-negative augmentation trades recall for precision (not net positive).
ottema/gliner2-ptbr v0.4 for that).@software{ottema_gliner2_ptbr_2026,
author = {Ottema},
title = {GLiNER2-PTBR: Open-source Brazilian Portuguese NER},
year = {2026},
version = {0.12b},
}
ottema/gliner2-ptbr (v0.4): generalist for informal PT-BR (chat, atendimento)fastino/gliner2-multi-v1: base multilingual GLiNER2