Downloads ยท 30 days
9
4% of all-time downloads
nahiar/xlm-roberta-ner
xlm-roberta-ner is a token classification model from nahiar. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
Indonesian ๐ฎ๐ฉ & English ๐ฌ๐ง | XLM-RoBERTa Base
Downloads ยท 30 days
9
4% of all-time downloads
All-time downloads
253
Public
Parameters
277M
3.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt2.2 GB ยท 67%
From the Hugging Face model README
Indonesian ๐ฎ๐ฉ & English ๐ฌ๐ง | XLM-RoBERTa Base
A fine-tuned XLM-RoBERTa-Base model for Named Entity Recognition (NER) on noisy social media text.
This model is optimized for multilingual informal content commonly found on:
It supports both Bahasa Indonesia and English, making it suitable for moderation systems, social listening, and content intelligence pipelines.
FacebookAI/xlm-roberta-baseThis model detects the following entity types:
| Label | Description |
|---|---|
| PER | Person |
| ORG | Organization |
| NOR | Political Organization |
| GPE | Geopolitical Entity |
| LOC | Location |
| FAC | Facility |
| LAW | Legal Entity (e.g., Undang-Undang) |
| EVT | Event |
| WOA | Work of Art |
BIO tagging format is used:
B-XXX โ Beginning of an entityI-XXX โ Inside an entityO โ Outside any entityEvaluated on held-out validation dataset:
| Metric | Score |
|---|---|
| F1 Score | 0.8387 |
| Precision | 0.8203 |
| Recall | 0.8580 |
| Training Loss | 0.0021 |
| Validation Loss | 0.1310 |
Evaluation Details
seqeval| Parameter | Value |
|---|---|
| Base Model | xlm-roberta-base |
| Training Samples | 695,108 |
| Validation Samples | 106,197 |
| Epochs | 5 |
| Learning Rate | 4e-5 |
| Batch Size | 32 |
| Optimizer | AdamW |
| Scheduler | Linear Warmup |
| Framework | Hugging Face Transformers |
from transformers import pipeline
ner = pipeline(
"token-classification",
model="nahiar/xlm-roberta-ner",
aggregation_strategy="simple"
)
text_id = "Jokowi menghadiri World Economic Forum di Davos."
text_en = "Apple is opening a new office in Jakarta next month."
print(ner(text_id))
print(ner(text_en))
"simple" โ Recommended (merges subword tokens)"first" โ Uses first token representation"average" โ Averages token scores"max" โ Takes maximum token scoreNOR, GPE, PER, ORG, EVT, LOC, LAW, FAC, WOAThis model may reflect demographic, geopolitical, or cultural biases present in the training dataset.
It is not intended to replace human judgment in high-risk or sensitive decision-making systems.
Human-in-the-loop review is strongly recommended for moderation or governance-related deployments.
Released under the Apache 2.0 License.
Free for commercial and research use.
@misc{hidayatuloh2026multilingualner,
author = {Nuri Hidayatuloh},
title = {Multilingual Named Entity Recognition for Social Media},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/nahiar/xlm-roberta-ner}
}