Downloads · 30 days
0
rafmacalaba/gliner-large-data-mentions
gliner-large-data-mentions is a token classification model from rafmacalaba. Use it when you need labels on individual words, such as names. It is set up for gliner. The card lists the license as apache-2.0.
Fine-tuned from urchade/glinerlarge-v2.1 for extracting data source mentions in social science and global health research papers.
Downloads · 30 days
0
Access
Public
Updated Jun 4, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py27 KB · 76%
From the Hugging Face model README
Fine-tuned from urchade/gliner_large-v2.1 for extracting data source mentions in social science and global health research papers.
Status: Training job pending. See
train_gliner.pyfor the complete self-contained training script.
| Label | Description | Examples |
|---|---|---|
SURVEY | Named survey programs | Demographic and Health Survey, DHS, MICS, LSMS, Afrobarometer |
DATASET | Specific named datasets | Census microdata, 2010 Population and Housing Census, LSMS-ISA panel dataset |
DATABASE | Named databases/repositories | World Development Indicators, IHME GBD, IPUMS-DHS, FAOSTAT, GRID3 |
GEOCODED_DATA | Geocoded/spatial data mentions | GPS coordinates, geocoded DHS cluster coordinates, geo-referenced household data |
VAGUE_MENTION | Vague/informal references | "a survey from Ghana", "a nationally representative household survey" |
from gliner import GLiNER
model = GLiNER.from_pretrained("rafmacalaba/gliner-large-data-mentions")
text = "We used GPS-tagged household locations from the 2018 DHS and linked them to the WorldPop gridded population database."
entities = model.predict_entities(
text,
labels=["SURVEY", "DATASET", "DATABASE", "GEOCODED_DATA", "VAGUE_MENTION"],
threshold=0.4,
)
for e in entities:
print(f"[{e['label']}] '{e['text']}' (score={e['score']:.3f})")
# [GEOCODED_DATA] 'GPS-tagged household locations' (score=0.87)
# [SURVEY] '2018 DHS' (score=0.92)
# [DATABASE] 'WorldPop gridded population database' (score=0.83)
train_gliner.py for portabilityEntity distribution:
SURVEY : 55 spans (31.4%)
GEOCODED_DATA : 40 spans (22.9%)
DATABASE : 38 spans (21.7%)
DATASET : 24 spans (13.7%)
VAGUE_MENTION : 18 spans (10.3%)
Based on the original GLiNER paper (AAAI 2024, arxiv:2311.08526) fine-tuning recipe:
| Parameter | Value |
|---|---|
| Base model | urchade/gliner_large-v2.1 (DeBERTa-v3-large) |
| Epochs | 25 |
| Batch size | 8 |
| Encoder LR | 1e-5 |
| Head LR (others_lr) | 5e-5 |
| LR scheduler | cosine |
| Warmup | 8% of total steps |
| Loss | sum reduction, focal_loss_alpha=-1 |
| Negatives ratio | 1.0 |
| Masking | global |
| FP16 | True (GPU) |
from hf_jobs import run
run(
script="https://huggingface.co/rafmacalaba/gliner-large-data-mentions/raw/main/train_gliner.py",
dependencies=["gliner==0.2.26", "torch", "trackio"],
hardware_flavor="t4-small",
timeout="3h",
env={
"TRACKIO_SPACE_ID": "rafmacalaba/ml-intern-gliner01",
"TRACKIO_PROJECT": "gliner-data-mentions",
}
)
Or via the ML Intern agent tool:
hf_jobs(
operation="run",
script="https://huggingface.co/rafmacalaba/gliner-large-data-mentions/raw/main/train_gliner.py",
dependencies=["gliner==0.2.26", "torch", "trackio"],
hardware_flavor="t4-small",
timeout="3h",
trackio_space_id="rafmacalaba/ml-intern-gliner01",
trackio_project="gliner-data-mentions"
)
pip install gliner==0.2.26 torch trackio
python train_gliner.py
# or smoke test:
python train_gliner.py --smoke_test
For an autonomous agent completing this task, follow these steps in order:
gliner==0.2.26 API validated: Trainer, TrainingArguments, model._create_data_collator()TrainingArguments has others_lr, loss_reduction, negatives, masking, focal_loss_* paramsGLiNER.push_to_hub() confirmed availablereport_to="trackio" natively supported in TrainingArgumentsrafmacalaba/gliner-large-data-mentionshf_jobs(
operation="run",
script="https://huggingface.co/rafmacalaba/gliner-large-data-mentions/raw/main/train_gliner.py",
dependencies=["gliner==0.2.26", "torch", "trackio"],
hardware_flavor="t4-small",
timeout="3h",
trackio_space_id="rafmacalaba/ml-intern-gliner01",
trackio_project="gliner-data-mentions"
)
Expected: job starts, downloads urchade/gliner_large-v2.1 (~900MB), trains 25 epochs (~103 examples, 13 steps/epoch = 325 total steps), ~20-40 min on T4.
hf_jobs(operation="logs", job_id="<job_id_from_step2>")
Look for: === Starting training ===, then {'loss': ..., 'epoch': ...} every 5 steps.
Watch for OOM (reduce batch_size to 4, increase gradient_accumulation_steps to 2).
hf_repo_files(operation="list", repo_id="rafmacalaba/gliner-large-data-mentions")
Expect: pytorch_model.bin or model.safetensors, config.json, tokenizer_config.json.
Run inference on test sentences and confirm entities are extracted with score > 0.4:
[SURVEY] Demographic and Health Survey[GEOCODED_DATA], [SURVEY][VAGUE_MENTION] A survey from GhanaIf entity scores are low (< 0.5) or entities are missed:
This model repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.