Downloads · 30 days
26
100% of all-time downloads
davanstrien/hub-task-tagger-gliner2.5-base
hub-task-tagger-gliner2.5-base is a text classification model from davanstrien. Use it when you need a label for a piece of text. It is set up for gliner2. The card lists the license as apache-2.0.
Suggests task tags for a Hugging Face dataset, such as text-generation, robotics or automatic-speech-recognition, from its column names and first row. It picks from the 52 task tags the Hub offers, and gives a probabi…
Downloads · 30 days
26
100% of all-time downloads
All-time downloads
26
Public
Parameters
194M
774 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors774 MB · 99%
From the Hugging Face model README
Suggests task tags for a Hugging Face dataset, such as text-generation, robotics or
automatic-speech-recognition, from its column names and first row. It picks from the 52 task
tags the Hub offers, and gives a probability for each.
TypeSafe AI's Jev has got people interested in "System One" models: models that don't generate text, but read a state and return typed answers with probabilities. This is a small, open example of the same idea for one narrow job, fine-tuned in 17 minutes on Hugging Face Jobs.
Try it: davanstrien/hub-task-tagger. Paste a dataset id and see the suggested tags, the owner's tags, and the exact text the model read.
The model reads a short text built from the dataset viewer's preview: a Columns: line with each
column's name and type, then the first row, cut to 370 tokens. It must be built exactly as in
training, so this repo ships the builder,
tagger_core.py,
next to the weights. The snippet below downloads it with huggingface_hub and uses it.
import json, math, os, sys
from huggingface_hub import hf_hub_download
from tokenizers import Tokenizer
from gliner2.classification import Classifier, ClassificationSchema, ClassificationConfig
repo = "davanstrien/hub-task-tagger-gliner2.5-base"
# 1. Build the model's input for a dataset, with the code that built the training data.
path = hf_hub_download(repo, "tagger_core.py")
sys.path.insert(0, os.path.dirname(path)) # make the downloaded file importable
import tagger_core as core
builder = core.StateBuilder(Tokenizer.from_file(hf_hub_download("Qwen/Qwen3.5-4B-Base", "tokenizer.json")), 370)
first_rows, config, split, error = core.fetch_first_rows("fancyzhx/ag_news")
state, truncated = builder.build(core.render_first_rows(first_rows))
print(state)
# 2. Score all 52 tags.
labels = json.load(open(hf_hub_download(repo, "classification_schema.json")))["tasks"][0]["labels"]
model = Classifier.from_pretrained(repo).eval()
schema = ClassificationSchema().multi("labels", labels)
logits = model.batch_score([state[:2000]], schema, config=ClassificationConfig(batch_size=1))[0].tasks["labels"]
# 3. Calibrated probabilities: softmax over the 52 logits at temperature 1.30 (fitted on held-out data).
z = [logits[label] / 1.30 for label in labels]
m = max(z)
e = [math.exp(v - m) for v in z]
probs = sorted(zip(labels, (v / sum(e) for v in e)), key=lambda x: -x[1])
print(probs[:3])
# A set of tags: everything at 0.225 or above, and always the top tag.
suggested = [label for label, p in probs if p >= 0.225] or [probs[0][0]]
print(suggested)
The Qwen tokenizer is only used to measure and cut the input to the length the model was trained
on. fetch_first_rows returns an error instead of rows when a dataset has no viewer preview.
Install with pip install "gliner2[local]==2.0.0". Weights are fp32. On a free 2-vCPU Space one
prediction takes about 0.7–1 s; on an L4 GPU in fp16, about 20 ms.
train-gliner2.py from
uv-scripts/classification, run as one hf jobs uv run command on an rtx-pro-6000:
--base-model fastino/gliner2.5-base-v1 --labels-file labels.json --label-column labels --label-augmentation off,
5 epochs, seed 0, bf16. About 17 minutes of training, roughly $1.50.task_categories. Each example is the
dataset's column names and first row (from the dataset viewer), cut to about 370 tokens, and its
declared tags (multi-label). Retired tags were mapped to current ones where the task is the same
(for example text2text-generation → text-generation); others were dropped.On the 3,000 development datasets. A prediction counts as correct if it is one of the owner's tags.
| Top-1 | Top-3 | Macro recall (24 tags) | |
|---|---|---|---|
| This model | 0.695 | 0.881 | 0.525 |
| Same recipe, second seed | 0.686 | 0.879 | 0.511 |
| GLiNER2.5-small, same recipe | 0.653 | 0.855 | 0.458 |
| GLiNER2.5-base, zero-shot | 0.102 | 0.213 | 0.099 |
| Always "text-generation" | 0.320 | – | – |
tabular-regression with tabular-classification,
because the input does not say which column is the target.text-to-3d or voice-activity-detection) had too few
training examples to measure.text-generation, or the reverse.Built on GLiNER2 by Fastino. Training, evaluation and the demo were run on Hugging Face Jobs and Spaces, with much of the work done by coding agents (Claude) and checked by a person. More detail: GLINER2-NOTES.md.