Downloads · 30 days
11
13% of all-time downloads
PeytonT/metadata-category-classifier
metadata-category-classifier is a text classification model from PeytonT. Use it when you need a label for a piece of text. It is set up for transformers.
Classifies paper metadata from title and abstract text into arXiv-style category labels.
Downloads · 30 days
11
13% of all-time downloads
All-time downloads
86
Public
Parameters
110M
883 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors440 MB · 100%
From the Hugging Face model README
Classifies paper metadata from title and abstract text into arXiv-style category labels.
allenai/scibert_scivocab_uncasedM2T1_metadatav2This model is part of the Repository Library stack, a research system for indexing, retrieving, aligning, and reasoning over scientific papers, structured paper content, repositories, and cross-domain links between them.
The training inputs for this package were assembled from the following Repository Library data sources:
arxiv_metadata: arXiv metadata records containing titles, abstracts, authors, categories, and update metadata.The v2 hardening pass restricts the classifier target space to a denser category set:
max_labels: 32min_per_label: 16max_per_label: 128arxiv_metadatatitle, abstractcategories[2000, 2025][0.9, 0.1, 0.0]bf16cross_entropy5e-05full_finetuneddpDeclared metrics: accuracy, macro_f1
Tracked v2 eval metrics on the current held-out split:
eval_loss: 0.003253802889958024eval_accuracy: 1.0eval_macro_f1: 1.0eval_balanced_accuracy: 1.0eval_label_count: 31Status: this card reflects the current tracked experiment configuration and packaged weights in the Repository Library model stack. The v2 scores are very high and should receive a leakage audit before being treated as a final benchmark.
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo_id = "PeytonT/metadata-category-classifier"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)
text = "TITLE: Example paper title
ABSTRACT: This paper studies representation learning for scientific documents."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
predicted_id = outputs.logits.argmax(dim=-1).item()
label = model.config.id2label.get(predicted_id, str(predicted_id))
print(label)