Downloads · 30 days
73
8% of all-time downloads
NIHRDataInsights/HRCSHealthCategories
HRCSHealthCategories is a text classification model from NIHRDataInsights. Use it when you need a label for a piece of text. The card lists the license as mit.
This model, developed by the National Institute for Health and Care Research (NIHR), assigns HRCS Health Categories (HCs) to research awards using the award title and abstract (micro F1 = 0.81). It is a multi-label tr…
Downloads · 30 days
73
8% of all-time downloads
All-time downloads
905
Public
Parameters
335M
1.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
This model, developed by the National Institute for Health and Care Research (NIHR), assigns HRCS Health Categories (HCs) to research awards using the award title and abstract (micro F1 = 0.81). It is a multi-label transformer classifier built on BiomedBERT-large, domain-adapted (DAPT) on healthcare grant titles and abstracts, then fine-tuned on cross-funder labelled HRCS data. The goal is to support portfolio analysis, automated tagging, and reproducible classification of biomedical research funding.
microsoft/BiomedNLP-BiomedBERT-large-uncased-abstractThe model was trained in two stages using a 24GB GPU.
We continued masked language modelling on grant titles and abstracts to adapt the encoder to research funding language as opposed to publications. This data used was a healthcare funder specific subset of Gomez Magenti, J. (2025) ‘Harmonised datasets of research project grants from UK and European funders’. Zenodo. doi:10.5281/zenodo.15479412.
Settings:
The adapted checkpoint was then used for supervised training.
The adapted model was fine-tuned for multi-label classification using sigmoid outputs and binary cross-entropy loss.
Input format:
AwardTitle + newline + AwardAbstract
Tokenisation:
Handling class imbalance:
A per-label weighting vector (pos_weight) is applied in the loss to reduce bias toward common categories.
Training configuration:
Data was split into three disjoint sets:
The test set was not used during training, checkpoint selection, or threshold tuning. The dataset used is listed at the top of the model card. Predictions are converted to labels using per-category probability thresholds tuned on the validation set. These thresholds are included in metadata.json.
Overall Metrics:
For a comprehensive breakdown of the model's performance, including Overall Metrics, Metrics per Category across both validation and test sets and Metrics per Funder across the validation set, please refer to the detailed evaluation spreadsheet included in this repository.
Download/View the Evaluation Results (Located in the Files and versions tab of this repository).
This model is intended for:
It is not intended to completely replace expert review.
We have provided a ready-to-use Python script that runs both this model (Health Categories) and a Research Activity Codes model on new award data simultaneously.
You can download the script and a sample dataset directly from the 'inference' subfolder in the Files and version tab of this repository.
Prerequisites:
pip install torch pandas numpy tqdm transformers huggingface_hubAwardTitle and AwardAbstract.Running the Code:
# --- USER SETTINGS --- section, update DATA_FOLDERS to point to the folder containing your CSV. (Leave it as ["./"] if your CSV is in the same folder as the script).TEST_FILENAME to match the name of your CSV.The script will automatically download the necessary AI models, process your text, and output a new CSV containing the predicted categories and an "AI Certainty Score" (SmallestLogitDiff) to help you identify which borderline grants require human review.
In addition to predicted labels, the inference script reports how close each prediction is to the model’s decision boundary in logit space. This is computed as the smallest absolute difference between any category’s logit and its corresponding decision threshold.
Records with logits close to the threshold represent borderline cases where the model is uncertain. These can be prioritised for human review, while higher-confidence predictions can be automated.
When progressively excluding records whose predictions lie closest to the decision boundary, the remaining high-confidence subset shows increasing accuracy:
| % of records excluded for human review | Micro-F1 on remaining subset |
|---|---|
| 0% | 0.81 |
| 10% | 0.83 |
| 20% | 0.85 |
| 30% | 0.87 |
| 40% | 0.89 |
| 50% | 0.90 |
| 60% | 0.91 |
| 70% | 0.92 |
| 80% | 0.94 |
| 90% | 0.96 |
This demonstrates that the model supports hybrid workflows in which uncertain cases are reviewed by experts while confident predictions can be automated.
NIHR, 2026. HRCS Health Category Classifier (BiomedBERT, DAPT). [Model]. Developed by Banks, A., Baghurst, D., Carter, J., Manville, C., Wang, K. and Downes, N. Available from: https://huggingface.co/NIHRDataInsights/HRCSHealthCategories