Downloads · 30 days
0
Shengyu/bioproc-benchmark
bioproc-benchmark is a machine learning model from Shengyu. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository accompanies our study on robustness-oriented biomedical named entity recognition (BioNER). It provides the resources needed to reproduce the benchmark construction, baseline experiments, and HybridNER…
Downloads · 30 days
0
Access
Public
Updated Apr 3, 2026
Repo size
18.9 GB
Likes
1
Public
Click a slice to open those files.
.bin9.8 GB · 52%
From the Hugging Face model README
This repository accompanies our study on robustness-oriented biomedical named entity recognition (BioNER). It provides the resources needed to reproduce the benchmark construction, baseline experiments, and HybridNER results reported in the paper.
The repository includes:
BioNER is commonly evaluated on homogeneous benchmarks, which can mask issues with robustness and cross-domain generalization. We introduce BioProCorpus, a task-oriented corpus from iProX proteomics project descriptions, designed to emphasize lexical variability, boundary ambiguity, and long-range context. Under a unified protocol, we benchmark BigBird, BioBERT, DeBERTa, KeBioLM, and BINDER across four datasets and propose HybridNER, which achieves the best performance on BioProCorpus.
├─ 1.Data_comparison/
│ ├─ 1.1_Datasets/
│ │ ├─ RevisedJNLPBA/
│ │ ├─ BioRED/
│ │ ├─ BioProCorpus/
│ │ └─ AnatEM/
│ │
│ ├─ 1.2_Train_code/
│ │ ├─ training_and_testing_RevisedJNLPBA.ipynb
│ │ ├─ training_and_testing_BioRED.ipynb
│ │ ├─ training_and_testing_BioProCorpus.ipynb
│ │ ├─ training_and_testing_AnatEM.ipynb
│ │ └─ BINDER-model/ # BINDER training/validation/testing on 4 datasets
│ │
│ └─ 1.3_Best_models_and_predictions/
│ ├─ RevisedJNLPBA/
│ ├─ BioRED/
│ ├─ BioProCorpus/
│ └─ AnatEM/
│
├─ 2.1HybridNER_training/ # HybridNER training code
│ ├─ train_stage_a.py
│ ├─ train_stage_b_kd_multi.py
│ ├─ train_stage_c_finetune.py
│ └─ eval_full_report.py
│
├─ 2.2HybridNER_trained/
│ ├─ model.pt
│ └─ model-PG_emphasis.pt
│
├─ 2.3Original_data_processing_code/
│ ├─ training_and_testing.ipynb
│ ├─ check_data_distribution_balance.ipynb
│ ├─ query_entity_terms_by_type.ipynb
│ ├─ entity_term_statistics.ipynb
│ ├─ data_split_processing_stratified.ipynb
│ ├─ data_split_processing_cross_file.ipynb
│ ├─ data_whitespace_cleaning.ipynb
│ └─ data_splitting.ipynb
│
├─ Entity Annotation Guidelines.pdf
└─ README.md
Four datasets are provided with train/development/test splits:
These datasets are organized for unified BioNER benchmarking under the same training and evaluation protocol.
The folder for baseline experiments contains notebook-based training, validation, and testing code for:
It also includes the BINDER-model/ subfolder for BINDER experiments on all four datasets.
This folder stores:
2.1 HybridNER_training/ contains HybridNER training code.
2.2 Trained_models/ provides two released model checkpoints:
model.ptmodel-PG_emphasis.ptThe folder 2.3Original_data_processing_code/ contains the preprocessing notebooks used to construct and inspect BioProCorpus before model training. These notebooks cover:
This part is included to support reproducibility from the original annotated data to the final benchmark-ready train/development/test partitions.
The experiments were developed and tested in a Linux-based environment. A typical setup is:
A GPU environment is strongly recommended for both baseline training and HybridNER training.
conda create -n bioproc python=3.10 -y
conda activate bioproc
Please install the required packages manually or from your own pinned requirement file. A typical environment includes:
pip install torch torchvision torchaudio
pip install transformers datasets tokenizers sentencepiece
pip install scikit-learn seqeval pandas numpy tqdm matplotlib jupyter notebook
pip install accelerate optuna
If your code uses additional packages such as torchcrf, evaluate, or domain-specific tokenizers, please also install them as needed:
pip install torchcrf evaluate
jupyter notebook
Then open the notebook files in:
1.Data_comparison/1.2_Train_code/
or
2.3Original_data_processing_code/
depending on whether you are reproducing preprocessing or baseline experiments.
To reproduce the processing pipeline for BioProCorpus, first go to:
2.3Original_data_processing_code/
Recommended execution order:
data_whitespace_cleaning.ipynb
Clean redundant whitespace and formatting artifacts.entity_term_statistics.ipynb
Compute corpus-level and entity-level statistics.query_entity_terms_by_type.ipynb
Inspect entity terms by category.check_data_distribution_balance.ipynb
Check data distribution and category balance.data_split_processing_stratified.ipynb
Generate within-file stratified train/development/test splits.data_split_processing_cross_file.ipynb
Analyze or compare cross-file split settings if needed.data_splitting.ipynb
Finalize the benchmark-ready split files.This stage reconstructs the corpus from cleaned annotated data to the final benchmark partitions used in the paper.
Navigate to:
1.Data_comparison/1.2_Train_code/
Open the notebook corresponding to the target dataset and run all cells in order.
Example notebooks include:
training_and_testing_RevisedJNLPBA.ipynbtraining_and_testing_BioRED.ipynbtraining_and_testing_BioProCorpus.ipynbtraining_and_testing_AnatEM.ipynbThese notebooks implement training, validation, and testing under the unified protocol for BigBird, BioBERT, DeBERTa, and KeBioLM.
For BINDER, use the scripts or notebooks in:
1.Data_comparison/1.2_Train_code/BINDER-model/
Navigate to:
2.1HybridNER_training/
Run the following scripts in order:
python train_stage_a.py
python train_stage_b_kd_multi.py
python train_stage_c_finetune.py
python eval_full_report.py
Suggested interpretation of the stages:
train_stage_a.py: initial supervised trainingtrain_stage_b_kd_multi.py: multi-teacher knowledge distillation stagetrain_stage_c_finetune.py: final fine-tuningeval_full_report.py: evaluation and report generationPlease ensure that dataset paths, pretrained model paths, output directories, and GPU settings are configured correctly in each script before execution.
If you do not want to retrain HybridNER from scratch, you can directly use the released checkpoints in:
2.2HybridNER_trained/
Available checkpoints:
model.ptmodel-PG_emphasis.ptPlease load the desired checkpoint in the evaluation script and update the checkpoint path accordingly.
BioProCorpus is organized into fixed training, development, and test sets for all downstream experiments. The benchmark split was designed to support fair comparison across all baseline models and HybridNER. The preprocessing notebooks in 2.3Original_data_processing_code/ document the main steps used to prepare the final corpus partitions and inspect benchmark composition.
The BioProCorpus dataset, HybridNER source code, preprocessing notebooks, split-generation scripts, released checkpoints, and related benchmark resources used in this study are freely available for non-commercial use at:
https://huggingface.co/Shengyu/bioproc-benchmark/
If you use this repository or BioProCorpus in your research, please cite the corresponding paper.
@article{your_paper_here,
title = {BioProCorpus: A corpus for robustness-oriented biomedical NER benchmarking and HybridNER evaluation under heterogeneous conditions},
author = {Liu, Shengyu and others},
journal = {To be updated},
year = {2026}
}
For questions about the dataset, code, or reproducibility, please contact: