Downloads · 30 days
0
0% of all-time downloads
estnafinema0/active-learning-nerc-models-kfold
active-learning-nerc-models-kfold is a machine learning model from estnafinema0. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains a series of NER models trained using an active learning framework. Our approach leverages a low-quality (cheap) dataset combined with high-quality expert annotations to iteratively improve ent…
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
1
Public
Repo size
27.6 GB
Likes
0
Public
Click a slice to open those files.
.md7.4 KB · 83%
From the Hugging Face model README
This repository contains a series of NER models trained using an active learning framework. Our approach leverages a low-quality (cheap) dataset combined with high-quality expert annotations to iteratively improve entity recognition performance. The core idea is to begin with a model trained solely on the cheap dataset (model_llm_pure) and then incrementally fine-tune it by selecting the most uncertain expert examples based on an uncertainty estimation module.
Our baseline model, model_llm_pure, achieves limited performance, while the model model_init_12, fine-tuned on the cheap dataset plus an additional 12% of expert examples, demonstrates a significant improvement. The active learning loop further refines the model by iteratively adding the most informative examples and saving each intermediate model in a dedicated branch.
This module provides an improved metric to evaluate model performance at the entity level, which is crucial for NER. A correct prediction requires that the entire entity (with proper boundaries and correct labels) is recognized correctly.
Key Steps:
Prediction Collection:
The evaluation function processes each batch from the evaluation DataLoader and, for each sentence, collects predicted and true labels in a list-of-lists format.
Metric Calculation:
Using the seqeval library, we compute:
This module estimates the uncertainty of each sentence by computing the average entropy of its tokens. A high average entropy indicates that the model is less confident in its predictions for that sentence.
Process:
Pass each sentence (example) through the model in evaluation mode (with gradients disabled).
Retrieve logits and apply softmax to obtain a probability distribution over labels for each token.
Compute the entropy for each valid token (i.e., where ner_tag_mask == 1):
$$ H(token) = - \sum_{y} P(y\mid token) \log P(y\mid token) $$
The average entropy across valid tokens serves as the sentence’s uncertainty measure.
Before initiating the active learning loop, we run a preliminary experiment using k-fold (5-fold) cross-validation on the expert dataset. This experiment determines the minimal volume of expert examples that yield a significant improvement over the baseline model.
Procedure:

Below is a comparison of the initial evaluation metrics for two baseline models:
| Model | Validation Loss | Seqeval Accuracy | F1-Score |
|---|---|---|---|
| model_llm_pure | 0.53443 | 0.85185 | 0.47493 |
| model_init_12 | 0.33402 | 0.93084 | 0.65344 |
Model model_init_12 is obtained by fine-tuning the base model on the cheap dataset combined with an additional 12% of expert examples, demonstrating significantly improved performance.
The core active learning loop starts from a pre-trained model (typically model_init_12) and iteratively:
batch_to_add).Note: Each intermediate model is saved in its own branch (e.g., active_iter_1_added_20, active_iter_2_added_40, etc.), which allows for easy comparison and retrieval later.
Each time the model is fine-tuned in the active learning loop, it is saved to a dedicated branch on Hugging Face. For example, to save the current model in a branch, use:
branch_name = "active_iter_1_added_20" # Example branch name
save_model_to_branch(model, REPO_NAME, branch_name)
To load an intermediate model from a specific branch:
loaded_model = load_model_from_branch(REPO_NAME, "active_iter_1_added_20")
Preliminary Experiment:
Run the preliminary threshold experiment (with k-fold cross-validation) to determine the optimal percentage of expert data to start with. For instance, if the analysis indicates that adding 7% of expert examples provides a stable improvement, use that as your baseline for active learning.
Initialize with Expert Data:
Fine-tune the base model (model_llm_pure) with the selected percentage (e.g., 7%) to produce model_init_12 (or a similar variant). Save this model in a dedicated branch (e.g., model_percentage_12).
Active Learning Loop:
Start the active learning loop with the pre-trained model_init_12 (by setting use_initial_training=False) and iteratively add batches of expert examples selected by the uncertainty estimation module.
Graph Analysis:
After the active learning loop completes, plot the graph of F1-score vs. the total number of added expert examples. This graph illustrates the improvement (or saturation) of the model as more high-quality data is incorporated.
This repository documents a complete active learning workflow for NER. Our approach includes:
By following this workflow, one can observe the improvement in model performance (primarily measured by entity-level F1-score) as additional expert data is added. The saved intermediate models allow for comprehensive analysis and comparison.