Downloads · 30 days
8
14% of all-time downloads
AI4PH/EpiMistral-7B
EpiMistral-7B is a machine learning model from AI4PH. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
EpiMistral-7B is a fine-tuned version of Open-Orca/Mistral-7B-OpenOrca specialized for extracting structured epidemiological information from unstructured disease outbreak reports. The model was trained on the WHO Dis…
Downloads · 30 days
8
14% of all-time downloads
All-time downloads
59
Public
Repo size
168 MB
Likes
0
Public
Click a slice to open those files.
.safetensors168 MB · 98%
From the Hugging Face model README
EpiMistral-7B is a fine-tuned version of Open-Orca/Mistral-7B-OpenOrca specialized for extracting structured epidemiological information from unstructured disease outbreak reports. The model was trained on the WHO Disease Outbreak News (DONs) curated database (Carlson et al., 2023) to automatically extract key epidemiological features including disease classification, geographical locations, case counts, temporal information, and outbreak characteristics.
This repository contains LoRA adapter weights only, not the full model weights.
Base Model (Mistral 7B / Open-Orca/Mistral-7B-OpenOrca): Licensed under Apache License 2.0
LoRA Adapter Weights: Released under CC0 1.0 Universal (Public Domain Dedication)
Attribution: When using this model, please include appropriate attribution:
Mistral 7B and Open-Orca/Mistral-7B-OpenOrca are licensed under the Apache License 2.0,
Copyright (c) Mistral AI and OpenOrca.
EpiMistral-7B LoRA adapter weights are released under CC0 1.0 Universal (Public Domain).
The model achieved the following results on the evaluation set:
| Metric | Score |
|---|---|
| Rouge-1 | 0.899 ± 0.046 |
| Rouge-2 | 0.853 ± 0.057 |
| Rouge-L | 0.889 ± 0.049 |
| Rouge-Lsum | 0.887 ± 0.047 |
These scores represent overall performance across 5-fold stratified cross-validation, demonstrating strong accuracy in extracting structured epidemiological information from unstructured outbreak reports.
This model is designed for:
The model extracts the following structured epidemiological information:
Disease Information:
Geographical Information:
Case Counts:
Temporal Information:
The model was trained on the WHO Disease Outbreak News curated database (Carlson et al., 2023), which contains:
The training followed an instruction-tuning paradigm where unstructured outbreak report text is paired with structured JSON output containing extracted epidemiological features. The prompt format used was:
Below is an instruction that describes a task, paired with an input that provides further context.
Write a response that appropriately completes the request.
### Instruction:
Extract disease outbreak information from the given text and format it as JSON.
Return a list containing one JSON object per outbreak mentioned.
Use "None" for missing information. Never invent or guess data.
### Input:
[Outbreak report text]
### Response:
[Extracted JSON with epidemiological features]
LoRA (Low-Rank Adaptation) Parameters:
Training Hyperparameters:
Evaluation Strategy:
Hardware:
Note: EpiMistral-7B demonstrated rapid convergence during training, reaching optimal performance within approximately 2 epochs, showcasing the efficiency of the Mistral architecture combined with the OpenOrca instruction-tuning for domain adaptation tasks.
The model uses 8-bit quantization with LoRA during training:
pip install transformers==4.52.4
pip install torch==2.3.1
pip install peft==0.12.0
pip install accelerate==1.7.0
pip install bitsandbytes==0.43.3
Note: You need access to the base Mistral-7B-OpenOrca model (available under Apache 2.0) to use these adapter weights.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
# Load base model (Apache 2.0 licensed)
base_model_id = "Open-Orca/Mistral-7B-OpenOrca"
adapter_model_id = "jrc-ai/EpiMistral-7B" # LoRA adapters
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load tokenizer from base model
tokenizer = AutoTokenizer.from_pretrained(base_model_id, trust_remote_code=True)
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
device_map="auto",
torch_dtype=torch.bfloat16,
)
# Load and apply LoRA adapters
model = PeftModel.from_pretrained(base_model, adapter_model_id)
# Example outbreak report
outbreak_text = """
WHO has reported 3 suspected cases of yellow fever in Maryland county,
in the south-eastern part of the country. One case with disease onset on
1 August has been confirmed (IgM positive) by the Institut Pasteur in
Abidjan, Côte d'Ivoire. All three cases have died.
"""
# Format prompt
prompt = f"""Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.
### Instruction:
Extract disease outbreak information from the given text and format it as JSON.
Return a list containing one JSON object per outbreak mentioned.
Always return a list of JSON objects, even for single outbreaks.
Use "None" for missing information. If no outbreak information is found, return an empty list [].
Never invent or guess data.
### Input:
{outbreak_text}
### Response:
"""
# Tokenize and generate
inputs = tokenizer(prompt, return_tensors="pt", truncation=True).to(device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=600,
temperature=0.1,
do_sample=False,
pad_token_id=tokenizer.eos_token_id
)
# Decode output
extracted_info = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(extracted_info)
[{
"DiseaseLevel1": "Yellow fever",
"DiseaseLevel2": "",
"Country": "Liberia",
"ISO": "LBR",
"OutbreakEpicenter": "Maryland county",
"CasesTotal": 3,
"CasesSuspected": 2,
"CasesProbable": null,
"CasesConfirmed": 1,
"Deaths": 3,
"OutbreakStartYear": 2001,
"OutbreakStartMonth": 8,
"OutbreakStartDay": 1,
"OutbreakDetectionYear": null,
"OutbreakDetectionMonth": null,
"OutbreakDetectionDay": null,
"OutbreakVerificationYear": null,
"OutbreakVerificationMonth": null,
"OutbreakVerificationDay": null,
"OutbreakEnd": null,
"OutbreakEndYear": null,
"OutbreakEndMonth": null,
"OutbreakEndDay": null
}]
This fine-tuned model dramatically outperforms in-context learning (iCL) approaches:
| Approach | Rouge-1 | Rouge-2 | Rouge-L | Rouge-Lsum |
|---|---|---|---|---|
| EpiMistral-7B (fine-tuned) | 0.899 | 0.853 | 0.889 | 0.887 |
| Mistral-7B-OpenOrca (16-shot iCL) | 0.600 | 0.475 | 0.580 | 0.598 |
| LLaMA 3.3-70B (16-shot iCL) | 0.840 | 0.698 | 0.824 | 0.841 |
Performance gain from fine-tuning: ~30 percentage points compared to its iCL baseline, demonstrating the substantial benefit of parameter-efficient fine-tuning for this specialized task.
| Model | Parameters | Rouge-1 | Rouge-2 | Rouge-L |
|---|---|---|---|---|
| EpiLLaMA 3.3-70B | 70B | 0.937 | 0.896 | 0.928 |
| EpiQwen 2.5-7B | 7B | 0.918 | 0.864 | 0.908 |
| EpiMistral-7B | 7B | 0.899 | 0.853 | 0.889 |
All pairwise comparisons are statistically significant (p < 0.001, Nemenyi post-hoc test with Bonferroni correction).
Key Characteristics:
If you use this model in your research, please cite:
@article{consoli2026generative,
title={Generative AI for Structured Epidemiological Information Extraction: Comparing In-Context Learning and Fine-Tuning Approaches},
author={Consoli, Sergio and Bertolini, Lorenzo and Stefanovitch, Nicolas and Spagnolo, Luigi and Espinosa, Laura and Stilianakis, Nikolaos I.},
journal={PLoS Digital Health},
volume={submitted, currently under revision},
year={2026}
}
Please also acknowledge the base models:
@article{jiang2023mistral,
title={Mistral 7B},
author={Jiang, Albert Q and Sablayrolles, Alexandre and Mensch, Arthur and Bamford, Chris and Chaplot, Devendra Singh and Casas, Diego de las and Bressand, Florian and Lengyel, Gianna and Lample, Guillaume and Saulnier, Lucile and others},
journal={arXiv preprint arXiv:2310.06825},
year={2023}
}
Upon evaluation, we identified no dual-use implications for this model. The model is designed specifically for public health surveillance and epidemic intelligence applications to support global health initiatives.
Important Notes:
We acknowledge:
Disclaimer: The views expressed are purely those of the authors and may not in any circumstance be regarded as stating an official position of the European Commission.