Downloads · 30 days
0
Raiff1982/healdette
healdette is a text classification model from Raiff1982. Use it when you need a label for a piece of text. It is set up for adapter-transformers. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Sep 27, 2025
Repo size
24.7 MB
Likes
0
Public
Click a slice to open those files.
.exe13.6 MB · 97%
From the Hugging Face model README
https://doi.org/10.5281/zenodo.17213886
A secure and flexible computational pipeline for generating and validating antibody sequences with multi-ethnic support. The pipeline integrates ProtGPT2 for sequence generation with BioPython for structural analysis and includes multi-ethnic HLA frequency data for immunogenicity assessment, with optimizations for various population-specific binding motifs.
requirements.txtgit clone https://github.com/Raiff1982/healdette.git
cd healdette
python -m venv .venv
# On Windows:
.venv\Scripts\activate
# On Unix/MacOS:
source .venv/bin/activate
pip install -r requirements.txt
Healdette now supports ancestry-weighted validation for multiple ethnic populations. The system uses:
Configuration files follow this structure:
{
"global_params": {
"sequence_length": {
"min": 40,
"max": 70
},
"structural_params": {
"helix_propensity": {
"min": 20,
"max": 50
},
"sheet_propensity": {
"min": 10,
"max": 40
}
},
"homopolymer_threshold": 4
},
"populations": {
"french_german": {
"ancestry_weight": 0.298,
"binding_motifs": ["WY", "RF", "KH", "YF"],
"biophysical_params": {
"aromatic_content": {
"min": 16,
"max": 28
},
"hydrophobic_content": {
"min": 33,
"max": 43
},
"net_charge": {
"min": 4,
"max": 13
}
},
"hla_frequencies": {
"hla_a": {},
"hla_b": {},
"hla_c": {}
}
}
}
}
The validation system considers:
Each population can define:
examples/ directory):{
"global_params": {
"sequence_length": {
"min": 40,
"max": 70
}
},
"populations": {
"french_german": {
"ancestry_weight": 0.298,
"binding_motifs": ["WY", "RF", "KH", "YF"],
"biophysical_params": {
"aromatic_content": {
"min": 16,
"max": 28
}
}
},
"finnish": {
"ancestry_weight": 0.057,
"binding_motifs": ["WH", "RF", "KY", "FF"],
"biophysical_params": {
"aromatic_content": {
"min": 14,
"max": 26
}
}
}
}
}
from modules.weighted_validator import WeightedSequenceValidator
from modules.config_validator import ConfigValidator
# Load and validate configuration
config_validator = ConfigValidator()
config = "path/to/config.json"
if config_validator.validate_file(config)['valid']:
# Create validator with ancestry-weighted parameters
validator = WeightedSequenceValidator(sequence, config)
# Get detailed validation results
results = validator.validate_sequence()
# Check population-specific scores
pop_scores = results['population_scores']
for pop, score in pop_scores.items():
print(f"{pop}: {score['score']} (weight: {score['weight']})")
python main.py config.json output.json --num-candidates 15
Complete example configurations are available in the examples/ directory:
european_populations_config.json: Configuration for European population clustersmulti_ethnic_config.json: General multi-ethnic configuration templateceltic_test_input.json: Celtic-specific test configurationThe weighted validator provides detailed results:
{
"valid": true,
"warnings": [],
"metrics": {
"aromatic_content": 22.5,
"hydrophobic_content": 38.2,
"binding_motifs": {
"scores": {
"french_german": {"score": 0.75, "weighted_score": 0.223},
"finnish": {"score": 0.5, "weighted_score": 0.029}
},
"total_score": 0.252
}
},
"population_scores": {
"french_german": {
"score": 0.8,
"weight": 0.298
},
"finnish": {
"score": 0.6,
"weight": 0.057
}
}
}
],
"num_sequences": 10,
"global_validation_params": {
"min_sequence_length": 40,
"max_sequence_length": 70,
"allow_homopolymers": false,
"structure_requirements": {
"helix_propensity": {
"min": 0.2,
"max": 0.5
},
"sheet_propensity": {
"min": 0.1,
"max": 0.4
}
}
}
}
python main.py --config input_config.json
The pipeline generates two types of output files in the output directory:
Detailed JSON output (antibody_designs_{timestamp}.json):
Summary report (antibody_summary_{timestamp}.txt):
To reproduce the results:
import torch
torch.manual_seed(42)
Ensure consistent data sources:
Run validation tests:
python -m unittest discover tests
MIT License. See LICENSE file for details.
If you use this software in your research, please cite:
@software{healdette2025,
title = {Healdette: Celtic-Optimized Antibody Generation Pipeline},
author = {Raiff, et al.},
year = {2025},
version = {1.0.0},
url = {https://github.com/Raiff1982/healdette}
}
Harrison, J. (2025). Healdette: A Population-Aware Antibody Design Pipeline. GitHub repository: https://github.com/Raiff1982/healdette
## Author
Jonathan Harrison (Raiff1982)