Downloads · 30 days
0
tron-mainz/extra_trees.indel
extra_trees.indel is a machine learning model from tron-mainz. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-nd-4.0.
This is an extra trees model for sensitive detection of somatic indel candidates.
Downloads · 30 days
0
Access
Public
Updated May 5, 2025
Repo size
8.5 MB
Likes
0
Public
Click a slice to open those files.
.joblib8.5 MB · 100%
From the Hugging Face model README
This is an extra trees model for sensitive detection of somatic indel candidates.
Using matched tumor-normal paired short read sequencing data, you can call a sensitive list of somatic small indels.
You can extract features from a matched tumor-normal sequencing pair using https://github.com/TRON-Bioinformatics/tronflow-vcf-postprocessing and use this model on its output. Specific features this model utilizes are:
This model is a part of VariantMedium somatic variant caller and is integrated directly into its workflow https://github.com/TRON-Bioinformatics/VariantMedium.
The model on its own is not intended to create a final list of variant calls, it is intended for filtering out noticeable false positives.
The model is trained on the output of cell line WES and an AML WGS, both originating from short read Illumina sequencing. It is tested for other cancer entities and for solid tumors, but is not tested for non-Illumina sequencing.
We recommend using this model for Illumina-based WES and WGS (paired, short read).
Matched tumor-normal sequencing published under https://ega-archive.org/studies/EGAS00001007633 and https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000159.v13.p5.
scikit-learn GridSearchCV. Hyperparameters are given under the relevant section.
Given matched tumor-normal BAM files:
hyperparams = [
{
'n_estimators': [100, 200],
'max_depth': [5, 10],
'criterion': ['entropy'],
'max_features': ['sqrt', 'log2'],
'bootstrap': [True, False]
},
{
'n_estimators': [300, 400],
'max_depth': [10, 15],
'criterion': ['entropy'],
'max_features': ['sqrt', 'log2'],
'bootstrap': [True, False]
}
]
Evaluation using CV and a left out cell line.
Tested on independent data sets:
Sensitivity and precision, with sensitivity as the primary metric since the aim was to filter out noticeable false positives instead of coming up with a final list of variants.
| Metric | Test Set |
|---|---|
| Precision | 0.8564 |
| Recall | 0.6048 |
| F1 Score | 0.7089 |
Test set recall dropped slightly compared to control (0.6167 -> 0.6048), but precision increased (0.8216 -> 0.8564).