Downloads · 30 days
0
k050506koch/SPECTR
SPECTR is a image-text-to-text model from k050506koch. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A deep learning model for predicting chemical formulas from mass spectrometry data using a CNN-Transformer architecture.
Downloads · 30 days
0
Access
Public
Updated Nov 11, 2025
Repo size
973 MB
Likes
0
Public
Click a slice to open those files.
.pth973 MB · 100%
From the Hugging Face model README
A deep learning model for predicting chemical formulas from mass spectrometry data using a CNN-Transformer architecture.
SPECTR is a machine learning system that accepts mass spectral data (m/z peaks and intensities) and predicts the corresponding chemical formula. The model uses a convolutional neural network encoder to process spectral peaks and a transformer decoder to generate chemical formulas token by token.
The model consists of two main components:
Additional dependencies for data preparation:
git clone https://github.com/krll-corp/SPECTR.git
cd SPECTR
pip install torch numpy pandas scikit-learn wandb tqdm requests beautifulsoup4
The model expects data in JSONL format with the following structure:
{"formula": "C6H12O6", "peaks": [{"m/z": 180.063, "intensity": 999}, {"m/z": 145.050, "intensity": 450}]}
Each line contains:
formula: Chemical formula as a string (e.g., "C6H12O6")peaks: List of peak objects with m/z (mass-to-charge ratio) and intensity valuesTrain the model using the train_conv.py script:
python train_conv.py
Training options:
--resume: Resume from previous run if available--checkpoint: Path to checkpoint file (default: checkpoint_last.pt)--save-every: Save checkpoint every N steps (default: 500)The script expects a data file at ../mona_massbank_dataset.jsonl. Modify the data_file variable in the script to point to your dataset.
Training hyperparameters (configurable in the script):
Evaluate a trained model using the eval_conv3.py script:
python eval_conv3.py --checkpoint <path_to_checkpoint> --data <path_to_data> --device cuda
Options:
--checkpoint: Path to trained model checkpoint (default: checkpoint_best.pt)--data: Path to evaluation data file (default: massbank_dataset.jsonl)--device: Device to use (choices: cpu, cuda, mps)--strategy: Decoding strategy (choices: greedy, beam, top_k, top_p; default: greedy)--beam-width: Beam width for beam search (default: 3)--top-k: Top-k value for top-k sampling (default: 5)--top-p: Top-p value for nucleus sampling (default: 0.9)--limit: Limit number of samples to evaluate (optional)The data/ directory contains utilities for data collection and preparation:
crawler.py: Extract mass spectrometry data from MassBank web recordsfiltering.py: Process and filter MoNA and MassBank datasets into training formatanalyze_massbank.py: Analyze MassBank dataset statisticsExample usage for data crawling:
cd data
python crawler.py
Default configuration:
Supported chemical elements: All elements from H (Hydrogen) to Og (Oganesson)
Special tokens:
<PAD>: Padding token<SOS>: Start of sequence<EOS>: End of sequence<UNK>: Unknown tokenThe model is evaluated using:
This project is licensed under the MIT License. See the LICENSE file for details.
Copyright (c) 2025 Kyryll Kochkin
If you use SPECTR in your research, please cite this repository:
@software{spectr2025,
author = {Kochkin, Kyryll},
title = {SPECTR: Mass Spectrometry to Chemical Formula Prediction},
year = {2025},
url = {https://github.com/krll-corp/SPECTR}
}
Contributions are welcome. Please open an issue or submit a pull request for any improvements or bug fixes.
This project uses data from: