Downloads · 30 days
73
8% of all-time downloads
gbyuvd/ChemMiniQ3-SAbRLo
ChemMiniQ3-SAbRLo is a text generation model from gbyuvd. Use it when you need the model to write or continue text. The card lists the license as mit.
ChemMiniQ3-SAbRLo is a lightweight experimental generative model for chemistry, built on mini Qwen2-like Mini Backbone, designed for rapid prototyping of HuggingFace AutoModel and AutoTokenizer compatibility, and fast…
Downloads · 30 days
73
8% of all-time downloads
All-time downloads
903
Public
Parameters
9.9M
434 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors158 MB · 100%
From the Hugging Face model README
ChemMiniQ3-SAbRLo is a lightweight experimental generative model for chemistry, built on mini Qwen2-like Mini Backbone, designed for rapid prototyping of HuggingFace AutoModel and AutoTokenizer compatibility, and fast iteration of Multi-Token Prediction (MTP) and RL fine-tuning algorithms/rewards.
It introduces a new reinforcement learning framework as the next iteration of ChemMiniQ3-HoriFIE, combining:
gbyuvd/synthaccess-chemselfies) to favor molecules that are easier to synthesize.The model can be trained on a laptop with only 2GB VRAM (NVIDIA 930M in my case), for NTP+MTP it took ~2h:40m per chunk and for RL (mix) it took ~1h for 4500 steps; and for ParetoControlled RL with mix took around 24m for 1125 steps.
Example of a generated molecule, found no identical mol in PubChem
O=C(O)CC=1CCCCC=1C2=CC=CC(=C2)NC(=O)CC3=CC=CC=C3CCC

Prototype research code — not production-ready. Built for speed, not scale - yet.
The information and model provided is for academic purposes only. It is intended for educational and research use, and should not be used for any commercial or legal purposes. The author do not guarantee the accuracy, completeness, or reliability of the information.
transformers.AutoModelForCausalLMAutoModelgbyuvd/synthaccess-chemselfies as a reward modelTrainer and PPOTrainer for easy RL experimentation💡 Note: The average SELFIES sequence length in our ~3M dataset is 33.41 ± 1.80 tokens — but for RL prototyping, we cap at 25 to accelerate training cycles and improve signal-to-noise in reward gradients.
💡 Target domain: molecular generation (SELFIES).
🔬 Goal: molecules that are valid, bioaware, and synthetically accessible.
🚀 Core innovation: fast, modular prototyping of MTP + RL fine-tuning pipelines using standard HuggingFace components.
Requirements:
datasets numpy pandas ranger21 rdkit scikit_learn selfies torch tqdm transformers
demo_usage.ipynb or download it to use (I am still learning abt HF API so please be patient.)train_withmtp.py for NTP-to-MTP trainingtrain_ppokl_withpareto.py for ParetoController adjusted rewards ratio/mixingtrain_ppokl_withsa.py with either "chemq3" (bioaware-only no SA), "sa" (SA-only no bioaware), or "mix" (combined rewards)mixusing the updated evaluate_molecular_model.py tested against chunk-4 data:
Example generated molecule (5):

...
Example 5:
Raw SELFIES : [O] [=C] [Branch1] [C] [O] [C] [C] [C] [S] [C] [C] [=C] [C] [=N] [C] [=C] [Ring1...
SMILES : O=C(O)CCCSCC1=CC=NC=C1
SA Label : Easy (confidence: 0.999)
Atoms : 14
Bonds : 14
🔍 SA Label Analysis (first 100 molecules):
Easy to synthesize: 87/100 (87%)
Hard to synthesize: 13/100 (13%)
=======================================================
📊 MOLECULAR GENERATION EVALUATION SUMMARY
=======================================================
Model Path : ./chunk-4
Generation Mode : MTP-aware
Samples Generated: 1000
-------------------------------------------------------
Validity : 0.9620 (962/1000)
Uniqueness : 0.9990 (unique valid)
Novelty (vs train): 0.9865 (space-free SELFIES)
Synthesis Labels : Easy: 886/961 (92.2%) | Hard: 75/961 (7.8%)
Internal Diversity: 0.8741 (1 - avg Tanimoto)
=======================================================
using the updated evaluate_molecular_model.py tested against chunk-4 data:
This serves as an important baseline before further adjusting the ratio of rewards; It appears lower in certain metrics but generate longer and more complex examples:

Raw SELFIES : [C] [C] [=N] [O] [C] [=C] [Ring1] [Branch1] [C] [=Branch1] [C] [=O] [N] [C@@H1] ...
SMILES : CC1=NOC=C1C(=O)N[C@@H1]2C(N)C3=CC=CC=C3CC2
SA Label : Easy (confidence: 0.999)
Atoms : 20
Bonds : 22
🔍 SA Label Analysis (first 100 molecules):
Easy to synthesize: 71/100 (71%)
Hard to synthesize: 29/100 (29%)
=======================================================
📊 MOLECULAR GENERATION EVALUATION SUMMARY
=======================================================
Model Path : ./checkpoints-1/model_step_9000
Generation Mode : MTP-aware
Samples Generated: 1000
-------------------------------------------------------
Validity : 0.9510 (951/1000)
Uniqueness : 1.0000 (unique valid)
Novelty (vs train): 0.9989 (space-free SELFIES)
Synthesis Labels : Easy: 633/951 (66.6%) | Hard: 318/951 (33.4%)
Internal Diversity: 0.8934 (1 - avg Tanimoto)
=======================================================
using the updated evaluate_molecular_model.py tested against chunk-4 data:

Example 1:
Raw SELFIES : [C] [C] [=N] [O] [C] [Branch2] [Ring1] [Branch1] [C@H1] [Branch1] [C] [C] [N] [C...
SMILES : C1C=NOC([C@H1](C)NC(=O)CNCC(F)(F)F)=N1
SA Label : Easy (confidence: 0.999)
Atoms : 18
Bonds : 18
🔍 SA Label Analysis (first 100 molecules):
Easy to synthesize: 81/100 (81%)
Hard to synthesize: 19/100 (19%)
=======================================================
📊 MOLECULAR GENERATION EVALUATION SUMMARY
=======================================================
Model Path : ./ppo_checkpoints_pareto/model_step_1125
Generation Mode : MTP-aware
Samples Generated: 1000
-------------------------------------------------------
Validity : 0.9600 (960/1000)
Uniqueness : 0.9990 (unique valid)
Novelty (vs train): 0.9875 (space-free SELFIES)
Synthesis Labels : Easy: 871/959 (90.8%) | Hard: 88/959 (9.2%)
Internal Diversity: 0.8741 (1 - avg Tanimoto)
=======================================================
We are actively working on scaling up ChemMiniQ3-SAbRLo with more ambitious experiments — all designed for rapid iteration:
pipeline(), generate(), Trainer)Training and scaling require significant computational resources.
If you’d like to support this research (e.g., helping us rent compute servers for rapid RL prototyping and MTP validation), you can contribute here:
Every bit of support helps us push ChemMiniQ3-SAbRLo further! 🚀🧬
Mix set on for 9000 stepsAutoModel and AutoTokenizer compatibility@misc{yang2024qwen2technicalreport,
title={Qwen2 Technical Report},
author={An Yang and Baosong Yang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Zhou and Chengpeng Li and Chengyuan Li and Dayiheng Liu and Fei Huang and Guanting Dong and Haoran Wei and Huan Lin and Jialong Tang and Jialin Wang and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Ma and Jianxin Yang and Jin Xu and Jingren Zhou and Jinze Bai and Jinzheng He and Junyang Lin and Kai Dang and Keming Lu and Keqin Chen and Kexin Yang and Mei Li and Mingfeng Xue and Na Ni and Pei Zhang and Peng Wang and Ru Peng and Rui Men and Ruize Gao and Runji Lin and Shijie Wang and Shuai Bai and Sinan Tan and Tianhang Zhu and Tianhao Li and Tianyu Liu and Wenbin Ge and Xiaodong Deng and Xiaohuan Zhou and Xingzhang Ren and Xinyu Zhang and Xipin Wei and Xuancheng Ren and Xuejing Liu and Yang Fan and Yang Yao and Yichang Zhang and Yu Wan and Yunfei Chu and Yuqiong Liu and Zeyu Cui and Zhenru Zhang and Zhifang Guo and Zhihao Fan},
year={2024},
eprint={2407.10671},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.10671},
}
@article{sorokina2021coconut,
title={COCONUT online: Collection of Open Natural Products database},
author={Sorokina, Maria and Merseburger, Peter and Rajan, Kohulan and Yirik, Mehmet Aziz and Steinbeck, Christoph},
journal={Journal of Cheminformatics},
volume={13},
number={1},
pages={2},
year={2021},
doi={10.1186/s13321-020-00478-9}
}
@article{zdrazil2023chembl,
title={The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods},
author={Zdrazil, Barbara and Felix, Eloy and Hunter, Fiona and Manners, Emma J and Blackshaw, James and Corbett, Sybilla and de Veij, Marleen and Ioannidis, Harris and Lopez, David Mendez and Mosquera, Juan F and Magarinos, Maria Paula and Bosc, Nicolas and Arcila, Ricardo and Kizil{\"o}ren, Tevfik and Gaulton, Anna and Bento, A Patr{\'i}cia and Adasme, Melissa F and Monecke, Peter and Landrum, Gregory A and Leach, Andrew R},
journal={Nucleic Acids Research},
year={2023},
volume={gkad1004},
doi={10.1093/nar/gkad1004}
}
@misc{chembl34,
title={ChemBL34},
year={2023},
doi={10.6019/CHEMBL.database.34}
}
@article{Gallo2023,
author = {Gallo, K and Kemmler, E and Goede, A and Becker, F and Dunkel, M and Preissner, R and Banerjee, P},
title = {{SuperNatural 3.0-a database of natural products and natural product-based derivatives}},
journal = {Nucleic Acids Research},
year = {2023},
month = jan,
day = {6},
volume = {51},
number = {D1},
pages = {D654-D659},
doi = {10.1093/nar/gkac1008}
}
@article{wright2021ranger21,
title={Ranger21: a synergistic deep learning optimizer},
author={Wright, Less and Demeure, Nestor},
year={2021},
journal={arXiv preprint arXiv:2106.13731},
}
@article{ropp2019gypsum,
title={Gypsum-DL: An Open-source Program for Preparing Small-molecule Libraries for Structure-based Virtual Screening},
author={Ropp, Patrick J. and Spiegel, Jacob O. and Walker, Jennifer L. and Green, Harrison and Morales, Guillermo A. and Milliken, Katherine A. and Ringe, John J. and Durrant, Jacob D.},
journal={Journal of Cheminformatics},
volume={11},
number={1},
year={2019},
doi={10.1186/s13321-019-0358-3}
}