Downloads · 30 days
38
15% of all-time downloads
Kiria-Nozan/ApexOracle
ApexOracle is a feature extraction model from Kiria-Nozan. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as mit.
This repository publishes the frozen molecule encoder used by ApexOracle for downstream embedding extraction. It is not the DLM pretraining repository and does not contain the guided molecule-generation pipeline.
Downloads · 30 days
38
15% of all-time downloads
All-time downloads
251
Public
Parameters
97.2M
389 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors389 MB · 100%
From the Hugging Face model README
This repository publishes the frozen molecule encoder used by ApexOracle for downstream embedding extraction. It is not the DLM pretraining repository and does not contain the guided molecule-generation pipeline.
The release contains a 12-block, 768-hidden-size diffusion transformer and the
SELFIES tokenizer files needed to reproduce its token-level hidden states.
The returned tensor includes tokenizer special-token positions. Use
attention_mask when pooling or selecting valid positions.
The runtime requires a CUDA GPU and FlashAttention:
pip install -r requirements.txt
git clone https://huggingface.co/Kiria-Nozan/ApexOracle
cd ApexOracle
python example.py
import torch
from transformers import AutoTokenizer
from DLM_emb_model import MolEmbDLM
model_dir = "Kiria-Nozan/ApexOracle"
device = torch.device("cuda")
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = MolEmbDLM.from_pretrained(model_dir).eval().to(device)
batch = tokenizer(
["[C] [C] [O]", "[C] [=C] [C] [=C] [C] [=C] [Ring1] [=Branch1]"],
padding=True,
truncation=False,
return_tensors="pt",
).to(device)
with torch.no_grad():
hidden_states = model(**batch)
print(hidden_states.shape) # [batch, padded_sequence_length, 768]
attention_mask may be the ordinary integer mask returned by Transformers;
the wrapper validates and converts it to the boolean mask required by the
non-padding FlashAttention backbone. A complete tokenizer batch, including
token_type_ids, can be passed directly with model(**batch).
b472f7508aaf0fdab4c935caf221415b48a5f8afd4d104a731c9d72d410c2c44ibm-research/materials.selfies-ted, audited revision
55e83392264cb998f7aa5014847df29868aefeb8The ApexOracle wrapper and frozen weights are released under the MIT License.
The attributed MDLM runtime and IBM tokenizer assets retain their Apache-2.0
terms; see THIRD_PARTY_NOTICES.md and LICENSES/Apache-2.0.txt.
@article{leng2025predicting,
title={Predicting and generating antibiotics against future pathogens with ApexOracle},
author={Leng, Tianang and Wan, Fangping and Torres, Marcelo Der Torossian and de la Fuente-Nunez, Cesar},
journal={arXiv preprint arXiv:2507.07862},
year={2025}
}