Downloads · 30 days
27
42% of all-time downloads
rtsh13/epigenetics-slm
epigenetics-slm is a machine learning model from rtsh13. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as llama3.2.
A Llama 3.2 1B Instruct model fine-tuned via QLoRA to generate five-category epigenetic health assessments from wearable/biomarker data, grounded in a Bio-RAG evidence retrieval pipeline.
Downloads · 30 days
27
42% of all-time downloads
All-time downloads
64
Public
Repo size
2.9 GB
Likes
0
Public
Click a slice to open those files.
.pt1.4 GB · 43%
From the Hugging Face model README
A Llama 3.2 1B Instruct model fine-tuned via QLoRA to generate five-category epigenetic health assessments from wearable/biomarker data, grounded in a Bio-RAG evidence retrieval pipeline.
Given a patient's biomarkers (HbA1c, NLR, circadian rest-activity metrics, sleep architecture, CosinorAge acceleration) and retrieved evidence chunks, the model produces a structured report with five sections: AGING, STRESS, METABOLISM, INFLAMMATION, SLEEP — each citing the evidence it was given.
| File | Description |
|---|---|
slm.q4_k_m.gguf | Quantized model (q4_k_m), ~771MB, for CPU inference via llama-cpp-python |
slm_lora/ | Full LoRA adapter + tokenizer + all training checkpoints (100–1491 steps), with optimizer/scheduler state for resuming training |
chroma_db/ | Populated Bio-RAG vector store (47 chunks, all-MiniLM-L6-v2 embeddings) — required alongside the model for grounded generation |
unsloth/llama-3.2-1b-instruct-unsloth-bnb-4bitSFTTrainer| Metric | Score |
|---|---|
| Category coverage (all 5 headers present) | 99.8% |
| Classification match rate | 98.5% |
| ROUGE-L | 0.822 |
These metrics check structural adherence (all five sections present) and classification-label accuracy against the training data's target responses. They do not measure citation faithfulness — see Limitations.
Requires the model to be prompted via the exact training-time template
(see slm_prompt.py in the source repo
for the canonical build_prompt()/build_inference_prompt() functions —
byte-identical prompt formatting between training and inference is
required for output quality).
from llama_cpp import Llama
llm = Llama(model_path="slm.q4_k_m.gguf", n_ctx=4096, verbose=False)
# prompt must be built via build_prompt() + build_inference_prompt()
# from the source repo — see link above
out = llm(prompt, max_tokens=512, temperature=0.2, stop=["<|eot_id|>"])
print(out["choices"][0]["text"])
For full end-to-end usage (XGBoost baseline + Bio-RAG retrieval + this
SLM), see scripts/demo.py in the source repo.
Training, export, and evaluation pipeline:
github.com/rtsh13/epigenetics-slm
(branch feat/week7-slm-finetuning)