Downloads · 30 days
67
15% of all-time downloads
qvac/genesis-i-model
genesis-i-model is a text generation model from qvac. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
- Pretrained on the Largest Synthetic Educational Dataset This model has been pretrained on Tether's QVAC Genesis I, the largest synthetic dataset released for educational LLM pre-training.
Downloads · 30 days
67
15% of all-time downloads
All-time downloads
442
Public
Parameters
1.7B
3.5 GB on disk
Likes
11
Public
Click a slice to open those files.
.safetensors3.4 GB · 100%
From the Hugging Face model README
Pretrained on the Largest Synthetic Educational Dataset
This model has been pretrained on Tether's QVAC Genesis I, the largest synthetic dataset released for educational LLM pre-training.
The model was trained from scratch on approximately 40B tokens of multi-domain educational text, using BF16 mixed precision and a 4,096-token context window. Training was made with a Qwen3-family 1.7B-parameter decoder-only transformer architecture.
Checkpoints are provided in standard Hugging Face format for easy inference, continual pre-training, and fine-tuning.
Multi-Domain Educational Coverage
Because the model is trained on QVAC Genesis I, it inherits curriculum-aligned coverage across:
Superior Benchmark Performance
Leveraging QVAC Genesis I as its training foundation, the model consistently outperforms baselines in:
First Publicly Released Education-Specific Pretrained Model
This is the first open-source pretrained model built directly on a rigorously validated synthetic dataset for education, offering deep and comprehensive STEM coverage.
abilities
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "qvac/genesisI-model"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16, # trained with BF16 mixed precision
device_map="auto"
)
prompt = "Explain precision vs. recall in one paragraph."
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
top_p=0.9,
temperature=0.7
)
print(tok.decode(out[0], skip_special_tokens=True))
Tip: On consumer GPUs, consider loading in float16 or using 4/8-bit quantization (e.g., bitsandbytes/AutoGPTQ).
4 × 8 × 480 = 15,360 samples/steptorch.compile (safe mode) after warmup stability.max_split_size_mb=512, expandable segments, GC threshold ~0.8).Cluster: ~60 nodes, each 8× NVIDIA H100 80GB (total 480 GPUs), ~800 GB RAM/node.
Scheduler: Slurm (priority partition, exclusive allocation, 72-hour limit).
Launch: srun + PyTorch DDP (world size 480; ranks bound via Slurm env).
Storage: Sharded checkpoints; periodic saves for robust resume.
Networking: NCCL over InfiniBand with UCX
NCCL_IB_DISABLE=0, NCCL_IB_HCA="mlx5*", NCCL_SOCKET_IFNAME=<ib0/enoX>, NCCL_BLOCKING_WAIT=1I/O: Async dataset prefetching; pinned FS threads.
Observability: W&B + structured logs (throughput, TFLOPs/GPU, mem, step time).
Reproducibility: Fixed seeds; exact launch scripts/env logged; effective tokens/step reported.
Final checkpoint converted to Hugging Face format for plug-and-play inference.
Suggested suite (edit as applicable):
Hardware
Software
# Slurm (illustrative)
srun -N 60 -n 480 --ntasks-per-node=8 --gpus-per-task=1 \
--cpus-per-task=8 --mem=0 \
bash -lc '
export NCCL_IB_DISABLE=0
export NCCL_IB_HCA="mlx5*"
export NCCL_SOCKET_IFNAME=ib0
export NCCL_BLOCKING_WAIT=1
export TORCH_DISTRIBUTED_DEBUG=DETAIL
python train.py \
--model qwen3_1p7b_from_scratch \
--tokenizer qwen3 \
--data_path /path/to/arrow \
--context_length 4096 \
--optimizer adamw --weight_decay 0.01 \
--lr 2e-4 --warmup_steps 600 \
--precision bf16-mixed \
--micro_batch_size 4 \
--grad_accum_steps 8 \
--eval_every 500 --log_every 50 \
--ckpt_every 1000 \
--activation_checkpointing \
--flash_attn 2 \
--compile safe \
--seed 42
'
AutoModelForCausalLM.safetensors for integrity.