Downloads · 30 days
29
16% of all-time downloads
auryn-macmillan/boostedv1
boostedv1 is a text generation model from auryn-macmillan. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Continued LoRA training of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B for improved reasoning and code generation. This model is the merged output of the boostedv1train pipeline: 400 MLX LoRA steps (Apple Silicon) + 550…
Downloads · 30 days
29
16% of all-time downloads
All-time downloads
180
Public
Parameters
1.8B
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
Continued LoRA training of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B for improved reasoning and code generation. This model is the merged output of the boostedv1train pipeline: 400 MLX LoRA steps (Apple Silicon) + 550 Phase-1 continuation steps (dual RTX 3090, bf16).
| Stage | Platform | Steps | Batch | Seq len | Data |
|---|---|---|---|---|---|
| MLX run 1 | Apple Silicon | 150 | 4 | 512 | OpenCodeInstruct |
| MLX run 2 | Apple Silicon | 250 | 2 | 1024 | OpenCodeInstruct |
| Phase 1 | 2x RTX 3090 | 550 | 16 (eff) | 2048 | OpenThoughts + OpenR1-Math + OpenCodeInstruct |
| Benchmark | Score |
|---|---|
| GSM8K | 46.0% |
| HumanEval (pass@1) | 7.3% |
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('auryn-macmillan/boostedv1') tok = AutoTokenizer.from_pretrained('auryn-macmillan/boostedv1') inputs = tok('What is 2+2?', return_tensors='pt') out = model.generate(**inputs, max_new_tokens=128) print(tok.decode(out[0]))