Downloads · 30 days
27
49% of all-time downloads
flamiinngo/math-code-mixtral-8x7b
math-code-mixtral-8x7b is a text generation model from flamiinngo. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A LoRA fine-tune of Mixtral-8x7B-Instruct-v0.1 for mathematical word problems and short code generation, built with Adaption Labs' AutoScientist for the AutoScientist Challenge (Math & Code category).
Downloads · 30 days
27
49% of all-time downloads
All-time downloads
55
Public
Repo size
420 MB
Likes
1
Public
Click a slice to open those files.
.safetensors218 MB · 51%
From the Hugging Face model README
A LoRA fine-tune of Mixtral-8x7B-Instruct-v0.1 for mathematical word problems and short code generation, built with Adaption Labs' AutoScientist for the AutoScientist Challenge (Math & Code category).
Head-to-head win rate against the base model:
| Evaluation | Base | Adapted |
|---|---|---|
| Training distribution | 33 | 67 |
| Math category (all tasks) | 40 | 60 |
These are wins in a paired comparison, not accuracy percentages.
An earlier checkpoint trained on 2,050 rows without augmentation scored 56–44 on the category evaluation. Expanding the training set with domain and general-purpose augmentation, and raising LoRA capacity to rank 64, gained 4 points. The platform's guidance suggested a larger gain from crossing 20,000 datapoints; the observed improvement was smaller than that.
The archive unpacks flat, so give it its own directory:
mkdir -p math-adapter
tar --zstd -xf math-code-mixtral-8x7b-weights.tgz -C math-adapter
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE = "mistralai/Mixtral-8x7B-Instruct-v0.1"
ADAPTER = "./math-adapter"
tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
model = AutoModelForCausalLM.from_pretrained(
BASE, torch_dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()
messages = [{"role": "user", "content":
"Jean eats one donut per 2 pages she writes. If she writes 12 pages and each "
"donut has 150 calories, how many calories does she eat?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
Hardware: Mixtral-8x7B is a 46.7B-parameter mixture-of-experts model — roughly 94 GB in bf16, or about 25 GB in 4-bit. The adapter is 208 MB.
Output style: trained to show brief working then state the answer explicitly, ending with "The answer is X."
| Parameter | Value |
|---|---|
| Base | mistralai/Mixtral-8x7B-Instruct-v0.1 |
Rank (r) | 64 |
lora_alpha | 128 |
lora_dropout | 0.0 |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Peak learning rate | 1e-4 |
| Schedule | cosine, warmup 0.1 |
| Epochs | 5 (145 steps) |
| Epoch | Step | Eval loss |
|---|---|---|
| 1.00 | 29 | 0.7350 |
| 2.00 | 58 | 0.7026 |
| 3.00 | 87 | 0.6889 |
| 4.00 | 116 | 0.6830 |
| 5.00 | 145 | 0.6805 |
Training loss fell from 1.248 to 0.764. Evaluation loss was still declining at epoch 5, so this run did not overfit — the last epoch still gained, if marginally.
flamiinngo/math-code-qa
— 5,200 rows (3,600 math, 1,600 code), every math answer ending in a result
verified against the upstream expected_answer column. Derived from
nvidia/OpenMathInstruct-2
and sahil2801/CodeAlpaca-20k,
both CC-BY-4.0. The training run additionally used AutoScientist domain and
general-purpose augmentation on top of this base.
The adapter is Apache 2.0, matching the base model. The training data is CC-BY-4.0 and requires attribution — see the dataset card.