Downloads · 30 days
49
12% of all-time downloads
artindnr/mochi
mochi is a text generation model from artindnr. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
49
12% of all-time downloads
All-time downloads
395
Public
Parameters
31.2B
62.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors62.4 GB · 100%
How the weights are stored.
BF1631.2B · 100%
From the Hugging Face model README
Mochi is a math-reasoning fine-tune of GLM-4.7-Flash, trained on the Open Math Reasoning (mini) dataset — the same chain-of-thought data used in the winning submission to the AI Mathematical Olympiad Progress Prize 2 (AIMO-2) on Kaggle.
The goal of this fine-tune is to sharpen GLM-4.7-Flash's step-by-step mathematical reasoning while keeping the small, fast footprint of the Flash base model.
Looking for a quantized/local version? See mochi-gguf for GGUF builds you can run with llama.cpp, Ollama, or LM Studio.
Mochi is intended for:
It is not intended as a general-purpose reasoning or safety-critical decision-making tool. As with any LLM, verify important calculations independently.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "artindnr/mochi"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [
{"role": "user", "content": "If x^2 - 5x + 6 = 0, what are the values of x?"}
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
You can also load Mochi with Unsloth for faster inference and further fine-tuning:
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="artindnr/mochi",
max_seq_length=4096,
load_in_4bit=True,
)
Mochi was fine-tuned starting from GLM-4.7-Flash on the CoT split of unsloth/OpenMathReasoning-mini, using Unsloth for efficient LoRA/QLoRA training.
| Base model | GLM-4.7-Flash |
| Dataset | unsloth/OpenMathReasoning-mini (CoT split) |
| Task | Supervised fine-tuning (chain-of-thought math) |
| Framework | Unsloth |
If you use this model, please also credit the underlying dataset and competition it draws from:
@misc{openmathreasoning,
title = {OpenMathReasoning},
author = {NVIDIA},
year = {2025},
url = {https://huggingface.co/datasets/nvidia/OpenMathReasoning}
}