Downloads · 30 days
16
27% of all-time downloads
lunovian/Qwen2.5-Math-7B-Instruct-4bit
Qwen2.5-Math-7B-Instruct-4bit is a machine learning model from lunovian. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Qwen2.5-Math-7B-Instruct-4bit is a 4-bit quantized version of the Qwen/Qwen2.5-Math-7B-Instruct model using GPTQ quantization (W4A16 - 4-bit weights, 16-bit activations).
Downloads · 30 days
16
27% of all-time downloads
All-time downloads
59
Public
Parameters
2B
5.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.5 GB · 100%
How the weights are stored.
I326.5B · 85%
From the Hugging Face model README
Qwen2.5-Math-7B-Instruct-4bit is a 4-bit quantized version of the Qwen/Qwen2.5-Math-7B-Instruct model using GPTQ quantization (W4A16 - 4-bit weights, 16-bit activations).
This model is optimized to:
This model is designed for direct use in mathematical and reasoning tasks, including:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "your-username/qwen2.5-math-7b-instruct-4bit"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
device_map="auto",
dtype="float16",
trust_remote_code=True,
low_cpu_mem_usage=False, # Important for compressed models
)
# Create prompt
prompt = "<|im_start|>user\nSolve for x: 3x + 5 = 14<|im_end|>\n<|im_start|>assistant\n"
# Generate
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200, do_sample=False)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
This model can be further fine-tuned for specific mathematical tasks or integrated into educational applications.
This model is NOT designed for:
Users should:
pip install transformers torch accelerate
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "your-username/qwen2.5-math-7b-instruct-4bit"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name,
device_map="auto",
dtype="float16",
trust_remote_code=True,
low_cpu_mem_usage=False,
)
# Use the model
prompt = "<|im_start|>user\nWhat is 2+2?<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The model was quantized using:
lm_headThe model was evaluated on the GSM8K test set.
The compressed model maintains high accuracy for mathematical tasks while significantly reducing size and memory requirements.
If you use this model, please cite:
Base Model:
@article{qwen2.5,
title={Qwen2.5: A Large Language Model for Mathematics},
author={Qwen Team},
year={2024}
}
Quantization Method:
@article{gptq,
title={GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers},
author={Frantar, Elias and Ashkboos, Saleh and Hoefler, Torsten and Alistarh, Dan},
journal={arXiv preprint arXiv:2210.17323},
year={2022}
}
To report issues or ask questions, please open an issue on the repository.