Downloads · 30 days
3
21% of all-time downloads
dhanasekharB/Zephyr-7B-quantized
Zephyr-7B-quantized is a machine learning model from dhanasekharB. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
3
21% of all-time downloads
All-time downloads
14
Public
Parameters
7.5B
4.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
How the weights are stored.
U87.2B · 96%
From the Hugging Face model README
zephyr-7b-beta-quantizeddhanasekharB/Zephyr-7B-quantized (Quantized)bitsandbytesThis quantized version of HuggingFaceH4/zephyr-7b-beta is optimized for low-memory usage and fast inference on CPUs. By reducing the model size through 4-bit quantization, it’s suitable for applications that require text generation on resource-constrained devices.
Quantization was performed using the bitsandbytes library, specifically with 4-bit NormalFloat quantization (nf4). This model supports tasks like text generation and interactive dialogue, and is tailored for efficient, on-device deployment.
This model is loaded with quantization configuration, making it efficient for CPU-based inference. Below is an example to load and run it for inference.
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load model and tokenizer
model_ckpoint = "dhanasekharB/Zephyr-7B-quantized"
tokenizer = AutoTokenizer.from_pretrained(model_ckpoint)
model = AutoModelForCausalLM.from_pretrained(model_ckpoint)
# Sample inference
my_text = "Hi, how are you?"
inputs = tokenizer(my_text, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=10)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
bitsandbytesbnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4"The base model, zephyr-7b-beta, was trained on diverse text data, but this quantized version was not further fine-tuned. Quantization was applied to adapt the model for CPU inference.
This model should not be used for generating harmful, biased, or misleading information. Users should monitor outputs, especially for public-facing applications, to ensure content aligns with ethical guidelines.
The base model was developed by [Model Original Authors/Organization]. Special thanks to the bitsandbytes library for making efficient quantization possible.
If you use this model, please cite it as follows:
@misc{your_model_name,
title={Quantized zephyr-7b-beta},
author={Your Name/Organization},
year={2024},
url={your Hugging Face Hub URL}
}