Downloads · 30 days
82
9% of all-time downloads
kousw/bitnet_b1_58-3B_quantized
bitnet_b1_58-3B_quantized is a text generation model from kousw. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This repository contains a quantized version of the 1bitLLM/bitnetb158-3B model. While the original repository showcases impressive validation results, it emulates BitNet's Linear layers, resulting in memory usage sim…
Downloads · 30 days
82
9% of all-time downloads
All-time downloads
904
Public
Parameters
332M
1.1 GB on disk
Likes
12
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
How the weights are stored.
I32204M · 61%
From the Hugging Face model README
This repository contains a quantized version of the 1bitLLM/bitnet_b1_58-3B model. While the original repository showcases impressive validation results, it emulates BitNet's Linear layers, resulting in memory usage similar to fp16 models. By leveraging the QuantLinear module from AutoGPTQ, this repository enables the output and execution of a 2-bit quantized model.
The quantized model offers significant advantages in terms of model size and memory consumption. With a model size of just 1GB , the quantized 3B model can perform inference with a context size of 2048 while consuming only 4.5GB of VRAM. Furthermore, since the weights used during execution are the same as the original repository, the perplexity (PPL) output remains unchanged.
pip install -r requirements.txt
The quantized model is already provided in this repository. However, if you wish to quantize the model yourself, you can load it from 1bitLLM/bitnet_b1_58-3B and save the quantized version (2-bit) to ./bitnet_b1_58-3B_quantized by running the following command:
python quantization.py
python eval_ppl.py --hf_path ./ --seqlen 2048 --max_dataset_size 1000
python eval_task.py --hf_path ./ \
--batch_size 1 \
--tasks \
--output_path result.json \
--num_fewshot 0 \
--ctx_size 2048