Downloads · 30 days
6
20% of all-time downloads
thisisaditichouhan/Phi-2-4bit-Quantized-Model
Phi-2-4bit-Quantized-Model is a text generation model from thisisaditichouhan. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
This repository contains a 4-bit quantized version of Microsoft's Phi-2 model, prepared using bitsandbytes and wrapped with a LoRA adapter via the PEFT library.
Downloads · 30 days
6
20% of all-time downloads
All-time downloads
30
Public
Repo size
10.5 MB
Likes
1
Public
Click a slice to open those files.
.safetensors10.5 MB · 75%
From the Hugging Face model README
This repository contains a 4-bit quantized version of Microsoft's Phi-2 model, prepared using bitsandbytes and wrapped with a LoRA adapter via the PEFT library.
The original Phi-2 model runs on 16-bit precision (Float16) and consumes a lot of memory. To make it highly efficient and runnable on free-tier cloud GPUs (like Google Colab T4) or local machines with limited VRAM, this model has been compressed to 4-bit using NormalFloat4 (NF4) quantization.
microsoft/phi-2bitsandbytes)BitsAndBytesConfig with load_in_4bit=True and bnb_4bit_compute_dtype=torch.float16.LoraConfig targeting the standard query/value projection layers (q_proj, v_proj).adapter_model.safetensors) and tokenizer configuration.This repository is perfect for anyone looking to experiment with lightweight text generation or perform Parameter-Efficient Fine-Tuning (PEFT) on top of a 4-bit quantized version of Phi-2.