Downloads · 30 days
25
0% of all-time downloads
SweatyCrayfish/llama-3-8b-quantized
llama-3-8b-quantized is a text generation model from SweatyCrayfish. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
This repository hosts the 4-bit quantized version of the Llama 3 model. Optimized for reduced memory usage and faster inference, this model is suitable for deployment in environments where computational resources are…
Downloads · 30 days
25
0% of all-time downloads
All-time downloads
67.4K
Public
Parameters
8B
16.1 GB on disk
Likes
12
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This repository hosts the 4-bit quantized version of the Llama 3 model. Optimized for reduced memory usage and faster inference, this model is suitable for deployment in environments where computational resources are limited.
To utilize this model efficiently, follow the steps below:
Load the model with specific parameters to ensure it utilizes 4-bit precision:
from transformers import AutoModelForCausalLM
model_4bit = AutoModelForCausalLM.from_pretrained("SweatyCrayfish/llama-3-8b-quantized", device_map="auto", load_in_4bit=True)
Adjust the precision of other components, which are by default converted to torch.float16:
import torch
from transformers import AutoModelForCausalLM
model_4bit = AutoModelForCausalLM.from_pretrained("SweatyCrayfish/llama-3-8b-quantized", load_in_4bit=True, torch_dtype=torch.float32)
print(model_4bit.model.decoder.layers[-1].final_layer_norm.weight.dtype)
Original repository and citations: @article{llama3modelcard, title={Llama 3 Model Card}, author={AI@Meta}, year={2024}, url = {https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md} }