Downloads · 30 days
18
39% of all-time downloads
ProCreations/Booper-Big-Chat-INT8
Booper-Big-Chat-INT8 is a text generation model from ProCreations. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is the real INT8-on-disk export of ProCreations/Booper-Big-Chat. All matrix weights, including the 3-D MoE expert tensors, use symmetric per-output-channel INT8; norms remain BF16. The weight artifact is 164.7 MB…
Downloads · 30 days
18
39% of all-time downloads
All-time downloads
46
Public
Repo size
165 MB
Likes
0
Public
Click a slice to open those files.
.safetensors165 MB · 99%
From the Hugging Face model README
This is the real INT8-on-disk export of ProCreations/Booper-Big-Chat. All matrix weights,
including the 3-D MoE expert tensors, use symmetric per-output-channel INT8; norms remain BF16.
The weight artifact is 164.7 MB, a
2.00× reduction from the BF16 safetensors file.
Because native Transformers quantizers do not currently wrap Mixtral's 3-D expert tensors,
load_int8.py is included. It reconstructs the standard Mixtral model in BF16 for maximum
compatibility while retaining a compact downloadable INT8 artifact:
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("ProCreations/Booper-Big-Chat-INT8")
sys.path.insert(0, path)
from load_int8 import load_model
model = load_model(path, device="cuda")