Downloads · 30 days
273
7% of all-time downloads
AngelSlim/Hy3-GPTQ-Int4
Hy3-GPTQ-Int4 is a text generation model from AngelSlim. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
<p align="center" <picture <source media="(prefers-color-scheme: dark)" srcset="https://github.com/Tencent/AngelSlim/blob/main/docs/source/assets/logos/angelslimlogolight.png?raw=true" <img alt="AngelSlim" src="https:…
Downloads · 30 days
273
7% of all-time downloads
All-time downloads
4.2K
Public
Parameters
299B
169 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors169 GB · 100%
How the weights are stored.
I32290B · 97%
From the Hugging Face model README
We use GPTQ 4-bit quantization to compress Hy3 to ~1/4 size with minimal accuracy loss. See the benchmark below:
<p align="center"> <img src="assets/benchmark.png" width="95%"/> </p>Build vLLM from source:
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
git clone https://github.com/vllm-project/vllm.git
cd vllm
uv pip install --editable . --torch-backend=auto
Start the vLLM server:
# Switch to trtllm backend to work-around mnnvl workspace size issue.
export VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm
vllm serve AngelSlim/Hy3-GPTQ-Int4 \
--tensor-parallel-size 8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 2 \
--tool-call-parser hy_v3 \
--reasoning-parser hy_v3 \
--enable-auto-tool-choice \
--port 8000 \
--served-model-name hy3-gptq-int4