Downloads · 30 days
0
K024/ChatGLM-6b-onnx-u8s8
ChatGLM-6b-onnx-u8s8 is a machine learning model from K024. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model is exported from ChatGLM-6b with int8 quantization and optimized for ONNXRuntime inference. Export code in this repo.
Downloads · 30 days
0
Access
Public
Updated May 16, 2023
Repo size
6.7 GB
Likes
11
Public
Click a slice to open those files.
.bin6.7 GB · 100%
From the Hugging Face model README
This model is exported from ChatGLM-6b with int8 quantization and optimized for ONNXRuntime inference. Export code in this repo.
Inference code with ONNXRuntime is uploaded with the model. Install requirements and run streamlit run web-ui.py to start chatting. Currently the MatMulInteger (for u8s8 data type) and DynamicQuantizeLinear operators are only supported on CPU. Arm64 with Neon support (Apple M1/M2) should be reasonably fast.
安装依赖并运行 streamlit run web-ui.py 预览模型效果。由于 ONNXRuntime 算子支持问题,目前仅能够使用 CPU 进行推理,在 Arm64 (Apple M1/M2) 上有可观的速度。具体的 ONNX 导出代码在这个仓库中。
Clone with git-lfs:
git lfs clone https://huggingface.co/K024/ChatGLM-6b-onnx-u8s8
cd ChatGLM-6b-onnx-u8s8
pip install -r requirements.txt
streamlit run web-ui.py
Or use huggingface_hub python client lib to download the repo snapshot:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="K024/ChatGLM-6b-onnx-u8s8", local_dir="./ChatGLM-6b-onnx-u8s8")
Codes are released under MIT license.
Model weights are released under the same license as ChatGLM-6b, see MODEL LICENSE.