Downloads · 30 days
99
0% of all-time downloads
microsoft/bloom-deepspeed-inference-int8
bloom-deepspeed-inference-int8 is a feature extraction model from microsoft. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as bigscience-bloom-rail-1.0.
This is a custom INT8 version of the original BLOOM weights to make it fast to use with the DeepSpeed-Inference engine which uses Tensor Parallelism. In this repo the tensors are split into 8 shards to target 8 GPUs.
Downloads · 30 days
99
0% of all-time downloads
All-time downloads
36.8K
Public
Repo size
185 GB
Likes
28
Public
Click a slice to open those files.
.pt180 GB · 100%
From the Hugging Face model README
This is a custom INT8 version of the original BLOOM weights to make it fast to use with the DeepSpeed-Inference engine which uses Tensor Parallelism. In this repo the tensors are split into 8 shards to target 8 GPUs.
The full BLOOM documentation is here.
To use the weights in repo, you can adapt to your needs the scripts found here (XXX: they are going to migrate soon to HF Transformers code base, so will need to update the link once moved).