Downloads · 30 days
0
microsoft/falcon-7B-onnx
falcon-7B-onnx is a machine learning model from microsoft. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository hosts the optimized version of falcon-7b to accelerate inference with ONNX Runtime CUDA execution provider.
Downloads · 30 days
0
Access
Public
Updated Feb 2, 2024
Repo size
26.8 GB
Likes
1
Public
Click a slice to open those files.
.data14.4 GB · 100%
From the Hugging Face model README
This repository hosts the optimized version of falcon-7b to accelerate inference with ONNX Runtime CUDA execution provider.
See the usage instructions for how to inference this model with the ONNX files hosted in this repository.
Below is average latency of generating a token using a prompt of varying size using NVIDIA A100-SXM4-80GB GPU:
| Prompt Length | Batch Size | PyTorch 2.1 torch.compile | ONNX Runtime CUDA |
|---|---|---|---|
| 32 | 1 | 53.64ms | 15.68ms |
| 256 | 1 | 59.55ms | 26.05ms |
| 1024 | 1 | 89.82ms | 99.05ms |
| 2048 | 1 | 208.0ms | 227.0ms |
| 32 | 4 | 70.8ms | 19.62ms |
| 256 | 4 | 78.6ms | 81.29ms |
| 1024 | 4 | 373.7ms | 369.6ms |
| 2048 | 4 | N/A | 879.2ms |
git clone https://github.com/microsoft/onnxruntime
cd onnxruntime
python3 -m pip install -r onnxruntime/python/tools/transformers/models/llama/requirements-cuda.txt
from optimum.onnxruntime import ORTModelForCausalLM
from onnxruntime import InferenceSession
from transformers import AutoConfig, AutoTokenizer
sess = InferenceSession("falcon-7b.onnx", providers = ["CUDAExecutionProvider"])
config = AutoConfig.from_pretrained("tiiuae/falcon-7b")
model = ORTFalconForCausalLM(sess, config, use_cache = True, use_io_binding = True)
tokenizer = AutoTokenizer.from_pretrained("tiiuae/falcon-7b")
inputs = tokenizer("Instruct: What is a fermi paradox?\nOutput:", return_tensors="pt")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))