Downloads · 30 days
0
EmbeddedLLM/gemma-7b-it-onnx
gemma-7b-it-onnx is a text generation model from EmbeddedLLM. Use it when you need the model to write or continue text. The card lists the license as gemma.
This repository contains optimized versions of the gemma-7b-it model, designed to accelerate inference using ONNX Runtime. These optimizations are specifically tailored for CPU and DirectML. DirectML is a high-perform…
Downloads · 30 days
0
Access
Public
Updated Jun 20, 2024
Repo size
6.4 GB
Likes
1
Public
Click a slice to open those files.
.data6.4 GB · 100%
From the Hugging Face model README
This repository contains optimized versions of the gemma-7b-it model, designed to accelerate inference using ONNX Runtime. These optimizations are specifically tailored for CPU and DirectML. DirectML is a high-performance, hardware-accelerated DirectX 12 library for machine learning, offering GPU acceleration across a wide range of supported hardware and drivers, including those from AMD, Intel, NVIDIA, and Qualcomm.
Here are some of the optimized configurations we have added:
To use the Gemma-7B-Instruct-ONNX model on Windows with DirectML, follow these steps:
conda create -n onnx python=3.10
conda activate onnx
winget install -e --id GitHub.GitLFS
pip install huggingface-hub[cli]
huggingface-cli download EmbeddedLLM/gemma-7b-it-onnx --include="onnx/directml/gemma-7b-it-int4/*" --local-dir .\gemma-7b-it-onnx
pip install numpy==1.26.4
pip install onnxruntime-directml
pip install --pre onnxruntime-genai-directml
conda install conda-forge::vs2015_runtime
Invoke-WebRequest -Uri "https://raw.githubusercontent.com/microsoft/onnxruntime-genai/main/examples/python/phi3-qa.py" -OutFile "phi3-qa.py"
python phi3-qa.py -m .\gemma-7b-it-onnx
Minimum Configuration:
Tested Configurations:
Model Page: Gemma
This model card corresponds to the 7B instruct version of the Gemma model. You can also visit the model card of the 2B base model, 7B base model, and 2B instruct model.
Resources and Technical Documentation:
Terms of Use: Terms