Downloads · 30 days
0
dunegym/openvino-models
openvino-models is a machine learning model from dunegym. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Pre-converted OpenVINO GenAI models, ready to use with ovtool.
Downloads · 30 days
0
Access
Public
Updated Sep 29, 2026
Repo size
192 GB
Likes
0
Public
Click a slice to open those files.
.bin199 GB · 97%
From the Hugging Face model README
Pre-converted OpenVINO GenAI models, ready to use with ovtool.
<category>/<model-name>/<format-quantization>/
llm (text generation) / vlm (vision-language multimodal) / image (diffusion image generation) / tts (text-to-speech) / embed (text embeddings) / rerank (cross-encoder rerankers)int4-asym-g128 = INT4 asymmetric, group 128; int4-sym-g128 = INT4 symmetric, group 128 (required for NPU); int8 = INT8 weight compressionEvery llm/vlm model ships the full five-variant ladder:
int4-asym-g128 / int4-sym-g128 / int4-awq-g128 / int8 / fp16.
| Model | Source (Hugging Face) | Devices |
|---|---|---|
llm/Qwen3-0.6B/* | Qwen/Qwen3-0.6B | CPU / GPU (sym verified on NPU ~21 tok/s; awq ~78 tok/s on iGPU) |
llm/Qwen3-1.7B/* | Qwen/Qwen3-1.7B | CPU / GPU (sym verified on NPU ~16 tok/s; asym ~42 tok/s on iGPU) |
llm/Qwen3-4B/* | Qwen/Qwen3-4B | CPU / GPU (sym verified on NPU ~9 tok/s; asym ~22 / awq ~24 tok/s on iGPU) |
llm/MiniCPM5-1B/* | openbmb/MiniCPM5-1B | CPU / GPU (sym verified on NPU ~21 tok/s; awq ~62 tok/s on iGPU) |
llm/MiniCPM5-2B/* | openbmb/MiniCPM5-2B | CPU / GPU (awq ~37 tok/s on iGPU; sym verified on NPU ~14 tok/s) |
vlm/Qwen3.5-0.8B/* | Qwen/Qwen3.5-0.8B | CPU / GPU (text-only) |
vlm/Qwen3.5-2B/* | Qwen/Qwen3.5-2B | CPU / GPU (text-only) |
vlm/Qwen3.5-4B/* | Qwen/Qwen3.5-4B | CPU / GPU (text-only; ~19 tok/s int4 / ~21 tok/s awq on iGPU) |
vlm/Qwen3-VL-2B-Instruct/* | Qwen/Qwen3-VL-2B-Instruct | CPU / GPU (text-only) |
vlm/Qwen3-VL-4B-Instruct/* | Qwen/Qwen3-VL-4B-Instruct | CPU / GPU (text-only; ~24 tok/s int4 on iGPU) |
vlm/Gemma-4-E2B-it/* | google/gemma-4-E2B-it | CPU / GPU — text + image verified (en/zh; ~26 tok/s CPU, ~20 tok/s GPU @0.5s TTFT); full five-variant ladder; omni model, audio not wired in GenAI; base (non-it) sibling not shipped (no chat template, not instruction-aligned) |
image/SD-Turbo/{int8,int4-g64} | stabilityai/sd-turbo | GPU (t2i / i2i; g64 required: UNet channels 320 % 128 != 0) |
image/LCM-Dreamshaper-v7/{int8,int4-g64} | SimianLuo/LCM_Dreamshaper_v7 | GPU — LCM-distilled, 4-8 steps (8-step 512px ~19s on iGPU) |
image/SD-1.5/{int8,int4-g64} | stable-diffusion-v1-5/stable-diffusion-v1-5 | GPU — classic SD1.5 base (20-step 512px ~47s on iGPU) |
image/SSD-1B/{int8,int4-g64} | segmind/SSD-1B | GPU — SDXL-distilled, native 1024px (25-step ~89s on iGPU) |
tts/SpeechT5-TTS/fp16 | microsoft/speecht5_tts | CPU (4.1s speech in 2.3s, 16 kHz; default speaker built into GenAI) / GPU |
tts/Kokoro-82M/fp16 | hexgrad/Kokoro-82M | CPU (4.7s speech in 2.7s, 24 kHz, --speaker af_heart; 54 voice packs) / GPU |
embed/BGE-small-en-v1.5/fp16 | BAAI/bge-small-en-v1.5 | CPU / GPU — 384-dim English embedder, CLS pooling auto-applied |
rerank/BGE-Reranker-v2-M3/fp16 | BAAI/bge-reranker-v2-m3 | CPU / GPU — multilingual (incl. zh) cross-encoder, sigmoid relevance scores |
embed/Qwen3-Embedding-0.6B/fp16 | Qwen/Qwen3-Embedding-0.6B | CPU / GPU — 1024-dim multilingual decoder embedder; LAST_TOKEN pooling + query instruction auto-applied by ovtool embed |
rerank/Qwen3-Reranker-0.6B/int4-asym-g128 | Qwen/Qwen3-Reranker-0.6B | CPU / GPU — converted like an LLM (convert llm); ovtool rerank auto-applies the official yes/no template (P(yes) 0.999 relevant / 0.010 irrelevant) |
pip install -e . # install ovtool
ovtool chat -m models/llm/Qwen3-0.6B/int4-sym-g128 -d NPU
ovtool image -m models/image/SD-Turbo/int8 -d GPU "a corgi surfing a wave" --steps 4
ovtool tts -m models/tts/Kokoro-82M/fp16 --speaker af_heart --language en-us "Hello!"
ovtool embed -m models/embed/BGE-small-en-v1.5/fp16 --query "what is OpenVINO?" "OpenVINO is a toolkit" "A cat video"
ovtool rerank -m models/rerank/BGE-Reranker-v2-M3/fp16 "what is OpenVINO?" "OpenVINO is an inference toolkit" "A cat video"
zh/ja voice packs ship in voices/ but are not supported end-to-end; non-English languages (es/fr-fr/hi/it/pt-br) require espeak-ngconvert llm); never feed it raw query+doc pairs — ovtool rerank applies the required instruction template automatically