Downloads · 30 days
0
AsadIsmail/whisper-medium-ternary
whisper-medium-ternary is a automatic speech recognition model from AsadIsmail. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
Ternary-quantized version of openai/whisper-medium.
Downloads · 30 days
0
Access
Public
Updated Apr 17, 2026
Repo size
454 MB
Likes
1
Public
Click a slice to open those files.
.npz454 MB · 100%
From the Hugging Face model README
Ternary-quantized version of openai/whisper-medium.
| Property | Value |
|---|---|
| Base Model | openai/whisper-medium |
| Parameters | 769M |
| Quantization | tritplane3 (240 decoder layers) |
| Audio encoder | FP16 (preserved) |
| Stored size | 453 MB |
| FP16 size | ~3.1 GB |
| Compression | 1.30× |
from ternary_quant.inference import load_ternary_model
import torch, numpy as np
model, proc = load_ternary_model("AsadIsmail/whisper-medium-ternary", runtime_mode="cached", device="cpu")
model = model.float() # Required for encoder compat
# Transcribe audio
import soundfile as sf
audio, sr = sf.read("audio.flac")
inputs = proc(audio.astype(np.float32), sampling_rate=sr, return_tensors="pt")
inputs = {k: v.float() for k, v in inputs.items()}
with torch.no_grad():
ids = model.generate(**inputs, max_new_tokens=100)
print(proc.batch_decode(ids, skip_special_tokens=True)[0])
Part of ternary-models.