Downloads · 30 days
23
0% of all-time downloads
meghanamakkapati/MistralAI_INT4_quantization
MistralAI_INT4_quantization is a automatic speech recognition model from meghanamakkapati. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains an INT4 NF4 quantized version of:
Downloads · 30 days
23
0% of all-time downloads
All-time downloads
23.6K
Public
Parameters
4.6B
2.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.9 GB · 99%
How the weights are stored.
U84.1B · 91%
From the Hugging Face model README
This repository contains an INT4 NF4 quantized version of:
mistralai/Voxtral-Mini-4B-Realtime-2602
mistralai/Voxtral-Mini-4B-Realtime-2602The model remains based on Voxtral Realtime. The compression approach focuses on weight quantization.
The challenge evaluation is expected to focus on ASR quality using WER, followed by energy efficiency ranking among qualifying submissions.
A tiny FLEURS smoke test was run on three languages:
en_usfr_frhi_inThe smoke test used one sample per language and was intended only to verify that the quantized model loads and transcribes.
Observed smoke-test macro WER:
0.6681290.650585Because this test used only one sample per language, these numbers should not be interpreted as a final benchmark. They only indicate that the INT4 checkpoint is functional and not obviously broken.
Intended serving command:
vllm serve --config vllm_config.yaml
Important note: this checkpoint was produced as a Transformers/BitsAndBytes INT4 NF4 checkpoint. Final vLLM compatibility should be verified in the official evaluation environment.
The clean submission folder was prepared from:
/content/mistral_voxtral_quant/voxtral-mini-4b-realtime-int4-nf4-bnb
This repository should include:
config.jsonREADME.mdvllm_config.yamlSame as the original base model: Apache-2.0.