Downloads · 30 days
6
26% of all-time downloads
Luigi/asr-468m-apache-base
asr-468m-apache-base is a automatic speech recognition model from Luigi. Use it when you need speech turned into text. The card lists the license as apache-2.0.
A 7-language (Chinese, English, French, German, Japanese, Korean, Cantonese) speech-to-text model distilled from Qwen3-ASR-0.6B (Apache-2.0). This is the Stage-1 base checkpoint — the quality target reached before any…
Downloads · 30 days
6
26% of all-time downloads
All-time downloads
23
Public
Parameters
468M
1.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.9 GB · 99%
From the Hugging Face model README
A 7-language (Chinese, English, French, German, Japanese, Korean, Cantonese) speech-to-text model distilled from Qwen3-ASR-0.6B (Apache-2.0). This is the Stage-1 base checkpoint — the quality target reached before any parameter compression, at 467.81M parameters. It statistically ties Audio8-ASR-0.1B (macro 15.36 vs 15.31) at ~1.4x its size, fully Apache-2.0 where Audio8 is CC-BY-NC and unusable commercially.
If you want the smaller, size-matched release (323.77M, Audio8's exact parameter budget, at a
quality cost — see its model card for the honest tradeoff), use
Luigi/asr-324m-apache instead. This base
checkpoint is also the required starting point for reproducing that model's compression pipeline.
Code, full training pipeline, and every finding: github.com/vieenrose/asr-324m-apache.
15.36 vs Audio8-ASR-0.1B's 15.31 — a statistical tie.
Unlike the 324M release, this checkpoint's vocabulary is not pruned (full 151,936-id Qwen3 tokenizer), so no id remapping is needed.
import torch
from qwen_asr.core.transformers_backend.modeling_qwen3_asr import Qwen3ASRForConditionalGeneration
from qwen_asr.core.transformers_backend.processing_qwen3_asr import Qwen3ASRProcessor
path = "Luigi/asr-468m-apache-base"
proc = Qwen3ASRProcessor.from_pretrained(path)
model = Qwen3ASRForConditionalGeneration.from_pretrained(path, dtype=torch.bfloat16).cuda().eval()
NATIVE = ("<|im_start|>system\n<|im_end|>\n<|im_start|>user\n<|audio_pad|><|im_end|>\n"
"<|im_start|>assistant\n")
def transcribe(wav_16k_float32, language="Chinese", max_new_tokens=128):
e = proc(text=NATIVE + f"language {language}<asr_text>", audio=[wav_16k_float32],
sampling_rate=16000, return_tensors="pt")
e = {k: (v.cuda() if torch.is_tensor(v) else v) for k, v in e.items()}
if "input_features" in e:
e["input_features"] = e["input_features"].to(torch.bfloat16)
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
out = model.generate(**e, max_new_tokens=max_new_tokens, do_sample=False)
ids = out[0][e["input_ids"].shape[1]:].tolist()
return proc.tokenizer.decode(ids, skip_special_tokens=True)
language accepts: Chinese, English, French, German, Japanese, Korean, Cantonese.
Trained on Common Voice 17 (CC0), WenetSpeech4TTS, Multilingual LibriSpeech, LibriSpeech, and FLEURS (all CC-BY-4.0). This model was trained in part on WenetSpeech4TTS, Multilingual LibriSpeech, LibriSpeech, and FLEURS, each licensed CC-BY-4.0 by their respective creators.
Audio8-ASR-0.1B is used only as a measurement reference throughout — its weights are never loaded, merged, or distilled from.