Downloads · 30 days
15.7K
100% of all-time downloads
FunAudioLLM/Fun-ASR-Nano-2512-vllm
Fun-ASR-Nano-2512-vllm is a automatic speech recognition model from FunAudioLLM. Use it when you need speech turned into text. It is set up for vllm. The card lists the license as apache-2.0.
This is the official vLLM-native packaging of FunAudioLLM/Fun-ASR-Nano-2512. It preserves the official checkpoint tensors and adds the layout required by vLLM's FunASRForConditionalGeneration implementation.
Downloads · 30 days
15.7K
100% of all-time downloads
All-time downloads
15.7K
Public
Parameters
985M
2 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors2 GB · 99%
From the Hugging Face model README
This is the official vLLM-native packaging of
FunAudioLLM/Fun-ASR-Nano-2512.
It preserves the official checkpoint tensors and adds the layout required by
vLLM's FunASRForConditionalGeneration implementation.
This repository does not define a new model and does not add LoRA weights. The
1,261 tensors in model.safetensors are bitwise equal to the tensors in the
official source model.pt at revision
272c57b82523ada6fd87095e955f8e29100979ab.
The validated path uses float32 for the highest transcription fidelity:
python -m pip install "vllm==0.27.1"
vllm serve FunAudioLLM/Fun-ASR-Nano-2512-vllm \
--revision vllm-0.27.1-20260830 \
--served-model-name fun-asr-nano \
--dtype float32 \
--gpu-memory-utilization 0.40 \
--enforce-eager
Send an OpenAI-compatible transcription request:
curl -fL \
https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-vllm/resolve/vllm-0.27.1-20260830/example/zh.mp3 \
-o zh.mp3
curl -sS http://127.0.0.1:8000/v1/audio/transcriptions \
-F [email protected] \
-F model=fun-asr-nano \
-F language=zh \
-F temperature=0 \
-F response_format=json
Expected text for the pinned sample:
开饭时间早上九点至下午五点。
| Item | Value |
|---|---|
| Official source | FunAudioLLM/Fun-ASR-Nano-2512 |
| Source revision | 272c57b82523ada6fd87095e955f8e29100979ab |
Source model.pt SHA-256 | 55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb |
| Tensor count | 1,261 |
| LoRA tensors | 0 |
Converted model.safetensors SHA-256 | 96dfbec48282dd24d3334369a01e9e909f321ee39a1b0003c528c5379f68c1a6 |
Sample example/zh.mp3 SHA-256 | 0e64de19e4ff9a02e682955c9112f32d2317cfdbb5bc2f3504664044c993f195 |
The complete machine-readable record is in
MODEL_PROVENANCE.json. The conversion can be
reproduced with convert_from_official.py.
The vLLM-native layout originated in the community work by
allendou/Fun-ASR-Nano-2512-vllm
and vLLM PR #33247, with
subsequent format and initialization fixes in vLLM PRs
#36108 and
#44215. This official
packaging keeps that attribution while anchoring the weights to the official
FunAudioLLM checkpoint.
The published evidence covers vLLM 0.27.1, PyTorch 2.13.0+cu129, Transformers 5.15.0, and one NVIDIA H100 80 GB GPU. The pinned Chinese sample returned the expected text in three consecutive deterministic requests. Other vLLM releases, accelerators, quantizations, and model quality across broader datasets require separate validation.
For the regular FunASR Python runtime, timestamps, speaker diarization, and
streaming services, use the canonical
modelscope/FunASR toolkit and the
original checkpoint.
The official source model declares Apache License 2.0. See
LICENSE. Third-party software such as vLLM remains subject to
its own license.