Downloads · 30 days
4
15% of all-time downloads
ngoctham/VocalParse
VocalParse is a automatic speech recognition model from ngoctham. Use it when you need speech turned into text. The card lists the license as apache-2.0.
VocalParse is a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Fine-tuned from Qwen3-ASR-1.7B, it transcribes singing audio into a structured autoregressive token seque…
Downloads · 30 days
4
15% of all-time downloads
All-time downloads
27
Public
Parameters
2B
4.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
From the Hugging Face model README
VocalParse is a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Fine-tuned from Qwen3-ASR-1.7B, it transcribes singing audio into a structured autoregressive token sequence that jointly encodes lyrics, pitch, note values, and global tempo (BPM).
Singing Audio (16kHz) → Whisper Encoder → Qwen LLM Decoder → AST Token Sequence
感 <P_68> <NOTE_4> 受 <P_60> <NOTE_8> 到 <P_65> <NOTE_8> ... <BPM_89>
It is recommended to use uv for setup:
uv venv --python 3.10
source .venv/bin/activate
uv pip install torch torchaudio --index-url https://download.pytorch.org/whl/cu124
uv pip install git+https://github.com/pymaster17/VocalParse.git
from vocalparse import transcribe_one
text = transcribe_one(
audio="path/to/song.wav",
checkpoint="pymaster/VocalParse",
)
print(text)
# Example output: 感 <P_68> <NOTE_4> 受 <P_60> <NOTE_8> ... <BPM_89>
| Property | Value |
|---|---|
| Base model | Qwen3-ASR-1.7B (Whisper encoder + Qwen LLM decoder) |
| Fine-tuning task | Automatic Singing Transcription (AST) |
| Training mode | CoT (asr_cot=true, bpm_position=last) |
| New vocabulary tokens | ~400 AST tokens (pitch, note value, BPM) |
| Input | Mono 16 kHz singing audio |
| Output | Interleaved lyric + pitch + note sequence with global BPM |
The base Qwen3-ASR vocabulary is extended with:
<P_0> – <P_127>) representing MIDI notes.<NOTE_4>, <NOTE_8>, <NOTE_DOT_8>).<BPM_0> – <BPM_255>).Standard interleaved format (bpm_position=last):
感 <P_68> <NOTE_4> 受 <P_60> <NOTE_8> 到 <P_65> <NOTE_8> ... <BPM_89>
CoT format produced during generation (asr_cot=true): the model first outputs plain lyrics, then the full interleaved score, separated by <|file_sep|>:
感受到<|file_sep|>感 <P_68> <NOTE_4> 受 <P_60> <NOTE_8> 到 <P_65> <NOTE_8> ... <BPM_89>
Metrics are computed with two-stage Needleman-Wunsch alignment: word-level alignment for lyrics, then pair-level alignment inside each matched word for pitch and note.
@article{vocalparse2026,
title = {VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models},
author = {Yukun Chen and Tianrui Wang and Zhaoxi Mu and Xinyu Yang and EngSiong Chng},
journal = {arXiv preprint arXiv:2605.04613},
year = {2026},
url = {http://arxiv.org/abs/2605.04613}
}
This model is licensed under Apache 2.0.