Skip to content

ngoctham

VocalParse

ngoctham/VocalParse

VocalParse is a automatic speech recognition model from ngoctham. Use it when you need speech turned into text. The card lists the license as apache-2.0.

VocalParse is a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Fine-tuned from Qwen3-ASR-1.7B, it transcribes singing audio into a structured autoregressive token seque…

Downloads · 30 days

4

15% of all-time downloads

All-time downloads

27

Public

Parameters

2B

4.1 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors4.1 GB · 100%

At a glance

Task
Automatic Speech Recognition
License
apache-2.0
Model type
qwen3_asr
Access
Public
Created
Aug 5, 2026
Updated
Aug 5, 2026
SHA
fe4d6e94

Base models

Task
Automatic Speech Recognition
Type
qwen3_asr
License
apache-2.0
Languages
zh
Created
Aug 5, 2026
Updated
Aug 5, 2026
VocalParse — AI Model — AIMarketly