Downloads · 30 days
0
baixuanzhang/Step-Audio-Tokenizer
Step-Audio-Tokenizer is a machine learning model from baixuanzhang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Step-Audio LLM is the industry’s first 130-billion parameter hu-manlike unified end-to-end model that integrates multimodal speech un-derstanding and generation capabilities, including singing voice synthesis, tool ut…
Downloads · 30 days
0
Access
Public
Updated May 16, 2026
Repo size
1.4 GB
Likes
0
Public
Click a slice to open those files.
.pt881 MB · 62%
From the Hugging Face model README
Step-Audio LLM is the industry’s first 130-billion parameter hu-manlike unified end-to-end model that integrates multimodal speech un-derstanding and generation capabilities, including singing voice synthesis, tool utilization, role-play and multilingual/dialectal comprehension and synthesis.
This repository provides the speech tokenizer component of Step-Audio LLM. For linguistic tokenization, we utilize the output from the Paraformer encoder, which is quantized into discrete representations at a token rate of 16.7 Hz. For semantic tokenization, we employ CosyVoice’s tokenizer, specifically designed to efficiently encode features essential for generating natural and expressive speech outputs, operating at a token rate of 25 Hz.
For more information, please refer to our repository: Step-Audio.