Downloads · 30 days
11
30% of all-time downloads
ZhangXinWhut/BiMTokenizer
BiMTokenizer is a audio-to-audio model from ZhangXinWhut. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
BiMTokenizer is a bidirectional Mamba speech tokenizer for low-bitrate neural speech coding. This repository provides five 16 kHz checkpoints from the same model family.
Downloads · 30 days
11
30% of all-time downloads
All-time downloads
37
Public
Repo size
5 GB
Likes
0
Public
Click a slice to open those files.
.pt5 GB · 100%
From the Hugging Face model README
BiMTokenizer is a bidirectional Mamba speech tokenizer for low-bitrate neural speech coding. This repository provides five 16 kHz checkpoints from the same model family.
Xin Zhang, Lin Li, Chuanbo Liu, Jianquan Liu, and Kong Aik Lee.
| Model | Checkpoint | Codebook | Bitrate |
|---|---|---|---|
| BiMTokenizer-Whisper | whisper/bimtokenizer_whisper_librispeech.pt | 5 x 196560 | 1100 bps |
| BiMTokenizer-SenseVoice | sensevoice/bimtokenizer_sensevoice_librispeech.pt | 5 x 196560 | 1100 bps |
| BiMTokenizer-SenseVoice (32768+4096) | sensevoice-32768-4096/bimtokenizer_sensevoice_32768_4096_librispeech.pt | 32768 + 6 x 4096 | 1087.5 bps |
| BiMTokenizer-SenseVoice (8x2048) | sensevoice-2048/bimtokenizer_sensevoice_2048_librispeech.pt | 1 semantic + 7 acoustic, each 2048 | 1100 bps |
| BiMTokenizer-SenseVoice Emilia-2W (16x2048) | sensevoice-2048-emilia2w/bimtokenizer_sensevoice_2048_emilia2w_2200bps.pt | 1 semantic + 15 acoustic, each 2048 | 2200 bps |
Each checkpoint is stored next to its matching config.yaml. Do not mix a
checkpoint with a configuration from another variant.
| Model | SIM ↑ | STOI ↑ | PESQ-NB ↑ | PESQ-WB ↑ | UTMOS ↑ | WER ↓ |
|---|---|---|---|---|---|---|
| Ground Truth | 1.00 | 1.00 | 4.55 | 4.64 | 4.09 | 2.16 |
| BiMTokenizer-Whisper | 0.87 | 0.95 | 3.56 | 3.03 | 4.21 | 2.44 |
| BiMTokenizer-SenseVoice | 0.85 | 0.94 | 3.45 | 2.85 | 4.18 | 2.53 |
| BiMTokenizer-SenseVoice (32768+4096) | 0.86 | 0.943 | 3.459 | 2.893 | 4.20 | 2.48 |
Install the Hub client:
pip install -U huggingface_hub
Download one model variant together with its configuration and the repository
manifest. Including the root config.yaml also allows Hugging Face to count
the download at repository level.
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="ZhangXinWhut/BiMTokenizer",
allow_patterns=["config.yaml", "whisper/*"],
local_dir="./BiMTokenizer-Whisper",
)
Replace whisper/* with sensevoice/*, sensevoice-32768-4096/*,
sensevoice-2048/*, or sensevoice-2048-emilia2w/* for the other checkpoints.
The complete repository can also be downloaded with:
huggingface-cli download ZhangXinWhut/BiMTokenizer \
--local-dir ./BiMTokenizer-weights
Clone the code repository and place the downloaded checkpoint under
./weights/. Example:
git clone https://github.com/ZhangXinWhut/BiMTokenizer.git
cd BiMTokenizer
python inference.py \
--config_path /path/to/whisper/config.yaml \
--checkpoint_path /path/to/whisper/bimtokenizer_whisper_librispeech.pt \
--input_dir /path/to/input_wavs \
--output_dir output_wavs \
--device cuda \
--batch_size 1
The fixed Leech codebooks required by the configurations are included in the
GitHub code repository under bimtokenizer/modules/quantizer/cache/.