Downloads · 30 days
0
DataoceanAI1/dolphin-cn-dialect-small-streaming
dolphin-cn-dialect-small-streaming is a machine learning model from DataoceanAI1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated May 13, 2026
Repo size
1.8 GB
Likes
2
Public
Click a slice to open those files.
.pt1.8 GB · 100%
From the Hugging Face model README
Paper Github Huggingface Modelscope
This model is officially maintained by Dataocean AI.
To ensure compatibility with existing user code and download links, we keep two official repositories for the same model:
Both repositories are maintained by the same team and contain the same model files.
DataoceanAI1 is the newly created enterprise organization account, while DataoceanAI is kept to avoid breaking existing user download scripts and links.
Please do not regard either repository as an unofficial copy or unauthorized redistribution.
Dolphin-CN-Dialect is a multi-dialect ASR model developed by Dataocean AI and Tsinghua University, with a strong focus on Chinese dialect recognition and real-world deployment scenarios. Compared with the previous Dolphin series, Dolphin-CN-Dialect introduces significant improvements in tokenizer design, dialect-balanced training, streaming capability, hotword biasing, and deployment efficiency.
The model supports Mandarin Chinese and 22 Chinese dialects, while also maintaining multilingual ASR capability inherited from Dolphin. Dolphin-CN-Dialect supports both streaming and non-streaming inference, enabling practical deployment in latency-sensitive applications such as real-time transcription and industrial speech recognition systems.
Dolphin-CN-Dialect is built upon the Dolphin architecture and follows a joint CTC-Attention framework with:
Compared to Dolphin, Dolphin-CN-Dialect introduces several important improvements:
Experimental results show that Dolphin-CN-Dialect achieves:

See details in the Paper.
Dolphin-CN-Dialect requires FFmpeg to convert audio files into WAV format. Please install FFmpeg first if it is not already installed on your system.
# Ubuntu / Debian
sudo apt update && sudo apt install ffmpeg
# MacOS
brew install ffmpeg
# Windows
choco install ffmpeg
Install Dolphin with pip:
pip install -U dolphin
Alternatively, install from source:
pip install git+https://github.com/DataoceanAI/Dolphin.git
Currently, Dolphin-CN-Dialect provides multiple model sizes optimized for different deployment scenarios.
| Model | Parameters | Hotwords |
|---|---|---|
| base.cn | 0.1 B | ❌ |
| base.cn.streaming | 0.1 B | ❌ |
| small.cn | 0.4 B | Encoder-biased Hotwords |
| small.cn.streaming | 0.4 B | Encoder-biased Hotwords |
| small.cn.prompt | 0.4 B | Prompt-based Hotwords |
Dolphin-CN-Dialect supports two hotword biasing approaches.
Encoder-Level Contextual Biasing
Prompt-Based Hotword Biasing
Experimental results show significant reductions in hotword error rates while maintaining strong overall ASR performance.
Dolphin-CN-Dialect primarily focuses on:
Supported dialects include:
For the complete language and dialect list, see languages.md.
| Device Type | Support Status |
|---|---|
| CUDA | ✅Supported |
| MPS (Apple) | ✅Supported |
| CPU | ✅Supported |
dolphin audio.wav
# Download model and specify the model path
dolphin audio.wav --model small.cn --model_dir /data/models/dolphin/
# Specify language and region
dolphin audio.wav --model small.cn --model_dir /data/models/dolphin/ --lang_sym "zh" --region_sym "CN"
# Specify the hotwords file with Encoder-biased method
dolphin audio.wav --model small.cn --model_dir /data/models/dolphin/ --hotword_list_path hotwords.txt --use_deep_biasing true
# Using prompt-based model
dolphin audio.wav --model small.cn.prompt --model_dir /data/models/dolphin/ --hotword_list_path hotwords.txt --use_prompt_hotword true --use_two_stage_filter true
import dolphin
from dolphin import transcribe
model_name = 'small.cn'
model = dolphin.load_model(model_name, device="cuda")
result = transcribe(model, 'audio.wav')
print(result.text)
# Specify language
result = transcribe(model, 'audio.wav', lang_sym="zh")
print(result.text)
# Specify language and region and encoder-biased hotwords
result = transcribe(model, 'audio.wav', lang_sym="zh", region_sym="CN", hotwords=['诺香丹青牌科研胶囊'], use_deep_biasing=True, use_two_stage_filter=True)
print(result.text)
## prompt-based hotwords
model_name = 'small.cn.prompt'
model = dolphin.load_model(model_name, device="cuda")
result = transcribe(model, 'audio.wav', hotwords=['诺香丹青牌科研胶囊'], use_prompt_hotword=True, use_two_stage_filter=True, decoding_method='attention')
print(result.text)
Dolphin-CN-Dialect is released under the Apache 2.0 License.