Downloads · 30 days
0
USTCPhonetics/FlexAligner
FlexAligner is a machine learning model from USTCPhonetics. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Downloads · 30 days
0
Access
Public
Updated Sep 18, 2026
Repo size
13.9 GB
Likes
0
Public
Click a slice to open those files.
.safetensors5 GB · 100%
From the Hugging Face model README
GitHub · PyPI · English · 简体中文
</div>This repository hosts the model assets used by
USTCPhonetics/FlexAligner.
Use an immutable tag or commit SHA for reproducible work; main is the candidate
asset branch for the next package release.
| Language | Assets | Status |
|---|---|---|
| English | en/chunker, en/aligner | Published and release-tested by FlexAligner |
| Mandarin | zh/chunker, zh/aligner, zh/word.dict | Updated AISHELL-1 candidate for the next FlexAligner release |
The English model files are unchanged from the release-tested bundle. English pronunciations use a user lexicon with the package's optional local G2P fallback; this repository does not ship an English system dictionary.
The Mandarin bundle contains tone-bearing rimes, with 208 Chunker tokens and 207
Aligner tokens. The Aligner supports both sil and sph. The supplied
zh/word.dict is an AISHELL-1 reference lexicon. The current model-compatible
pronunciation for ri uses r iz4, for example 日 r iz4.
0.3.0a1 remains pinned to the older immutable model commit and is
not compatible with the updated Mandarin files on main.sph, updated language detection, and a new
exact manifest pin before using this bundle automatically.v0.2.0a1 are retained unchanged..
├── README.md
├── model_manifest.json
├── en
│ ├── chunker
│ └── aligner
└── zh
├── chunker
├── aligner
└── word.dict
model_manifest.json records the exact file set, sizes, SHA-256 digests,
compatibility boundary, and provenance statements.
USTCPhonetics publishes and redistributes the model weights with authorization to make these artifacts public. It does not claim ownership of the underlying LibriSpeech or AISHELL-1 datasets.
The Mandarin dictionary is the pre-supplied reference lexicon corresponding to AISHELL-1's supplementary resources. It is provided for reference under the AISHELL-1 resource's Apache-2.0 license; USTCPhonetics, OpenPhonetics, and FlexAligner make no ownership claim over that lexicon.
The repository uses license: other because it contains assets with different
provenance and license layers. The dataset licenses do not silently replace the
terms applicable to model weights, and the MIT license of the FlexAligner source
repository is not represented as ownership of the datasets or lexicon.
本仓库保存 FlexAligner 使用的中英文模型资产。为保证可复现性,请固定不可变 tag
或完整 commit SHA;main 是下一软件版本的候选模型分支。
sil 和 sph。zh/word.dict 是预先提供的 AISHELL-1 对应参考词典;本项目不声明对该词典的所有权。ri -> r iz4,例如 日 r iz4。0.3.0a1 仍固定旧模型 revision,不能直接使用 main 上的新
普通话 bundle;下一软件版本完成兼容改造和真实模型 E2E 后才能切换。本仓库采用 license: other,用于准确表达模型权重、训练数据和参考词典的分层来源;
FlexAligner 源代码的 MIT 许可不等于对底层数据或词典主张所有权。