Downloads ยท 30 days
0
niobures/ctc_forced_aligner
ctc_forced_aligner is a machine learning model from niobures. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
We are open-sourcing the CTC forced aligner used in Deskpai.
Downloads ยท 30 days
0
Access
Public
Updated Feb 20, 2026
Repo size
1.3 GB
Likes
0
Public
Click a slice to open those files.
.onnx1.3 GB ยท 100%
From the Hugging Face model README
We are open-sourcing the CTC forced aligner used in Deskpai.
With focus on production-ready model inference, it supports 18 different alignment models, including multilingual models(German, English, Spanish, French and Italian etc), and provides SRT and WebVTT alignment and generation out of box. It supports both ONNXRuntime and PyTorch for model serving.
pip install ctc_forced_aligner
pip install ctc_forced_aligner[gpu]
pip install ctc_forced_aligner[torch]
pip install ctc_forced_aligner[all]
from ctc_forced_aligner import AlignmentSingleton
alignment_service = AlignmentSingleton()
input_audio_path = "audio.mp3"
input_text_path = "input.txt"
output_srt_path = "output.srt"
ret = alignment_service.generate_srt(input_audio_path,
input_text_path,
output_srt_path)
if ret:
print(f"Aligned SRT is generated at {output_srt_path}")
output_vtt_path = "output.vtt"
ret = alignment_service.generate_webvtt(input_audio_path,
input_text_path,
output_vtt_path)
if ret:
print(f"aligned VTT is generated to {output_vtt_path}")
from ctc_forced_aligner import AlignmentTorch
at = AlignmentTorch()
ret = at.generate_srt(input_audio_path, input_text_path, output_srt_path)
if ret:
print(f"aligned srt is generated to {output_srt_path}")
ret = at.generate_webvtt(input_audio_path, input_text_path, output_vtt_path)
if ret:
print(f"aligned VTT is generated to {output_vtt_path}")
from ctc_forced_aligner import AlignmentTorch
at = AlignmentTorch()
ret = at.generate_srt(input_audio_path, input_text_path, output_srt_path, model_type='WAV2VEC2_ASR_BASE_960H')
if ret:
print(f"aligned srt is generated to {output_srt_path}")
ret = at.generate_webvtt(input_audio_path, input_text_path, output_vtt_path, model_type='WAV2VEC2_ASR_BASE_960H')
if ret:
print(f"aligned VTT is generated to {output_vtt_path}")
These are fine-tuned models with a CTC-based ASR head:
WAV2VEC2_ASR_BASE_960HWAV2VEC2_ASR_BASE_100HWAV2VEC2_ASR_BASE_10MWAV2VEC2_ASR_LARGE_10MWAV2VEC2_ASR_LARGE_100HWAV2VEC2_ASR_LARGE_960HWAV2VEC2_ASR_LARGE_LV60K_10MWAV2VEC2_ASR_LARGE_LV60K_100HWAV2VEC2_ASR_LARGE_LV60K_960HThese models are fine-tuned for specific languages:
VOXPOPULI_ASR_BASE_10K_DE (German ASR)
VOXPOPULI_ASR_BASE_10K_EN (English ASR)
VOXPOPULI_ASR_BASE_10K_ES (Spanish ASR)
VOXPOPULI_ASR_BASE_10K_FR (French ASR)
VOXPOPULI_ASR_BASE_10K_IT (Italian ASR)
Fine-tuned on VoxPopuli speech corpus.
HUBERT_ASR_LARGEHUBERT_ASR_XLARGEFor PyTorch serving, use AlignmentTorch or AlignmentTorchSingleton.
WAV2VEC2_ASR_LARGE_960H or HUBERT_ASR_LARGEVOXPOPULI_ASR_BASE_10K_*WAV2VEC2_ASR_BASE_10M (smallest model)WAV2VEC2_ASR_LARGE_LV60K_960H or HUBERT_ASR_XLARGEFor ONNXRuntime serving with minimum dependencies, use Alignment or AlignmentSingleton.
Please contact us if you want to integrate your model into this package.
BSD-2-Clause license.BSD license.This project is licensed under the BSD License, note that the default model has CC-BY-NC 4.0 License, so make sure to use a different model for commercial usage.MIT License and redistributed with the same license.
WAV2VEC2_ASR_BASE_960HWAV2VEC2_ASR_BASE_100HWAV2VEC2_ASR_BASE_10MWAV2VEC2_ASR_LARGE_10MWAV2VEC2_ASR_LARGE_100HWAV2VEC2_ASR_LARGE_960HWAV2VEC2_ASR_LARGE_LV60K_10MWAV2VEC2_ASR_LARGE_LV60K_100HWAV2VEC2_ASR_LARGE_LV60K_960HVOXPOPULI_ASR_BASE_10K_DEVOXPOPULI_ASR_BASE_10K_ENVOXPOPULI_ASR_BASE_10K_ESVOXPOPULI_ASR_BASE_10K_FRVOXPOPULI_ASR_BASE_10K_ITHUBERT_ASR_LARGEHUBERT_ASR_XLARGEMMS_FA is published by the authors of Scaling Speech Technology to 1,000+ Languages Pratap et al., 2023 under CC-BY-NC 4.0 License.CC-BY-NC 4.0 License.๐ Note: It's essential to verify the licensing terms from the official repositories or documentation before using these models.