Downloads · 30 days
0
inference4j/wav2vec2-base-960h
wav2vec2-base-960h is a automatic speech recognition model from inference4j. Use it when you need speech turned into text. It is set up for onnx. The card lists the license as mit.
ONNX export of wav2vec2-base-960h, a Wav2Vec2 model fine-tuned on 960 hours of LibriSpeech for automatic speech recognition using CTC decoding.
Downloads · 30 days
0
Access
Public
Updated Feb 13, 2026
Repo size
378 MB
Likes
1
Public
Click a slice to open those files.
.onnx378 MB · 100%
From the Hugging Face model README
ONNX export of wav2vec2-base-960h, a Wav2Vec2 model fine-tuned on 960 hours of LibriSpeech for automatic speech recognition using CTC decoding.
Mirrored for use with inference4j, an inference-only AI library for Java.
try (Wav2Vec2 model = Wav2Vec2.fromPretrained("models/wav2vec2-base-960h")) {
Transcription result = model.transcribe(Path.of("audio.wav"));
System.out.println(result.text());
}
| Property | Value |
|---|---|
| Architecture | Wav2Vec2 Base (12 transformer layers) |
| Task | Automatic speech recognition (CTC decoding) |
| Training data | LibriSpeech 960h |
| Input | 16kHz mono audio (float32 waveform) |
| Output | CTC logits → greedy-decoded text |
| Original framework | PyTorch (HuggingFace Transformers) |
| ONNX export | By Xenova (Transformers.js) |
This model is licensed under the MIT License. Original model by Facebook AI, ONNX export by Xenova.