Downloads · 30 days
954
31% of all-time downloads
cstr/data2vec-audio-960h-GGUF
data2vec-audio-960h-GGUF is a automatic speech recognition model from cstr. Use it when you need speech turned into text. The card lists the license as apache-2.0.
GGUF conversion of facebook/data2vec-audio-base-960h for use with CrispASR.
Downloads · 30 days
954
31% of all-time downloads
All-time downloads
3.1K
Public
Repo size
412 MB
Likes
0
Public
Click a slice to open those files.
.gguf412 MB · 100%
From the Hugging Face model README
GGUF conversion of facebook/data2vec-audio-base-960h for use with CrispASR.
# Uses the wav2vec2 backend (auto-detected from GGUF architecture)
crispasr --backend wav2vec2 -m data2vec-audio-base-960h-q4_k.gguf -f audio.wav
Data2Vec Audio differs from standard wav2vec2 in three ways handled by the converter:
All three are auto-detected from the HuggingFace model config and stored as GGUF metadata flags.
| File | Size | JFK Transcription |
|---|---|---|
| data2vec-audio-base-960h-f16.gguf | 196 MB | perfect |
| data2vec-audio-base-960h-q4_k.gguf | 79 MB | perfect |
| data2vec-audio-base-960h-q8_0.gguf | 120 MB | perfect |
Tested on JFK inaugural address (11s):
AND SO A MY FELLOW AMERICANS ASK NOT WHAT YOUR COUNTRY CAN DO FOR YOU
ASK WHAT YOU CAN DO FOR YOUR COUNTRY
Identical to the Python HuggingFace reference output. All quantized variants produce the same transcription.
@inproceedings{baevski2022data2vec,
title={data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language},
author={Baevski, Alexei and Hsu, Wei-Ning and Xu, Qiantong and Babu, Arun and Gu, Jiatao and Auli, Michael},
booktitle={ICML},
year={2022}
}
facebook.apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.