Downloads · 30 days
11
9% of all-time downloads
TuKoResearch/AuriStreamDistill_100M40PredTeacher_librispeech960
AuriStreamDistill_100M40PredTeacher_librispeech960 is a feature extraction model from TuKoResearch. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
A Data2Vec-style bidirectional speech encoder trained via distillation from AuriStream models.
Downloads · 30 days
11
9% of all-time downloads
All-time downloads
120
Public
Repo size
717 MB
Likes
0
Public
Click a slice to open those files.
.bin359 MB · 100%
From the Hugging Face model README
A Data2Vec-style bidirectional speech encoder trained via distillation from AuriStream models.
TuKoResearch/AuriStream100M_40Pred_BigAudioDataset_500kfrom transformers import AutoModel, Wav2Vec2FeatureExtractor
import torch
# Load model and feature extractor
model = AutoModel.from_pretrained("TuKoResearch/AuriStreamDistill_100M40PredTeacher_librispeech960", trust_remote_code=True)
model.eval() # Important for inference!
feature_extractor = Wav2Vec2FeatureExtractor.from_pretrained("TuKoResearch/AuriStreamDistill_100M40PredTeacher_librispeech960")
# Prepare audio (16kHz, mono)
audio = torch.randn(16000).numpy() # 1 second of audio
# Extract features
inputs = feature_extractor(audio, return_tensors="pt", sampling_rate=16000)
with torch.no_grad():
outputs = model(inputs.input_values, output_hidden_states=True)
# Get representations
last_hidden = outputs.last_hidden_state # (1, 50, 768) for 1 second
all_hidden = outputs.hidden_states # Tuple of 13 tensors
When output_hidden_states=True, the model returns hidden states from all layers:
hidden_states[0]: Feature projection output (after conv encoder + projection)hidden_states[1] to hidden_states[12]: Transformer layer outputshidden_states[12]: Final layer output (same as last_hidden_state)This makes the model suitable for linear probing experiments at different layers.
This model was trained using Data2Vec-style distillation:
If you use this model, please cite:
@misc{distilled_speech_encoder,
title={Distilled Speech Encoder},
author={TuKo Research},
year={2025},
url={https://huggingface.co/TuKoResearch/AuriStreamDistill_100M40PredTeacher_librispeech960}
}