Downloads · 30 days
9
21% of all-time downloads
TuKoResearch/AuriStream100M_80Pred_BigAudioDataset_500k-randinit
AuriStream100M_80Pred_BigAudioDataset_500k-randinit is a feature extraction model from TuKoResearch. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
AuriStream is a speech language model by Greta Tuckute and Klemen Kotar.
Downloads · 30 days
9
21% of all-time downloads
All-time downloads
42
Public
Parameters
595M
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 100%
From the Hugging Face model README
AuriStream is a speech language model by Greta Tuckute and Klemen Kotar.
This model predicts cochlear tokens from a tokenizer such as WavCochCausalV8192.
Native training step-zero initialization for the 100M 80-prediction model. This exactly uses origin seed 11101994 and historical source commit 6b68c1b639726a9d82b76539d5e769a7fe7e29b9, matching the start of W&B run 2zdi8htl. The weights are untrained FP32 values produced before XLA/FSDP wrapping.
| Parameter | Value |
|---|---|
| Parameters | ~0.59B |
| Layers | 12 |
| Hidden Size | 768 |
| Attention Heads | 12 |
| Vocab Size | 8192 |
| Prediction Steps | 80 |
from transformers import AutoModel, AutoConfig
# Load with trust_remote_code for custom model
model = AutoModel.from_pretrained(
"TuKoResearch/AuriStream100M_80Pred_BigAudioDataset_500k-randinit",
trust_remote_code=True,
)
# Or load config first
config = AutoConfig.from_pretrained("TuKoResearch/AuriStream100M_80Pred_BigAudioDataset_500k-randinit", trust_remote_code=True)
This checkpoint uses shared model code from TuKoResearch/AuriStream-base.
This model uses cochlear tokens from WavCochCausalV8192.