Downloads · 30 days
69
41% of all-time downloads
soniqo/ReDimNet2-B6-ONNX-FP32
ReDimNet2-B6-ONNX-FP32 is a audio classification model from soniqo. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as mit.
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.
Downloads · 30 days
69
41% of all-time downloads
All-time downloads
169
Public
Repo size
105 MB
Likes
2
Public
Click a slice to open those files.
.onnx54.4 MB · 52%
From the Hugging Face model README
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.
| Property | Value |
|---|---|
| Parameters | 12.3 million |
| Format | ONNX opset 18, Float32 |
| Model size | 48.9 MiB |
| Input | audio, [1, 96000] mono Float32 samples |
| Sample rate | 16 kHz |
| Window | 6 seconds |
| Output | embedding, [1, 192] L2-normalized Float32 |
Applications should repeat clean two-to-six-second speech to fill the input and center-crop longer samples. Do not use overlapping, mixed, or unalignable speech as identity evidence.
The export is rejected unless its embedding has cosine similarity at least 0.9999 with the pinned PyTorch checkpoint and remains unit-normalized.
| Measurement | Result |
|---|---|
| Warm six-second CPU inference | 245.5 ms |
Latency is measured on the export host and is not a Windows hardware claim.
The supported native host is speech-core:
#include <speech_core/models/onnx_redimnet_speaker_embedding.h>
speech_core::OnnxReDimNetSpeakerEmbedding model(
"ReDimNet2B6.onnx");
auto embedding = model.embed(samples.data(), samples.size(), 16000);
| File | Description |
|---|---|
ReDimNet2B6.onnx | Fixed-shape speaker encoder |
config.json | Graph contract, provenance, hashes, and parity |
README.md | This model card |
LICENSE | Upstream MIT license |
Converted from the official
PalabraAI/ReDimNet2 B6
vb2+vox2_v0 large-margin checkpoint. The pinned source revision and
checkpoint SHA-256 are recorded in config.json.
Speaker embeddings are useful for labeling; they are not biometric authentication and do not protect against voice spoofing.