Downloads · 30 days
284
5% of all-time downloads
aufklarer/ReDimNet2-B6-CoreML
ReDimNet2-B6-CoreML is a audio classification model from aufklarer. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for coreml. The card lists the license as mit.
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.
Downloads · 30 days
284
5% of all-time downloads
All-time downloads
6.1K
Public
Repo size
25.4 MB
Likes
2
Public
Click a slice to open those files.
.bin25.4 MB · 98%
From the Hugging Face model README
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.
| Property | Value |
|---|---|
| Parameters | 12.3 million |
| Format | Compiled Core ML, Float16 weights |
| Compiled size | 24.7 MiB |
| Input | 96,000 mono Float32 samples |
| Sample rate | 16 kHz |
| Window | 6 seconds |
| Output | 192-dimensional L2-normalized embedding |
| Minimum deployment | macOS 15 / iOS 18 |
The checkpoint was trained on VoxBlink2 and VoxCeleb2. The fixed six-second shape avoids the slow Core ML fallback observed with a flexible waveform shape. Applications should repeat clean two-to-six-second speech to fill the input and center-crop longer samples.
| File | Size | Description |
|---|---|---|
ReDimNet2B6.mlmodelc/ | 24.7 MiB | Precompiled Core ML model |
config.json | <2 KiB | Input, output, source revision, checksum, and validation metadata |
README.md | <4 KiB | This model card |
LICENSE | 1.0 KiB | MIT license from the upstream implementation |
Measured on an Apple M2 Max after two warm-up predictions:
| Measurement | Result | Meaning |
|---|---|---|
| Warm six-second inference | 13.8 ms | One voice-profile embedding |
| Warm throughput | 72.6 embeddings/s | Repeated six-second windows after warm-up |
| Meeting pilot equal-error rate, 2-second clips | 1.50% | Lower is better; WeSpeaker Core ML was 5.17% |
| Meeting pilot equal-error rate, 3-second clips | 0.00% | Lower is better; WeSpeaker Core ML was 1.50% |
| LibriSpeech test-clean equal-error rate, 40 speakers | 0.00% | Two- and three-second controls |
The meeting pilot contains five recurring speakers and is not a universal quality claim. Thresholds must be calibrated for the intended microphones, languages, and acoustic conditions. Speaker embeddings are useful for labeling; they are not biometric authentication and do not protect against voice spoofing.
import coremltools as ct
import numpy as np
model = ct.models.CompiledMLModel("ReDimNet2B6.mlmodelc")
audio = np.zeros((1, 96_000), dtype=np.float32)
embedding = model.predict({"audio": audio})["embedding"]
speech embed-speaker voice.wav --engine redimnet2 --json
import SpeechVAD
let model = try await ReDimNet2SpeakerModel.fromPretrained()
let embedding = try model.embed(audio: samples, sampleRate: 16_000)
Converted from the official
PalabraAI/ReDimNet2 B6
vb2+vox2_v0 large-margin checkpoint. The source revision and checkpoint
SHA-256 are recorded in config.json.