Downloads · 30 days
11
17% of all-time downloads
aitytech/WeSpeaker-ResNet34-LM-MLX
WeSpeaker-ResNet34-LM-MLX is a audio classification model from aitytech. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as mit.
MLX-compatible weights for WeSpeaker ResNet34-LM, converted from the pyannote speaker embedding model with BatchNorm fused into Conv2d.
Downloads · 30 days
11
17% of all-time downloads
All-time downloads
66
Public
Parameters
6.6M
26.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors26.5 MB · 100%
From the Hugging Face model README
MLX-compatible weights for WeSpeaker ResNet34-LM, converted from the pyannote speaker embedding model with BatchNorm fused into Conv2d.
WeSpeaker ResNet34-LM is a speaker embedding model (~6.6M params) that produces 256-dimensional L2-normalized speaker embeddings from audio. Trained on VoxCeleb for speaker verification and diarization.
Architecture:
Input: [B, T, 80, 1] log-mel spectrogram (80 fbank, 16kHz)
│
├─ Conv2d(1→32, k=3, p=1) + ReLU
├─ Layer1: 3× BasicBlock(32→32)
├─ Layer2: 4× BasicBlock(32→64, stride=2)
├─ Layer3: 6× BasicBlock(64→128, stride=2)
├─ Layer4: 3× BasicBlock(128→256, stride=2)
│
├─ Statistics Pooling: mean + std → [B, 5120]
├─ Linear(5120→256) → L2 normalize
│
Output: [B, 256] speaker embedding
BatchNorm is fused into Conv2d at conversion time — no BN layers in the MLX model.
import SpeechVAD
// Speaker embedding
let model = try await WeSpeakerModel.fromPretrained()
let embedding = model.embed(audio: samples, sampleRate: 16000)
// embedding: [Float] of length 256, L2-normalized
// Compare speakers
let similarity = WeSpeakerModel.cosineSimilarity(embeddingA, embeddingB)
// Full speaker diarization pipeline
let pipeline = try await DiarizationPipeline.fromPretrained()
let result = pipeline.diarize(audio: samples, sampleRate: 16000)
for seg in result.segments {
print("Speaker \(seg.speakerId): \(seg.startTime)s - \(seg.endTime)s")
}
Part of qwen3-asr-swift.
python3 scripts/convert_wespeaker.py --upload
Converts the original pyannote/wespeaker-voxceleb-resnet34-LM checkpoint using a custom unpickler (no pyannote.audio dependency required). Key transformations:
w_fused = w × γ/√(σ²+ε), b_fused = β − μ×γ/√(σ²+ε)[O, I, H, W] → [O, H, W, I] for MLX channels-lastresnet. prefix, seg_1 → embeddingnum_batches_tracked keys| PyTorch Key | MLX Key | Shape |
|---|---|---|
resnet.conv1.weight + resnet.bn1.* | conv1.weight | [32, 3, 3, 1] |
resnet.layer{L}.{B}.conv{1,2}.weight + bn{1,2}.* | layer{L}.{B}.conv{1,2}.weight | [O, 3, 3, I] |
resnet.layer{L}.0.shortcut.0.weight + shortcut.1.* | layer{L}.0.shortcut.weight | [O, 1, 1, I] |
resnet.seg_1.weight | embedding.weight | [256, 5120] |
resnet.seg_1.bias | embedding.bias | [256] |
The original WeSpeaker model is released under the MIT License.