Downloads · 30 days
534
71% of all-time downloads
OpenASR/diarizen-large-s80-v2
diarizen-large-s80-v2 is a voice activity detection model from OpenASR. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for openasr. The card lists the license as other.
Downloads · 30 days
534
71% of all-time downloads
All-time downloads
755
Public
Repo size
139 MB
Likes
1
Public
Click a slice to open those files.
.oasr139 MB · 100%
From the Hugging Face model README
DiariZen Large-s80-md-v2 — optional high-accuracy overlap-aware speaker segmentation for local Voice ID
Speaker-diarization support pack for the OpenASR runtime — pure-Rust inference, no Python at inference time.
</div>.oasr pack runs locally through OpenASR's persistent ggml graph without Python at inference time.oasr packs run with no Python at inference, engineered for peak performance on CPU & GPU# 1. Install the OpenASR CLI · https://openasr.org
# 2. Pull the pack
openasr pull diarizen-large-s80-v2:fp16 --accept-license
# 3. Diarize any transcription (works with every OpenASR ASR model)
openasr transcribe meeting.wav --model xasr-zh-en --diarize --format srt
| Quant | File (.oasr) | Size |
|---|---|---|
| fp16 | diarizen-large-s80-v2-fp16.oasr | 139 MB |
<sub>Single fp16 build: projection weights ship as fp16; norms/biases and other parity-sensitive tensors stay f32 inside the pack. No extra public quant tiers.</sub>
DiariZen Large-s80-md-v2 is BUT Speech@FIT's overlap-aware speaker-segmentation
checkpoint built from a 24-layer WavLM Large encoder and a Conformer EEND head.
OpenASR packages the pinned checkpoint as one fp16 .oasr capability pack and
uses it as an optional external segmenter in the universal local-file Voice ID
pipeline. It predicts recording-local speaker activity; ReDimNet2-B6 still
provides clustering and enrolled-person identity. The checkpoint is licensed
under CC BY-NC 4.0, so downloading and activating it require explicit
non-commercial acknowledgement. OpenASR does not select it merely because Voice
ID was enabled; the permissive segmentation-3.0 pack remains the default.
Converted from BUT-FIT/diarizen-wavlm-large-s80-md-v2 with the OpenASR importer:
python3 tooling/diarizen/convert_diarizen.py --checkpoint <pytorch_model.bin> --config <config.toml> --out <diarizen-large-s80-v2-fp16.oasr> --model-id diarizen-large-s80-v2 --quant fp16
The .oasr container is GGUF-backed; projection weights are stored as fp16 while
norms/biases and other parity-sensitive tensors remain f32.
This pack inherits the upstream model's license: CC BY-NC 4.0 (source). OpenASR packaging retains the upstream copyright; the only modification is format conversion.
This pack redistributes the pinned BUT-FIT/diarizen-wavlm-large-s80-md-v2
checkpoint in OpenASR's .oasr runtime format. Credit for the model,
architecture, training and original weights belongs to BUT Speech@FIT and the
DiariZen authors. The checkpoint is licensed under CC BY-NC 4.0; OpenASR's
format conversion does not broaden that license.