Downloads · 30 days
0
Reza2kn/Audio-Mosaicist-1
Audio-Mosaicist-1 is a machine learning model from Reza2kn. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Audio Mosaicist-1 is a multilingual text-to-AudioMosaic bridge for searching large unstructured audio corpora with natural-language prompts. It maps Qwen3 text embeddings into the frozen AudioMosaic acoustic z space,…
Downloads · 30 days
0
Access
Public
Updated Sep 9, 2026
Repo size
1.7 MB
Likes
1
Public
Click a slice to open those files.
.jsonl8.2 MB · 81%
From the Hugging Face model README
Audio Mosaicist-1 is a multilingual text-to-AudioMosaic bridge for searching large unstructured audio corpora with natural-language prompts. It maps Qwen3 text embeddings into the frozen AudioMosaic acoustic z space, then retrieves and localizes candidate audio events.
This repository contains the bridge artifacts and helper scripts. It does not contain the full private Persian audio corpus used to build the current VisualEars mining index.
mlx-community/Qwen3-Embedding-0.6B-4bit-DWQ for the current build.projection_category_qwen3_to_audiomosaic_z.npy: category-level Qwen3-text to AudioMosaic-z bridge.projection_exact_anchor_qwen3_to_audiomosaic_z.npy: exact-anchor bridge.text_category_prototypes.npy and text_category_prototypes.jsonl: text-side category prototypes.audio_category_centroids.npy: AudioMosaic-side category centroids.audio_anchor_index.with_inferred_categories.parquet: anchor metadata used for the current bridge.categories.json: current category list.audiomosaic_text_bridge_query.py: minimal query helper.scripts/asr_query_audiomosaic_text_bridge.py: full-corpus search helper used in the VisualEars mining run.scripts/asr_localize_audiomosaic_events.py: multi-scale timestamp localizer for retrieved candidates.scripts/asr_make_audio_mosaicist_extension_pack.py: helper for building local-language extension packs.Download the repo and point the query script at an existing AudioMosaic embedding index:
python scripts/asr_query_audiomosaic_text_bridge.py \
--bridge-dir Audio-Mosaicist-1 \
--index /path/to/audiomosaic_index \
--prompts prompts.jsonl \
--out search_results.jsonl
Then recover exact audio rows and localize the event windows:
python scripts/asr_localize_audiomosaic_events.py \
--candidates search_results.jsonl \
--manifest exact_manifest_rows.jsonl \
--out localized_events.jsonl \
--repo-dir /path/to/audiomosaic-vit-b16-pretrained \
--device cuda:0
The helper scripts are intentionally plain Python so people can adapt them to their own storage layout.
The bridge was calibrated with English and Persian prompts, so those two are validated. Because Qwen3 Embedding is multilingual, prompts in other Qwen-supported languages may work through cross-lingual alignment, but they are not yet measured. For non-English/Persian use, treat this as zero-shot until you add local-language calibration prompts or anchors.
You do not need a carefully structured dataset. The easiest extension is a JSONL/CSV with any of these fields when available:
audio_path or HF dataset locator fields (source, file_path, row_index, audio_col)label or caption in your languagelang, for example ar, hi, tr, decategory if you already know itRecommended extension workflow:
z vectors.For unlabeled audio, first mine clusters in AudioMosaic space, manually name a small number of representative clusters in your language, then use those labels as bridge anchors.
You can start an extension pack with:
python scripts/asr_make_audio_mosaicist_extension_pack.py \
--input your_audio_or_labels.csv \
--out audio_mosaicist_extension_your_lang.jsonl \
--lang your_language_code
The extension pack is a staging file: embed the listed audio with AudioMosaic, embed the labels with Qwen3 or another multilingual text encoder, then fit a small adapter or refit the bridge.
AudioMosaic retrieval is coarse. The production path is:
z query.Fine-window scores are relative, not fully calibrated probabilities. For VisualEars-grade event alerts, the next model should train a dedicated fine event localizer on mined/mixed spans.
This model distribution is licensed under the Apache License, Version 2.0. See LICENSE. Existing third-party copyright, license, and attribution notices remain applicable.