Downloads · 30 days
17
55% of all-time downloads
BUDDY-0912/streamsonic-v5.4-base
streamsonic-v5.4-base is a audio-text-to-text model from BUDDY-0912. Use it for the audio-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
On-device streaming audio perception — silent by default, speak once at a salient event, don't repeat, correctly alert on danger. 端侧流式音频感知 —— 平时沉默、显著事件处开口一次、不重复念、对危险正确告警。
Downloads · 30 days
17
55% of all-time downloads
All-time downloads
31
Public
Parameters
4.7B
9.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.4 GB · 100%
From the Hugging Face model README
On-device streaming audio perception — silent by default, speak once at a salient event, don't repeat, correctly alert on danger. 端侧流式音频感知 —— 平时沉默、显著事件处开口一次、不重复念、对危险正确告警。
This is the base model of StreamSonic v5.4: Qwen2.5-Omni-3B, Thinker only, fine-tuned to the v4loud stage (streaming, 1.5s chunk, "always-describe"). Weights are merged (ready to load). 这是 StreamSonic v5.4 的基座模型:Qwen2.5-Omni-3B,仅 Thinker,微调到 v4loud 阶段(流式、1.5s chunk、随时会描述)。权重已 merged,可直接加载。
v5.4 = this base + two small heads (hosted on GitHub): v5.4 = 本基座 + 两个小头(放在 GitHub):
| Part 部件 | Where 位置 | Size |
|---|---|---|
| Base 底座 (v4loud, frozen) | this repo 本仓库 | ~8.8 GB |
| Main decision head 主决策头 | GitHub: weights/head.pt | 4.1 M |
| Secondary suppressor 二级抑制头 | GitHub: weights/sup.pt | 0.9 K |
Code & slides / 代码与幻灯片: https://github.com/Buddy-svg/streamsonic
python src/train/s2_infer.py --audio <clip>.wav \
--base BUDDY-0912/streamsonic-v5.4-base \
--head weights/head.pt --sup weights/sup.pt \
--sup-thr 0.70 --thr 0.75
Runtime knobs / 运行时旋钮: thr (main head 主头), sup-thr=0.70 (suppressor 抑制强度).
Author: Wenbo Gao (Intern) · Manager: Yuzhou Liu · Mentor: Gordan Han · Week 1 → Week 12 (v5.4).