Downloads · 30 days
0
jatshi/Audio-Codec-LLM-Native-Audio-Projector-v3
Audio-Codec-LLM-Native-Audio-Projector-v3 is a feature extraction model from jatshi. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as apache-2.0.
audioprojector.pt is the trainable Whisper-to-Qwen continuous prefix projector from the v3 RTX 4090 smoke run. Whisper-small and Qwen2.5-1.5B-Instruct were loaded as real frozen base models; two optimizer steps reduce…
Downloads · 30 days
0
Access
Public
Updated Aug 2, 2026
Repo size
7.1 MB
Likes
0
Public
Click a slice to open those files.
.pt7.1 MB · 100%
From the Hugging Face model README
audio_projector.pt is the trainable Whisper-to-Qwen continuous prefix projector
from the v3 RTX 4090 smoke run. Whisper-small and Qwen2.5-1.5B-Instruct were loaded
as real frozen base models; two optimizer steps reduced the smoke loss from
3.31035 to 2.70892 (mean 3.00964), with 3,859.58 MiB peak VRAM.
This artifact proves the tensor path, backward pass, optimizer and serialization.
Two steps are not a convergence experiment and do not establish enhancement
quality. See run_manifest.json and the
source repository.