Downloads · 30 days
0
Jaylin0418/TASTE2-8B-EN-IF-TaskVector-Merged
TASTE2-8B-EN-IF-TaskVector-Merged is a machine learning model from Jaylin0418. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains slm.pt — the speech language model (SLM) backbone of TASTE2-8B-EN, merged with an instruction-following (IF) task vector derived from Qwen2-7B-Instruct.
Downloads · 30 days
0
Access
Public
Updated Jul 27, 2026
Repo size
18.7 GB
Likes
0
Public
Click a slice to open those files.
.pt18.7 GB · 100%
From the Hugging Face model README
This repository contains slm.pt — the speech language model (SLM) backbone of TASTE2-8B-EN, merged with an instruction-following (IF) task vector derived from Qwen2-7B-Instruct.
TASTE2-8B-EN is an English speech language model that generates speech tokens auto-regressively, conditioned on text instructions. Its SLM backbone is a Qwen2-7B-based language model trained on English speech data.
This model merges IF capability into the TASTE-EN backbone without any SFT, using task vector arithmetic.
Let:
slm.ptStep 1: Compute TIES-sparsified IF task vector
$$\tau = \theta_{\text{Instruct}} - \theta_{\text{Qwen2-7B}}$$
$$\hat{\tau} = \text{TopK}_{30%}(\tau) \quad \text{(keep top 30% by magnitude, zero out the rest)}$$
$$\delta_{\text{IF}} = 0.5 \times \hat{\tau}$$
Step 2: Apply task vector to TASTE-EN backbone
$$\theta_{\text{merged}} = \theta_{\text{TASTE-EN}} + \delta_{\text{IF}}$$
| Parameter | Value |
|---|---|
| Merge method | TIES |
| Density | 0.3 |
| Weight | 0.5 |
| SFT after merge | None |
| File | Description |
|---|---|
slm.pt | Merged SLM backbone (18.7 GB). Drop-in replacement for the original slm.pt in TASTE2-8B-EN. |
To run inference, place this slm.pt alongside the other TASTE2-8B-EN model files (flow.pt, hift.pt, llm.pt, speech_tokenizer_v2.onnx, campplus.onnx, etc.) and use the standard TASTE inference pipeline.
If you use this model, please cite the original TASTE paper:
@article{taste,
title={TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling},
author={Tseng, Liang-Hsuan and Chen, Yi-Chang and Lee, Kuan-Yi and Shiu, Da-Shan and Lee, Hung-yi},
journal={arXiv preprint arXiv:2504.07053},
year={2025},
url={https://arxiv.org/abs/2504.07053}
}