Downloads · 30 days
3
0% of all-time downloads
ChristianYang/dasheng-base-env-encoder
dasheng-base-env-encoder is a machine learning model from ChristianYang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
mispeech/dasheng-base fine-tuned (LoRA, merged) to extract environment/background embeddings from speech. Trained on ChristianYang/Env-TTS-Clean with conversation-level AAMSoftmax. Auto-uploaded from training step 100…
Downloads · 30 days
3
0% of all-time downloads
All-time downloads
1.1K
Public
Parameters
85.4M
356 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors342 MB · 96%
From the Hugging Face model README
mispeech/dasheng-base fine-tuned (LoRA, merged) to extract environment/background
embeddings from speech. Trained on ChristianYang/Env-TTS-Clean with
conversation-level AAMSoftmax. Auto-uploaded from training step 100000.
import torch
from transformers import AutoModel, AutoFeatureExtractor
from huggingface_hub import hf_hub_download
repo = "ChristianYang/dasheng-base-env-encoder"
backbone = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
fe = AutoFeatureExtractor.from_pretrained(repo, trust_remote_code=True)
head = torch.load(hf_hub_download(repo, "head.pt"), map_location="cpu", weights_only=True)
# pipeline: 16 kHz wav -> fe -> mel -> backbone.encoder tokens [B,T,768]
# -> Conv1d x2 (head["cnn"]) -> masked attentive pooling (head["pool"]) -> 768-d embedding
Head weights (head.pt): two Conv1d(768,768,k=3) layers + attentive pooling
(Linear(768,1)). Token count = mel frames // 4.