Downloads · 30 days
0
elitexp/Nepali-Swar-Experimental
Nepali-Swar-Experimental is a text-to-speech model from elitexp. Use it when you need text read aloud. It is set up for pytorch. The card lists the license as apache-2.0.
A 16.1 M parameter semantic language model for Nepali and English speech. Swar (स्वर) is Nepali for voice, sound, or musical note.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
65 MB
Likes
0
Public
Click a slice to open those files.
.pt64.5 MB · 99%
From the Hugging Face model README
A 16.1 M parameter semantic language model for Nepali and English speech. Swar (स्वर) is Nepali for voice, sound, or musical note.
Experimental. A research checkpoint, not a production system.
This is one component of a TTS system, not a complete one. It converts text into CosyVoice3 speech tokens. It does not produce audio by itself.
text -> [ Nepali-Swar, 16.1 M ] -> speech tokens -> [ CosyVoice3 flow + HiFT ] -> audio
this repo frozen, not ours
The decoder is the frozen flow-matching model and HiFT vocoder from
Fun-CosyVoice3-0.5B. This model replaces only the 0.5 B Qwen LM in that stack —
which is the point: 16.1 M parameters in place of 500 M, a 31× reduction in the
component that carries the language.
| dev loss | dev accuracy | |
|---|---|---|
| chance (6,569 classes) | 8.790 | 0.015% |
| Nepali-Swar (16.1 M) | 4.1980 | 18.23% |
| on call-centre register | 2.9938 | 31.53% |
| Fun-CosyVoice3-0.5B reference | 3.765 | 17.6% |
Accuracy is roughly 1,200× chance. That is less impressive than it sounds: next-speech-token prediction has many valid answers for the same text — small differences in pitch or timing are different token ids but equally correct speech — so ~18% sits near the intrinsic ceiling rather than measuring quality. Loss is the number with real headroom.
The shipped checkpoint is tuned for call-centre register, which is why its score on that slice (2.9938 / 31.53%) is far stronger than its general score.
| file | |
|---|---|
swar.pt | the model, 64 MB. Carries lm_config, char_vocab and control_buckets inside it. |
char_vocab.json | 154-character vocabulary — part of the model |
control_buckets.json | rate/pause bucket edges — part of the model |
sample.wav | 5.9 s sample decoded through CosyVoice3's vocoder |
⚠️ char_vocab.json and control_buckets.json are not metadata. Character ids are
positions in that file, and control ids are defined by those bucket edges. Running
these weights against different ones is silent corruption — no error, just wrong
sounds. They are embedded in the .pt for that reason.
import torch
from nvoice.lm import build
from nvoice.vocab import CharVocab, SOS, TASK, ctrl_prefix
ck = torch.load("swar.pt", map_location="cpu")
model = build("M16", **{k: v for k, v in ck["lm_config"].items() if k != "head_dim"})
model.load_state_dict(ck["model"]); model.to("cuda").eval()
vocab = CharVocab(ck["char_vocab"]["chars"]) # travels inside the checkpoint
ids = vocab.encode("नमस्कार, हजुरलाई कसरी सहयोग गर्न सक्छु?")
prefix = [SOS] + ctrl_prefix("callcenter", None, None, None) + ids + [TASK]
tokens = model.generate(prefix, max_new=760, min_new=12, device="cuda")
# tokens -> audio via CosyVoice3's flow + HiFT
Four optional dials precede the text: style (neutral / callcenter), rate,
pause, pitch — five buckets each, or masked. pitch is reserved and
non-functional. Rate control is weak.
Apache-2.0. The decoder it depends on (Fun-CosyVoice3-0.5B) carries its own terms.