Downloads · 30 days
579
75% of all-time downloads
Anthropic-ai/aivoice
aivoice is a text-to-speech model from Anthropic-ai. Use it when you need text read aloud. The card lists the license as apache-2.0.
Downloads · 30 days
579
75% of all-time downloads
All-time downloads
774
Public
Repo size
11.6 GB
Likes
1
Public
Click a slice to open those files.
.safetensors6.8 GB · 58%
From the Hugging Face model README
notebook https://cuty.io/pQGWF3c7
This directory contains the local model files used by viitor-ai/viitor-voice-nar.
ViiTorVoice-NAR is a non-autoregressive speech generation model for voice cloning, local speech editing, and emotion / paralinguistic speech control. The files in this directory are split by function so each model component can be loaded independently.
local_models/
├── aligner/
│ └── Qwen3-ForcedAligner-0.6B/
├── assets/
│ └── dualcodec_silence_2s.pt
├── dualcodec/
│ ├── dualcodec_ckpts/
│ └── w2v-bert-2.0/
└── llm/
└── 0p6_emotion/
| Component | Path | Purpose |
|---|---|---|
| ViiTorVoice-NAR LLM | llm/0p6_emotion/ | Generates target speech tokens from text, prompt speech tokens, edit masks, duration conditions, and emotion or non-verbal tags. |
| DualCodec | dualcodec/dualcodec_ckpts/ | Converts waveform audio into discrete speech codebook tokens and decodes generated tokens back into waveform audio. |
| W2V-BERT 2.0 | dualcodec/w2v-bert-2.0/ | Extracts semantic speech features used by the DualCodec encoder. |
| Qwen3 Forced Aligner | aligner/Qwen3-ForcedAligner-0.6B/ | Aligns speech audio with text and provides timestamps for local speech editing. |
| Runtime Assets | assets/ | Stores small auxiliary files, such as precomputed silence tokens used during generation or padding. |