Downloads · 30 days
0
Jegan022/Vocalis
Vocalis is a machine learning model from Jegan022. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
An End-to-End Visual Silent Speech Interface (VSSI) for Non-Vocal Speakers
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
75.3 MB
Likes
1
Public
Click a slice to open those files.
.pt75.3 MB · 100%
From the Hugging Face model README
An End-to-End Visual Silent Speech Interface (VSSI) for Non-Vocal Speakers
A multi-scale visual articulatory representation combining lip appearance, temporal lip motion, jaw dynamics and lower-face motion, coupled with phoneme-viseme multi-task supervision and speaker-adaptive neural speech synthesis, can produce lower-latency and more intelligible personalized speech than conventional lip-to-text-to-TTS pipelines for non-vocal speakers.
RGB CAMERA (30–60 FPS)
│
▼
Face Tracking & Alignment
(MediaPipe Face Mesh + Affine Warp)
│
▼
Multi-Scale Articulatory ROI
┌──────────────────────┼──────────────────────┐
▼ ▼ ▼
Lip Region Jaw & Chin Cheeks & Contour
(96 × 96) (112 × 112) (64 × 64)
│ │ │
└──────────────────────┼──────────────────────┘
▼
Multi-Stream Spatial Encoder
(ConvNeXt-VSSI / 2D-3D Hybrid Frontends)
│
▼
Cross-Articulatory Attention Fusion
(Lip Appearance + Dynamics + Jaw + Pose)
│
▼
Temporal Articulatory Encoder
(Conformer / Transformer / BiLSTM)
│
▼
Visual Speech Representation
│
┌───────────────┴───────────────┐
▼ ▼
Viseme Decoder (Auxiliary) Phoneme Decoder (CTC)
(Fisher/Jeffers Visemes) (ARPAbet / IPA Tokens)
│ │
└───────────────┬───────────────┘
▼
Multi-Task Articulatory Loss
L = λ₁L_ctc + λ₂L_phoneme + λ₃L_viseme + λ₄L_lang
│
▼
Streaming Prefix Beam Search + LM
+ Confidence & Uncertainty Estimation
│
┌───────────────┴───────────────┐
▼ ▼
[Pipeline A: Cascaded] [Pipeline B: Direct]
Reconstructed Text/Phonemes Articulatory Mel-Gen
│ │
▼ │
Speaker Adapter │
(LoRA / Emb Conditioning) │
│ │
▼ ▼
Personalized Neural TTS Acoustic Vocoder
(Acoustic Model + HiFi-GAN) (HiFi-GAN / Mel)
│ │
└───────────────┬───────────────┘
▼
Real-Time Audio Buffer
(<300 ms End-to-End Latency)
# 1. Clone repository and navigate
git clone https://github.com/vocalis/voxless.git
cd voxless
# 2. Create isolated virtual environment
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/macOS
# 3. Install PyTorch with CUDA 12.x support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124
# 4. Install dependencies
pip install -r requirements.txt
pip install -e .
python -m voxless.cli demo --camera 0 --pipeline pipeline_a
Keyboard controls:
q: Quit applicationr: Reset beam decoder historys: Toggle between Pipeline A (Cascaded) and Pipeline B (Direct)python -m voxless.cli enroll --speaker-id spk_01 --speaker-name "Participant"
python -m voxless.cli collect --speaker-id spk_01 --num-prompts 5
python -m voxless.cli ablate
python -m voxless.cli benchmark
streamlit run voxless/app/web_dashboard.py
Run the automated test suite covering vision, models, decoders, adaptation, synthesis, and streaming latency:
pytest tests/ -v
Voice cloning and personalized synthesis require cryptographic token verification (EthicalConsentGate). Voice profiles cannot be synthesized without participant-signed consent.