Downloads · 30 days
18
27% of all-time downloads
AyaanAhmed123/Vspark
Vspark is a text-to-video model from AyaanAhmed123. Use it when you need video from a text prompt. The card lists the license as apache-2.0.
VSpark is a fully trainable, scratch-initialized 726.2M parameter video generation model built with a Spatial-Temporal Diffusion Transformer architecture.
Downloads · 30 days
18
27% of all-time downloads
All-time downloads
67
Public
Parameters
726M
5.8 GB on disk
Likes
2
Trending 1
Click a slice to open those files.
.bin2.9 GB · 50%
From the Hugging Face model README
VSpark is a fully trainable, scratch-initialized 726.2M parameter video generation model built with a Spatial-Temporal Diffusion Transformer architecture.
| Component | Details |
|---|---|
| Video backbone | Spatial-Temporal DiT · 16 blocks · dim=1024 · 16 heads |
| Video VAE | 4-level encoder/decoder · latent_ch=4 · stride-8 |
| Text encoder | 6-layer transformer · dim=768 · BPE vocab=49408 |
| Audio decoder | 8-block mel-spectrogram DiT · dim=512 |
| Total | ~726.2M params |
from vspark_pipeline import VSparkPipeline
pipe = VSparkPipeline.from_pretrained("AyaanAhmed123/Vspark")
frames, audio_mel = pipe("A golden sunset over ocean waves", num_inference_steps=50)
python train_colab.py --data_dir /path/to/videos --epochs 50
Apache 2.0