Downloads · 30 days
0
Qualcomm-AI-Research/Neodragon
Neodragon is a text-to-video model from Qualcomm-AI-Research. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as bsd-3-clause-clear.
<script id="MathJax-script" async src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js"</script <div align="center" style="padding: 20px; border-radius: 10px;" <div style="display: flex; align-items: cente…
Downloads · 30 days
0
Access
Public
Updated Jul 2, 2026
Repo size
17.5 GB
Likes
6
Public
Click a slice to open those files.
.safetensors17.5 GB · 100%
From the Hugging Face model README
Animesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima, Adil Karjauv, Hanwen Xiong, Vancheeswaran Vaidyanathan, Will Zeng, Rafael Esteves, Tushar Singhal, Fatih Porikli, Mohsen Ghafoorian, Amirhossein Habibian
</div>
@inproceedings{
karnewar2026neodragon,
title={Neodragon: Mobile Video Generation Using Diffusion Transformer},
author={Animesh Karnewar and Denis Korzhenkov and Ioannis Lelekas and Noor Fathima and Adil Karjauv and Mohsen Ghafoorian and Amir Habibian},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=XBzIhhwv8d}
}
@article{karnewar2025neodragonTR,
title={Neodragon: Mobile Video Generation using Diffusion Transformer},
author={Karnewar, Animesh and Korzhenkov, Denis and Lelekas, Ioannis and Karjauv, Adil and Fathima, Noor and Xiong, Hanwen and Vaidyanathan, Vancheeswaran and Zeng, Will and Esteves, Rafael and Singhal, Tushar and Porikli, Fatih and Ghafoorian, Mohsen and Habibian Amirhossein},
journal={arXiv preprint arXiv:2511.06055},
url={https://qualcomm-ai-research.github.io/neodragon},
year={2025}
}
<section class="section hero is-light">
<div class="container is-max-widescreen">
<div class="columns is-centered has-text-centered">
<div class="column is-11">
<div class="content has-text-justified">
<p>
We introduce Neodragon, a text-to-video system capable of generating 2s (49 frames @24 fps) videos
at a resolution of <code>[640×1024]</code> directly on a <strong>Qualcomm Hexagon NPU</strong> in a
record <strong>~6.7s</strong> (7 FPS). Differing from existing transformer-based offline text-to-video
generation models, <strong>Neodragon</strong> is the first to have been specifically optimized for mobile
hardware to achieve efficient, low-cost, and high-fidelity video synthesis.
</p>
<ul>
<li>
<strong>Replacing the original large 4.762B <em>T5</em><sub>XXL</sub> Text-Encoder</strong>
with a much smaller 0.2B <em>DT5</em> (DistilT5) with minimal quality loss, enabling the entire model
to run without CPU offloading. This is enabled through a novel Text-Encoder Distillation
procedure which uses only generative text-prompt data and <em>does not</em> require any image or video data.
</li>
<li>
<strong>Proposing an Asymmetric Decoder Distillation approach</strong> which allows us to replace the native
codec-latent-VAE decoder with a more efficient one, without disturbing the generative latent-space of the
video generation pipeline.
</li>
<li>
<strong>Pruning of MMDiT blocks</strong> within the denoiser backbone based on their relative importance,
with recovery of original performance through a two-stage distillation process.
</li>
<li>
<strong>Reducing the NFE (Neural Functional Evaluation) requirement</strong> of the denoiser by performing
step distillation using a technique adapted from DMD for <em>pyramidal</em> flow-matching, thereby significantly
accelerating video generation.
</li>
</ul>
<p>
When paired with an optimized SSD1B first-frame image generator and QuickSRNet for 2×
super-resolution, our end-to-end <strong>Neodragon</strong> system becomes a highly parameter
(<strong>4.945B</strong> full model), memory (<strong>3.5GB</strong> peak RAM usage), and
runtime (<strong>6.7s</strong> E2E latency) efficient mobile-friendly model, while achieving a <em>VBench</em>
total score of <strong>81.61</strong>, yielding high-fidelity generated videos.
</p>
<p>
By enabling low-cost, private, and on-device text-to-video synthesis, <strong>Neodragon</strong> democratizes
AI-based video content creation, empowering creators to generate high-quality videos without reliance on cloud services.
</p>
<p>
Inference code is available at:
<a href="https://github.com/qualcomm-ai-research/neodragon">
https://github.com/qualcomm-ai-research/neodragon
</a>
</p>
</div>
</div>
</div>
</div>
</section>
Please Refer to: https://github.com/qualcomm-ai-research/neodragon
This model is released under the BSD 3-Clause Clear license and the Qualcomm responsible AI license: https://www.qualcomm.com/site/responsible-ai-license
The model is intended for research purposes. Possible research areas and tasks include:
While the capabilities of the presented mobile video generation model are impressive, they can also reinforce or exacerbate social biases strictly based on our foundational-base model Pyramidal-Flow.