Downloads · 30 days
0
leduy99/Mobile-OV
Mobile-OV is a text-to-video model from leduy99. Use it when you need video from a text prompt. The card lists the license as other.
One downloadable inference-only bundle for neomobileov-infer-clear. This release selects V11-balanced (not V11-control), Exp1-64K and the released NeoDragon Hybrid. It is not an experimental Joint V2/V3 or DMD checkpo…
Downloads · 30 days
0
Access
Public
Updated Sep 10, 2026
Repo size
10.3 GB
Likes
0
Public
Click a slice to open those files.
.pt6.2 GB · 60%
From the Hugging Face model README
One downloadable inference-only bundle for
neo_mobileov-infer-clear.
This release selects V11-balanced (not V11-control), Exp1-64K and the released
NeoDragon Hybrid. It is not an experimental Joint V2/V3 or DMD checkpoint.
The older mobile_ov_135k_full.pt in this repository is a separate legacy file
and is not needed for this release.
No optimizer, gradient scaler, RNG state, training history, teacher model, duplicate Smol backbone, native Qwen/T5/CLIP neural encoder, SSD1B or QuickSR weights are included. The small Qwen tokenizer remains for the trained image condition-length rule. All seven safetensors weight files retain their original bytes/dtypes; no quantization or lossy precision conversion was applied.
The .tar.gz is a single distribution file, not a pickle checkpoint to pass
to torch.load or a standard Transformers AutoModel directory. Its component
layout preserves lazy loading, shared weights and separate conversion targets.
It contains everything the clean inference runtime needs except Python packages
and the runtime code itself. No sibling research repo is needed.
Clone the clean code repo and install its tested dependencies first. Run the following from that repo root, using an empty destination for the new assets:
hf download leduy99/Mobile-OV \
v11-exp1/mobile-ov-v11-exp1-inference.tar.gz \
v11-exp1/mobile-ov-v11-exp1-inference.tar.gz.sha256 \
--local-dir downloads
(cd downloads/v11-exp1 && sha256sum -c mobile-ov-v11-exp1-inference.tar.gz.sha256)
mkdir -p assets
tar -xzf downloads/v11-exp1/mobile-ov-v11-exp1-inference.tar.gz -C assets
python generate.py --assets assets/v11-exp1 --device cpu --verify-assets \
--prompt 'A red toy car drives along a sunny coastal road.' \
--output output/example.mp4
The command generates the first frame with DreamLite and the continuation with NeoDragon. Default video: 49 frames at 24 fps, 320x512; image anchor: 640x1024, four steps. Logical DreamLite time IDs remain 800x1280, as in the validated recipe. The two VAEs are not interchangeable: the RGB image is encoded by the video VAE.
For GPU inference on a SLURM-managed machine, obtain an allocation before using
--device cuda, or use the clean repo's scripts/generate_local.sbatch.
python understand.py --assets assets/v11-exp1 --device cpu \
--image output/example.anchor.png --prompt 'Describe this image.'
Text understanding and sampled-frame video understanding are also supported. Instruction-based image/video editing is not validated.
v11-exp1/bundle_report.json records archive size, checksum, per-component tensor
counts and dtypes. manifest.json inside the archive verifies every bundled
asset. Preparation checks all input hashes and detects changes while packing.
Source checkpoint hashes are preserved, but machine-local paths are omitted.
The clean runtime's recorded CPU parity checks cover conditions, three images, three videos and both bridge export graphs. Packaging does not improve the known prompt-following or motion weaknesses. This is a Linux/PyTorch reference, not a finished iPhone/iPad/CoreML model; device latency, memory and quantization quality are not established by this release. See the code repo's validation log.
Read NOTICES.md and the notices inside the archive before use. In particular, DreamLite is CC BY-NC 4.0 (non-commercial). NeoDragon carries BSD-3-Clause-Clear plus the Qualcomm Responsible AI License; SmolVLM2 is Apache-2.0. Packaging does not replace these terms or grant patent rights. Use this bundle for non-commercial research consistent with all component terms.
Original generators and backbone are the work of their respective authors. The Mobile-OV research contributes the trained bridges and their integration; no claim of training these foundational generators from scratch is made.