Downloads · 30 days
0
AllGPTORG/VOSR_1.4B_Mobile
VOSR_1.4B_Mobile is a image-to-image model from AllGPTORG. Use it when you need one image transformed into another. The card lists the license as apache-2.0.
AllGPTORG/VOSR1.4BMobile is a mobile deployment package of the VOSR-1.4B one-step image super-resolution model, quantized and compiled as Qualcomm QNN DLC graphs for the Snapdragon 8 Gen 3 NPU.
Downloads · 30 days
0
Access
Public
Updated Aug 22, 2026
Repo size
15.4 GB
Likes
0
Public
Click a slice to open those files.
.bin7.6 GB · 60%
From the Hugging Face model README
AllGPTORG/VOSR_1.4B_Mobile is a mobile deployment package of the
VOSR-1.4B one-step image super-resolution model, quantized and compiled as
Qualcomm QNN DLC graphs for the Snapdragon 8 Gen 3 NPU.
The package is designed for high-quality image restoration and super-resolution, including text-rich images. It preserves the original one-step VOSR pipeline while splitting the network into smaller graphs suitable for mobile integration.
This repository contains QNN deployment artifacts. It is not a Transformers or Diffusers checkpoint and cannot be loaded with
from_pretrained().
| Property | Value |
|---|---|
| Base model | CSWRY/VOSR, VOSR-1.4B one-step |
| Task | Generative image restoration and super-resolution |
| Target SoC | Qualcomm Snapdragon 8 Gen 3 / SM8650 |
| Compile target | Samsung Galaxy S24 family, Android 14 |
| Runtime format | Qualcomm QNN DLC |
| Static tile size | 512 x 512 pixels |
| Batch size | 1 |
| Activations | FP16 (A16) |
| DiT block weights | INT4 (W4A16) |
| Auxiliary/final graph weights | INT8 (W8A16), except the FP16 VAE decoder |
| Total DLC size | 1,736,469,708 bytes / 1.617 GiB |
The graphs must be executed in the order shown below.
| Order | File | Precision | Size |
|---|---|---|---|
| 1 | vosr_dinov2l_layer17.dlc | W8A16 | 229.50 MiB |
| 2 | vosr_qwen_vae_encoder.dlc | W8A16 | 19.09 MiB |
| 3 | vosr_dit_prepare.dlc | W8A16 | 42.61 MiB |
| 4 | vosr_dit_blocks_00_12.dlc | W4A16 | 444.86 MiB |
| 5 | vosr_dit_blocks_12_18.dlc | W4A16 | 222.62 MiB |
| 6 | vosr_dit_blocks_18_24.dlc | W4A16 | 222.62 MiB |
| 7 | vosr_dit_blocks_24_36.dlc | W4A16 | 444.86 MiB |
| 8 | vosr_dit_final.dlc | W8A16 | 4.90 MiB |
| 9 | vosr_qwen_vae_decoder.dlc | Legacy W8A16 DLC; EPContext uses FP16 | 24.97 MiB |
manifest.json contains the same graph order, precision assignment, and exact
byte size for programmatic use.
sd8g3/qnn-context/ contains eleven pre-linked QNN context binaries for the
Snapdragon 8 Gen 3 HTP. They were linked with QAIRT 2.45.0.260326154327,
target DSP v75 / SoC model 57, and use O1 graph finalization to stay inside the
device's 8 MiB VTCM budget.
The DiT is split into six contiguous stages: 00_06, 06_12, 12_18,
18_24, 24_30, and 30_36. RMSNorm pointwise multiplications are divided
along the token axis before linking; a CPU reference comparison against the
unmodified QDQ block measured 72.9 dB PSNR.
The complete context package is 1,140,031,488 bytes (1.062 GiB). The small
EPContext ONNX wrappers and their exact SHA-256, byte-size, tensor, graph-order,
and immutable Hub-revision contracts are stored in
sd8g3/vosr_runtime_manifest.json.
The Qwen VAE decoder is compiled directly from the original FP16 ONNX graph. The earlier W8A16 QDQ decoder removed nearly all spatial detail. The FP16 SM8650 context matches the CPU ONNX decoder at 59.77 dB PSNR with a maximum per-channel error of one RGB level.
pi/w8a16/ contains the CPU-oriented VOSR-1.4B runtime used by AllCamera Cloud
AI. Its DINO and DiT MatMul weights use ONNX Runtime MatMulNBits INT8
quantization while activations remain FP16. VAE, prepare, and final graphs stay
FP16. The package is 1.663 GiB, down from 3.116 GiB for the split FP16 graphs,
and is pinned by the immutable Hub tag pi-w8a16-v1.
The graphs are deliberately executed one at a time so an 8 GB Raspberry Pi 5
does not hold the complete 1.4B pipeline in memory. The largest measured local
FP16 graph peak was about 1.84 GB RSS; its quantized counterpart used about
557 MB. ONNX Runtime 1.28 provides an ARM64 wheel with the required
MatMulNBits kernels.
Quantization was compared against the complete FP16 split-graph reference on a real zoom patch: 49.34 dB PSNR, 0.99986 RGB correlation, and matching edge energy. A full 40x 2 MP Cloud job passed geometry, color, border, noise, and detail gates after transferring the model reconstruction residual over the denoised OEM frame. This preserves real captured edges while still adding VOSR structure.
All image and latent tensors use NCHW layout.
| Graph | Inputs | Outputs |
|---|---|---|
| DINOv2-L layer 17 | lq_image: FP16 [1,3,512,512] | dino_features: FP16 [1,1024,1024] |
| Qwen VAE encoder | lq_image: FP16 [1,3,512,512]; posterior_noise: FP16 [1,16,64,64] | lq_latent: FP16 [1,16,64,64] |
| DiT prepare | latent_pair: FP16 [1,32,64,64]; timestep: FP32 [1]; next_timestep: FP32 [1]; dino_features: FP16 [1,1024,1024] | hidden: FP16 [1,1024,1536]; conditioning: FP16 [1,1536]; block_conditioning: FP16 [1,9216]; projected_dino: FP16 [1,1024,1536] |
| DiT block stages | hidden, block_conditioning, projected_dino | hidden_out: FP16 [1,1024,1536] |
| DiT final | hidden: FP16 [1,1024,1536]; conditioning: FP16 [1,1536] | velocity: FP16 [1,16,64,64] |
| Qwen VAE decoder | normalized_latent: FP16 [1,16,64,64] | sr_image: FP16 [1,3,512,512] |
Use the Qualcomm AI Engine Direct SDK / QNN runtime to load and execute the DLCs. The host application is responsible for preprocessing, graph orchestration, random noise generation, the one-step latent update, tiling, and image postprocessing.
[-1, 1].posterior_noise if reproducible output is required.z with shape [1,16,64,64], concatenate
lq_latent and z along the channel axis, and use the result as latent_pair.timestep = [1.0] and
next_timestep = [0.0].hidden sequentially through all four DiT block-stage DLCs, or through
all six sd8g3/qnn-context DiT stages when using the EPContext package.
Reuse block_conditioning and projected_dino for every stage.z = z - velocity.[-1, 1], convert it back to RGB, and blend overlapping tiles when tiling.For 4x super-resolution, bicubic-upscale the source to the final target resolution before creating the model tiles. The network then restores detail at that target resolution.
00_06 six-block context was profiled successfully on the Galaxy S24
target: 382.9 ms estimated warm inference, 230.9 ms warm load, and 119.5 MB
estimated peak memory. Every reported operator, including the VTCM chunks,
executed on the NPU.Performance, memory use, and image quality depend on the QNN SDK version, device firmware, thermal state, tiling implementation, and host-side orchestration. Test on the exact target device before shipping a production application.
This package is intended for research and mobile application development involving:
It is not intended for forensic reconstruction or for recovering information that is not present in the source image. Generative restoration can introduce plausible but incorrect details.
This is a quantized mobile derivative of VOSR. The original architecture, training, and checkpoints were created by the VOSR authors. See the official project and paper for full details.
If you use this model, please cite the original VOSR work:
@inproceedings{wu2026vosr,
title = {VOSR: A Vision-Only Generative Model for Image Super-Resolution},
author = {Wu, Rongyuan and Sun, Lingchen and Zhang, Zhengqiang and Kong, Xiangtao and Zhao, Jixin and Wang, Shihao and Zhang, Lei},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}
Released under the Apache License 2.0, following the upstream VOSR repository. Users are responsible for reviewing and complying with the licenses and terms of all upstream components and the Qualcomm QNN SDK/runtime used for deployment.