Downloads · 30 days
0
Kwai-Kolors/MetaView
MetaView is a machine learning model from Kwai-Kolors. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2026
Repo size
8 GB
Likes
13
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
MetaView is a diffusion-based framework for high-fidelity novel view synthesis that enables accurate rendering under large view changes from a single image.
Current generative novel view synthesis methods typically rely on restrictive explicit 3D reconstruction pipelines to enforce spatial consistency but inherently restrict generalization in large view changes. Morever, recent interactive generative methods suffers from scale drifting and poor geometry consistency. MetaView bridges this gap by combining implicit geometry modeling with minimal yet essential explicit 3D cues:
Scale-Aware Implicit Geometry Priors: We incorporate implicit geometry priors from a feed-forward geometry perception network (Depth Anything 3) to regularize structure without imposing restrictive reconstruction pipelines. These geometric signals are incorporated into the pretrained MM-DiT backbone via non-invasive parallel attention layers.
Metric Scale Anchoring via Modified RoPE: To overcome the scale drifting issue prevalent in fully implicit methods, we allocate an extra subspace for the z-axis to anchor the scene scales and it plays with camera pose via PRoPE. This explicitly injects scale cues, anchoring the generation to a consistent 3D metric space.
Given a single input image and a target camera pose, MetaView synthesizes the corresponding novel view with precise camera controllability, strong geometry consistency, and remarkable cross-domain generalization.