Downloads · 30 days
16
46% of all-time downloads
shreethar/Latent-Student-Spatial-Forcing
Latent-Student-Spatial-Forcing is a image-text-to-text model from shreethar. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This is the standalone Stage 4 inference package for the Latent Student. The Stage 4 LoRA adapter has been merged into shreethar/LatentStudent-ckpt-400.
Downloads · 30 days
16
46% of all-time downloads
All-time downloads
35
Public
Parameters
4.5B
9.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
From the Hugging Face model README
This is the standalone Stage 4 inference package for the Latent Student. The
Stage 4 LoRA adapter has been merged into shreethar/LatentStudent-ckpt-400.
spatial_parameters.pt: five learned spatial-slot embeddings and the
Stage 4 waypoint MLPlatent_student_config.json: packaging and provenance metadatastage4_config.json: training configuration, when present in the checkpointVGGT and the Spatial Forcing projection head were training-only supervision components. They are not needed for waypoint inference.
shreethar/LatentStudent-ckpt-400stage4_partial_run_2/step_002650best_checkpoint.json002650Use the project's LatentStudent wrapper so the spatial slots and waypoint
head are restored alongside the merged VLM:
from transformers import AutoTokenizer
from train.stage4.checkpointing import load_latent_student_checkpoint
repo_id = "shreethar/Latent-Student-Spatial-Forcing"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
end_think_token_id = tokenizer.convert_tokens_to_ids("</think>")
student = load_latent_student_checkpoint(
checkpoint=repo_id,
end_think_token_id=end_think_token_id,
trainable=False,
M=6,
K=5,
)
student.eval()
Loading only with AutoModelForImageTextToText restores the merged VLM but not
the external spatial slots or waypoint head. Use the wrapper above for the
complete Latent Student behavior.