Downloads · 30 days
48
100% of all-time downloads
donghao-zhou/PackLab-VLM-9B
PackLab-VLM-9B is a image-text-to-text model from donghao-zhou. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
PackLab-VLM-9B is a packing-specialized vision-language model for closed-loop robotic bin packing. It is fine-tuned from a Qwen3.5-9B backbone with packing-oriented supervised fine-tuning on PackData-20K.
Downloads · 30 days
48
100% of all-time downloads
All-time downloads
48
Public
Parameters
9.4B
18.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.8 GB · 100%
From the Hugging Face model README
PackLab-VLM-9B is a packing-specialized vision-language model for closed-loop robotic bin packing. It is fine-tuned from a Qwen3.5-9B backbone with packing-oriented supervised fine-tuning on PackData-20K.
The model predicts structured packing actions from multimodal packing states, including the container heightmap, candidate-object attributes, and observation-action history.
Please refer to the PackLab code repository for environment setup, inference, evaluation, and data formatting instructions.
By default, the released code expects the checkpoint at:
weights/PackLab-VLM-9B/
The checkpoint is released in HuggingFace format and contains:
PackLab-VLM-9B/
├── README.md
├── chat_template.jinja
├── config.json
├── generation_config.json
├── model.safetensors
├── processor_config.json
├── tokenizer.json
└── tokenizer_config.json
This checkpoint is intended for research on multimodal robotic bin packing, closed-loop packing policies, and physical packing evaluation. Users should evaluate generated actions carefully before applying the model to physical robotic systems.
The checkpoint is released under the Apache License 2.0.
Users should apply the model responsibly and avoid unsafe robotic deployment without appropriate validation, supervision, and hardware safeguards. Packing actions are determined by the model input, model weights, and runtime decoding settings.