Downloads · 30 days
276
8% of all-time downloads
chancharikm/CHAI_SFT_model_8b
CHAI_SFT_model_8b is a video-text-to-text model from chancharikm. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of Qwen/Qwen3-VL-8B-Instruct on the allsftformatsunbalanced20251122part1 dataset.
Downloads · 30 days
276
8% of all-time downloads
All-time downloads
3.5K
Public
Parameters
770K
452 GB on disk
Likes
2
Public
Click a slice to open those files.
.pt189 GB · 84%
From the Hugging Face model README
This model is a fine-tuned version of Qwen/Qwen3-VL-8B-Instruct on the all_sft_formats_unbalanced_20251122_part_1 dataset.
It was developed as part of the research paper: "Building a Precise Video Language with Human-AI Oversight" (CVPR 2026 Highlight).
This model belongs to a family of video-language models (VLMs) optimized for precise video captioning using the CHAI (Critique-based Human–AI Oversight) framework. CHAI pairs trained human experts with model-generated pre-captions: experts provide correctional critiques that guide revisions into improved post-captions.
This model is intended for research in:
More information needed
The following hyperparameters were used during training:
If you find this work useful, please cite:
@inproceedings{chai2026,
title = {Building a Precise Video Language with Human--AI Oversight},
author = {Zhiqiu Lin and Chancharik Mitra and Siyuan Cen and Isaac Li and Yuhan Huang and Yu Tong Tiffany Ling and Hewei Wang and Irene Pi and Shihang Zhu and Ryan Rao and George Liu and Jiaxi Li and Ruojin Li and Yili Han and Yilun Du and Deva Ramanan},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}