Downloads · 30 days
13
5% of all-time downloads
pdsdpo/SynthAlign-7B
SynthAlign-7B is a image-text-to-text model from pdsdpo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
13
5% of all-time downloads
All-time downloads
263
Public
Repo size
28.3 GB
Likes
1
Public
Click a slice to open those files.
.bin14.1 GB · 100%
From the Hugging Face model README
PDS-DPO-7B is a vision-language model built upon LLaVA 1.5 7B and trained using the proposed Preference Data Synthetic Direct Preference Optimization (PDS-DPO) framework. This approach leverages synthetic data generated using generative and reward models as proxies for human preferences to improve alignment, reduce hallucinations, and enhance reasoning capabilities.
@article{wijaya2024multimodal,
title={Multimodal Preference Data Synthetic Alignment with Reward Model},
author={Wijaya, Robert and Nguyen, Ngoc-Bao and Cheung, Ngai-Man},
journal={arXiv preprint arXiv:2412.17417},
year={2024}
}