Downloads · 30 days
25
12% of all-time downloads
tianyil1/MistralForCausalLM_Cal_DPO
MistralForCausalLM_Cal_DPO is a machine learning model from tianyil1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
25
12% of all-time downloads
All-time downloads
201
Public
Parameters
7.2B
29 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of alignment-handbook/zephyr-7b-sft-full on the HuggingFaceH4/ultrafeedback_binarized dataset.
The Cal-DPO algorithm effectively addresses the alignment problem between large language models and human preferences by calibrating the implicit rewards in comparative preference learning to match the real rewards. It has demonstrated excellent performance in multiple task benchmark tests.
The following hyperparameters were used during training:
We evaluate models on 6 key benchmarks using the Eleuther AI Language Model Evaluation Harness , a unified framework to test generative language models on a large number of different evaluation tasks.