Skip to content

Laurie

ChatGLM-Efficient-Tuning-RM

Laurie/ChatGLM-Efficient-Tuning-RM

ChatGLM-Efficient-Tuning-RM is a machine learning model from Laurie. Use it for the machine learning task on the model card, and read the license before you ship it in a product.

1. 首先采用监督学习的方式,使用人工标注的数据在预训练语言模型的基础上进行微调,得到一个初步的语言模型。 2. 然后收集人工反馈的数据,这些数据包含人工生成的文本及人工给出的评分。利用这些带评分的数据训练一个reward模型。 3. 最后,利用reward模型产生的奖励信号,采用强化学习的方式对第1步得到的语言模型进行优化。在这个步骤中,会生成大量文本,并采用reward模型给出的奖励进行更新。 所以整体流程分为监督微调、训练r…

Downloads · 30 days

0

Access

Public

Updated Jun 12, 2023

Repo size

44.2 MB

Likes

3

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt29.4 MB · 67%

At a glance

Access
Public
Created
May 2, 2023
Updated
Jun 12, 2023
SHA
c5f13841
Created
May 2, 2023
Updated
Jun 12, 2023
ChatGLM-Efficient-Tuning-RM — AI Model — AIMarketly