Downloads · 30 days
14
30% of all-time downloads
merve/peft-copy-test
peft-copy-test is a text generation model from merve. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as bigscience-openrail-m.
Downloads · 30 days
14
30% of all-time downloads
All-time downloads
46
Public
Repo size
34.1 MB
Likes
0
Public
Click a slice to open those files.
.bin33.6 MB · 99%
From the Hugging Face model README

Adapter weights of a Reinforcement Learning fine-tuned model based on the LLaMA model (see Meta's LLaMA release for the original LLaMA model). The model is designed to generate human-like responses to questions in Stack Exchange domains of programming, mathematics, physics, and more. For more info check out the blog post and github example.
Developed by: Hugging Face
Model type: An auto-regressive language model based on the transformer architecture, and fine-tuned with Stack Exchange datasets.
Languages: Predominantly English, with additional data from languages with the following ISO codes:
| bg | ca | cs | da | de | es | fr | hr | hu | it | nl | pl | pt | ro | ru | sl | sr | sv | uk |
|---|
License: bigscience-openrail-m
Finetuned from: LLaMA
Repository: https://huggingface.co/trl-lib/llama-7b-se-rl-peft/tree/main
Base Model Repository: https://github.com/facebookresearch/llama
Demo: https://huggingface.co/spaces/trl-lib/stack-llama
Original datasets are described in the LLaMA Model Card. Fine-tuning datasets for this model are based on Stack Exchange Paired, which consists of questions and answers from various domains in Stack Exchange, such as programming, mathematics, physics, and more. Specifically:
Traditional Fine-tuning: https://huggingface.co/datasets/lvwerra/stack-exchange-paired/tree/main/data/finetune
RL Fine-tuning: https://huggingface.co/datasets/lvwerra/stack-exchange-paired/tree/main/data/rl
Reward Model: https://huggingface.co/trl-lib/llama-7b-se-rm-peft
The model was first fine-tuned on the Stack Exchange question and answer pairs and then RL fine-tuned using a Stack Exchange Reward Model. It is trained to respond to prompts with the following template:
Question: <Query>
Answer: <Response>
BibTeX:
@misc {beeching2023stackllama,
author = { Edward Beeching and
Younes Belkada and
Kashif Rasul and
Lewis Tunstall and
Leandro von Werra and
Nazneen Rajani and
Nathan Lambert
},
title = { StackLLaMa: An RL Fine-tuned LLaMa Model for Stack Exchange Question and Answering },
year = 2023,
url = { https://huggingface.co/trl-lib/llama-7b-se-rl-peft },
doi = { 10.57967/hf/0513 },
publisher = { Hugging Face Blog }
}
Nathan Lambert, Leandro von Werra, Edward Beeching, Kashif Rasul, Younes Belkada, Margaret Mitchell