Downloads · 30 days
0
SebastianBodza/DElefant-MPT
DElefant-MPT is a machine learning model from SebastianBodza. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-sa-4.0.
<img src="https://huggingface.co/SebastianBodza/DElefant-MPT/resolve/main/badgegerlefant.png" style="max-width:200px" DElefant is a LLM developed for instruction tuned German interactions. This version is built on top…
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2023
Repo size
110 MB
Likes
2
Public
Click a slice to open those files.
.bin110 MB · 100%
From the Hugging Face model README
QLoRa-Finetuning of the MPT-30B model on two RTX 3090 with the translated WizardLM Dataset.
If there is sufficient demand, additional adjustments can be made:
Prompt-Template:
{instruction}\n\n### Response:
Code example for inference:
import torch
from peft import PeftModel, PeftConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load peft config for pre-trained checkpoint etc.
config = PeftConfig.from_pretrained("SebastianBodza/DElefant-MPT")
# load base LLM model and tokenizer
tokenizer = AutoTokenizer.from_pretrained( "mosaicml/mpt-30b",
padding_side="right",
use_fast=True)
model = AutoModelForCausalLM.from_pretrained("mosaicml/mpt-30b", device_map="auto", load_in_8bit=True)
# Load the Lora model
model = PeftModel.from_pretrained(model, "SebastianBodza/DElefant-MPT", device_map={"":0})
model.eval()
frage = "Wie heißt der Bundeskanzler?"
prompt = f"{frage}\n\n### Response:"
txt = tokenizer(prompt, return_tensors="pt").to("cuda")
txt = model.generate(**txt,
max_new_tokens=256,
eos_token_id=tokenizer.eos_token_id)
tokenizer.decode(txt[0], skip_special_tokens=True)
Gradient-Accumulation led to divergence after a couple of steps. Therefore we reduced the blocksize to 1024 and used two RTX 3090 to get a BS of 4. Probably too small to generalize well.
Training was based on Llama-X with the adaptions of WizardLMs training script and additional adjustments to QLoRa tune. MPT-Code from <a href="https://huggingface.co/SebastianBodza/mpt-30B-qlora-multi_GPU">SebastianBodza/mpt-30B-qlora-multi_GPU</a>
<img src="https://huggingface.co/SebastianBodza/DElefant-MPT/resolve/main/train_loss_DElefant.svg" style="max-width:350px">