Downloads · 30 days
0
jrc/llama3-8b-coedit
llama3-8b-coedit is a machine learning model from jrc. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This is a Llama3 8B based model trained using torchtune on the grammarly/coedit dataset.
Downloads · 30 days
0
Access
Public
Updated Apr 19, 2024
Repo size
16.1 GB
Likes
0
Public
Click a slice to open those files.
.pt16.1 GB · 100%
From the Hugging Face model README
This is a Llama3 8B based model trained using torchtune on the grammarly/coedit dataset.
The exact training script (lora_finetune_distributed) and config (8B_lora.yaml) are both included in this repository.
Training command: tune run --nproc_per_node 8 lora_finetune_distributed --config 8B_lora.yaml
Yes I used 8 GPUs :)
In order to add the dataset, I added the following lines to the config:
dataset:
_component_: torchtune.datasets.instruct_dataset
source: grammarly/coedit
template: GrammarErrorCorrectionTemplate
column_map: {"sentence": "src", "output": "tgt"}
train_on_input: False
split: train
Loss curve
