Downloads · 30 days
45
0% of all-time downloads
beomi/gemma-mling-7b
gemma-mling-7b is a text generation model from beomi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Update @ 2024.04.15: First release of Gemma-Mling 7B model
Downloads · 30 days
45
0% of all-time downloads
All-time downloads
21.2K
Public
Parameters
8.5B
17.1 GB on disk
Likes
14
Public
Click a slice to open those files.
.safetensors17.1 GB · 100%
From the Hugging Face model README
Update @ 2024.04.15: First release of Gemma-Mling 7B model
Original Gemma Model Page: Gemma
This model card corresponds to the 7B base version of the Gemma-Mling model, continual pretrained on mainly Korean/English/Chinese/Japanese + 500 multilingual corpus.
Resources and Technical Documentation:
Terms of Use: Terms
Citation
@misc {gemma_mling_7b,
author = { {Junbum Lee, Taekyoon Choi} },
title = { gemma-mling-7b },
year = 2024,
url = { https://huggingface.co/beomi/gemma-mling-7b },
publisher = { Hugging Face }
}
Model Developers: Junbum Lee (Beomi) & Taekyoon Choi (Taekyoon)
Below we share some code snippets on how to get quickly started with running the model. First make sure to pip install -U transformers, then copy the snippet from the section that is relevant for your usecase.
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("beomi/gemma-mling-7b")
model = AutoModelForCausalLM.from_pretrained("beomi/gemma-mling-7b")
input_text = "머신러닝과 딥러닝의 차이는"
input_ids = tokenizer(input_text, return_tensors="pt")
outputs = model.generate(**input_ids)
print(tokenizer.decode(outputs[0]))
# pip install accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("beomi/gemma-mling-7b")
model = AutoModelForCausalLM.from_pretrained("beomi/gemma-mling-7b", device_map="auto")
input_text = "머신러닝과 딥러닝의 차이는"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
print(tokenizer.decode(outputs[0]))
Details about the model internals.
Training was done using beomi/Gemma-EasyLM.
We trained a mixture of multiple language datasets and trained until 100B. The released model is the best performance model based on our Evaluation below from model checkpoints.
For Korean and English datasets, we utilized sampled llama2ko training dataset which combined 1:1 ratio in each language.
| Dataset | Jsonl (GB) | Sampled |
|---|---|---|
| range3/cc100-ja | 96.39 | No |
| Skywork/SkyPile-150B | 100.57 | Yes |
| llama2ko dataset (ko/en) | 108.5 | Yes |
| cis-lmu/Glot500 | 181.24 | No |
| Total | 486.7 | . |
Model evaluation metrics and results.
!git clone https://github.com/EleutherAI/lm-evaluation-harness.git
!cd lm-evaluation-harness && pip install -r requirements.txt && pip install -e .
!lm_eval --model hf \
--model_args pretrained=beomi/gemma-mling-7b,dtype="float16" \
--tasks "haerae,kobest,kmmlu_direct,cmmlu,ceval-valid,mmlu,xwinograd,xcopa \
--num_fewshot "0,5,5,5,5,5,0,5" \
--device cuda
jp-stable branch)
!git clone -b jp-stable https://github.com/Stability-AI/lm-evaluation-harness.git
!cd lm-evaluation-harness && pip install -e ".[ja]"
!pip install 'fugashi[unidic]' && python -m unidic download
!cd lm-evaluation-harness && python main.py \
--model hf-causal \
--model_args pretrained=beomi/gemma-mling-7b,torch_dtype='auto'"
--tasks "jcommonsenseqa-1.1-0.3,jnli-1.3-0.3,marc_ja-1.1-0.3,jsquad-1.1-0.3,jaqket_v2-0.2-0.3,xlsum_ja,mgsm"
--num_fewshot "3,3,3,2,1,1,5"
| Category | Metric | Shots | Score |
|---|---|---|---|
| Default Metric | ACC | ||
| Knowledge (5-shot) | MMLU | 61.76 | |
| KMMLU (Exact Match) | 42.75 | ||
| CMLU | 50.93 | ||
| JMLU | |||
| C-EVAL | 50.07 | ||
| HAERAE | 0-shot | 63.89 | |
| KoBest (5-shot) | BoolQ | 85.47 | |
| COPA | 83.5 | ||
| Hellaswag (acc-norm) | 63.2 | ||
| Sentineg | 97.98 | ||
| WiC | 70.95 | ||
| XCOPA (5-shot) | IT | 72.8 | |
| ID | 76.4 | ||
| TH | 60.2 | ||
| TR | 65.6 | ||
| VI | 77.2 | ||
| ZH | 80.2 | ||
| JP Eval Harness (Prompt ver 0.3) | JcommonsenseQA | 3-shot | 85.97 |
| JNLI | 3-shot | 39.11 | |
| Marc_ja | 3-shot | 96.48 | |
| JSquad (Exact Match) | 2-shot | 70.69 | |
| Jaqket (Exact Match) | 1-shot | 81.53 | |
| MGSM | 5-shot | 28.8 | |
| XWinograd (0-shot) | EN | 89.03 | |
| FR | 72.29 | ||
| JP | 82.69 | ||
| PT | 73.38 | ||
| RU | 68.57 | ||
| ZH | 79.17 |
These models have certain limitations that users should be aware of.
Open Large Language Models (LLMs) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.
The development of large language models (LLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
Risks identified and mitigations:
The training is supported by TPU Research Cloud program.
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 11.18 |
| IFEval (0-Shot) | 20.29 |
| BBH (3-Shot) | 17.63 |
| MATH Lvl 5 (4-Shot) | 4.15 |
| GPQA (0-shot) | 0.00 |
| MuSR (0-shot) | 6.85 |
| MMLU-PRO (5-shot) | 18.14 |