Downloads · 30 days
76
0% of all-time downloads
beomi/gemma-ko-2b
gemma-ko-2b is a text generation model from beomi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Update @ 2024.03.26: First release of Gemma-Ko 2B model
Downloads · 30 days
76
0% of all-time downloads
All-time downloads
77.2K
Public
Parameters
2.5B
5 GB on disk
Likes
43
Public
Click a slice to open those files.
.safetensors5 GB · 100%
From the Hugging Face model README
Update @ 2024.03.26: First release of Gemma-Ko 2B model
Original Gemma Model Page: Gemma
This model card corresponds to the 2B base version of the Gemma-Ko model.
Resources and Technical Documentation:
Terms of Use: Terms
Citation
@misc {gemma_ko_7b,
author = { {Junbum Lee, Taekyoon Choi} },
title = { gemma-ko-7b },
year = 2024,
url = { https://huggingface.co/beomi/gemma-ko-7b },
doi = { 10.57967/hf/1859 },
publisher = { Hugging Face }
}
Model Developers: Junbum Lee (Beomi) & Taekyoon Choi (Taekyoon)
Summary description and brief definition of inputs and outputs.
Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. They are text-to-text, decoder-only large language models, available in English, with open weights, pre-trained variants, and instruction-tuned variants. Gemma models are well-suited for a variety of text generation tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as a laptop, desktop or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone.
Below we share some code snippets on how to get quickly started with running the model. First make sure to pip install -U transformers, then copy the snippet from the section that is relevant for your usecase.
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("beomi/gemma-ko-2b")
model = AutoModelForCausalLM.from_pretrained("beomi/gemma-ko-2b")
input_text = "머신러닝과 딥러닝의 차이는"
input_ids = tokenizer(input_text, return_tensors="pt")
outputs = model.generate(**input_ids)
print(tokenizer.decode(outputs[0]))
# pip install accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("beomi/gemma-ko-2b")
model = AutoModelForCausalLM.from_pretrained("beomi/gemma-ko-2b", device_map="auto")
input_text = "머신러닝과 딥러닝의 차이는"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
print(tokenizer.decode(outputs[0]))
First make sure to install flash-attn in your environment pip install flash-attn
model = AutoModelForCausalLM.from_pretrained(
"beomi/gemma-ko-2b",
torch_dtype=torch.float16,
+ attn_implementation="flash_attention_2"
).to(0)
Details about the model internals.
Training was done using beomi/Gemma-EasyLM.
Model evaluation metrics and results.
TBD
These models have certain limitations that users should be aware of.
Open Large Language Models (LLMs) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.
The development of large language models (LLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
Risks identified and mitigations:
The training is supported by TPU Research Cloud program.