Downloads · 30 days
25
26% of all-time downloads
sciencemj/tinyllm-29m-chat
tinyllm-29m-chat is a text generation model from sciencemj. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as cc-by-nc-sa-4.0.
TinyStories 사전학습 후 DailyDialog 로 SFT 한 가중치. 짧은 대화를 주고받는다.
Downloads · 30 days
25
26% of all-time downloads
All-time downloads
97
Public
Parameters
29.6M
118 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors118 MB · 100%
From the Hugging Face model README
TinyStories 사전학습 후 DailyDialog 로 SFT 한 가중치. 짧은 대화를 주고받는다.
파라미터 29,577,728 개. RTX 3060 Ti 한 대에서 사전학습 3.06 시간 + SFT 2.4 분. 학습 코드와 설계 근거: https://github.com/sciencemj/tinyLLM
val loss 2.3975 (DailyDialog), 종료 토큰 준수율 100%
비상업 전용이다. DailyDialog 가 CC BY-NC-SA 4.0 이고 데이터셋 카드에 "Dataset provided for research purposes only" 라고 적혀 있다. 유료 서비스, 광고가 붙은 데모, 사내 제품 어디에도 쓸 수 없다. 재배포할 때는 출처를 표시하고 동일 조건으로 공개해야 한다.
상업적 이용이 필요하면 제약이 없는 사전학습 가중치(tinyllm-29m-tinystories)에서 시작해 2 단계 데이터만 허용적 라이선스로 교체하면 된다. SFT 는 2.4 분이다.
transformers 를 쓰지 않는다. 이 저장소의 modeling_tinyllm.py 하나면 된다.
import torch
from tokenizers import Tokenizer
from modeling_tinyllm import TinyLM
model = TinyLM.from_pretrained(".")
tok = Tokenizer.from_file("tokenizer.json")
ids = torch.tensor([tok.encode("<|user|>Hi, how are you today?<|eot|><|assistant|>").ids])
out = model.generate(ids, 60, temperature=0.6, top_k=20, stop_id=3)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
대화는 턴 표시가 필요하다. 모델이 실제로 읽는 것은 끊기지 않는 한 줄이다:
<|user|>Hi, how are you today?<|eot|><|assistant|>
끝의 <|assistant|> 가 있어야 모델이 이야기를 이어쓰는 대신 자기 차례로 답한다.
특수 토큰 id 는 <|endoftext|>=0, <|user|>=1, <|assistant|>=2, <|eot|>=3 이다.
ids (B, 512)
→ nn.Embedding(8000, 512) + nn.Embedding(512, 512)
→ nn.TransformerEncoder(
nn.TransformerEncoderLayer(512, nhead=8, dim_feedforward=2048,
activation="gelu", norm_first=True,
batch_first=True),
num_layers=8, norm=nn.RMSNorm(512))
→ nn.Linear(512, 8000, bias=False) # token embedding 과 tying
decoder-only 를 TransformerEncoderLayer 로 만든다. TransformerDecoderLayer 는
cross-attention 용 memory 를 필수로 요구해서 맞지 않는다.
토크나이저는 TinyStories 와 DailyDialog 합집합에서 학습한 자체 8k byte-level BPE 다.
같이 받은 tokenizer.json 을 반드시 써야 한다. 다른 토크나이저로는 동작하지 않는다.
된다 — 문법, 구두점, 따옴표 대화 형식, 문단 나누기, 인물 이름 유지, 인과 연결.
안 된다 — 턴 간 기억, 질문에 대한 직접 답변, 사실성, 문장 안 반복, 논리 일관성. 영어만 안다. 사실 정보를 얻는 용도로 쓰면 안 된다.
자세한 것은 MODEL_CARD.md.
@article{eldan2023tinystories,
title={TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
author={Eldan, Ronen and Li, Yuanzhi},
journal={arXiv preprint arXiv:2305.07759},
year={2023}
}
@InProceedings{li2017dailydialog,
author = {Li, Yanran and Su, Hui and Shen, Xiaoyu and Li, Wenjie and Cao, Ziqiang and Niu, Shuzi},
title = {DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset},
booktitle = {Proceedings of The 8th International Joint Conference on Natural Language Processing (IJCNLP 2017)},
year = {2017}
}