Downloads · 30 days
28
32% of all-time downloads
lumasik/SLIM-1-base
SLIM-1-base is a text generation model from lumasik. Use it when you need the model to write or continue text. The card lists the license as mit.
<div align="center" <img src="https://huggingface.co/lumasik/SLIM-1-base/resolve/main/prewiev.png" alt="Preview" width="400" <br <i"Small language intelligent models"</i </div
Downloads · 30 days
28
32% of all-time downloads
All-time downloads
88
Public
Parameters
185M
1.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.pt742 MB · 66%
From the Hugging Face model README
Task: Text generation (base to instruct & thinking later)
Parameters: 175M
Context window: 2048 tokens
Training data: FineWeb Edu, RefinedWeb, OpenWebMath, TinyCodes, CodeSearchNet, Cosmopedia (total: 7b tokens)
Hyperparams: LR 4e-4, effective batch 126, per_device_batch_size=6, gradient_accumulation_steps=21
Expected loss: train - 2.765, valid - 2.774
Framework: PyTorch + HuggingFace Transformers
Training progress:

How to launch:
import torch
from transformers import PreTrainedTokenizerFast, LlamaForCausalLM
from tokenizers import Tokenizer
tk = PreTrainedTokenizerFast(tokenizer_object=Tokenizer.from_file("tokenizer.json"), bos_token="<|bos|>", eos_token="<|eos|>", unk_token="<|unk|>", pad_token="<|pad|>")
md = LlamaForCausalLM.from_pretrained(".", torch_dtype=torch.bfloat16).to("cuda")
ids = tk.encode("<|bos|>The future of AI is", return_tensors="pt").to("cuda")
gen = md.generate(ids, max_new_tokens=150, do_sample=True, temperature=0.7, repetition_penalty=1.2, no_repeat_ngram_size=3, pad_token_id=0)
print(tk.decode(gen[0], skip_special_tokens=True))