Downloads · 30 days
23
23% of all-time downloads
helloadhavan/llara1.1-100M-base
llara1.1-100M-base is a text generation model from helloadhavan. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<img src="data:image/svg+xml;base64,PD94bWwgdmVyc2lvbj0iMS4wIiBlbmNvZGluZz0iVVRGLTgiPz4KPHN2ZyB2ZXJzaW9uPSIxLjEiIHhtbG5zPSJodHRwOi8vd3d3LnczLm9yZy8yMDAwL3N2ZyIgd2lkdGg9IjUwMCIgaGVpZ2h0PSIyMDAiIHN0eWxlPSJiYWNrZ3JvdW5kL…
Downloads · 30 days
23
23% of all-time downloads
All-time downloads
99
Public
Parameters
124M
2.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.pt1.7 GB · 56%
From the Hugging Face model README
Llara1.1 is a 124M parameter (33M params more than llara1.0) autoregressive language model trained from scratch on English web text. It follows the GPT-2 Small architecture and is trained entirely from random initialisation — no pretrained weights, no distillation, no fine-tuning of an existing model. but it does use GPT's tokenizer (sorta)
The name Llara is original and unrelated to LLaMA or LoRA.
Note: The model is stil undertrained according to The Chinchilla Laws (2022)
| Property | Value |
|---|---|
| Architecture | GPT-2 (decoder-only transformer) |
| Parameters | ~124.0M |
| Context length | 512 tokens |
| Embedding dim | - |
| Layers | 12 |
| Attention heads | 12 |
| Vocabulary | 50,257 (GPT-2 BPE) |
| Training data | FineWeb (HuggingFaceFW/fineweb) + Custom dataset |
| Training docs | 131M tokens |
| Epochs | 1.1 |
| Precision | fp16 |
from transformers import GPT2LMHeadModel, AutoTokenizer, pipeline
model = GPT2LMHeadModel.from_pretrained("helloadhavan/llara1.1-100M-base")
tokenizer = AutoTokenizer.from_pretrained("helloadhavan/llara1.1-100M-base")
gen = pipeline("text-generation", model=model, tokenizer=tokenizer)
output = gen(
"Once upon a time",
max_new_tokens=20,
do_sample=True,
temperature=0.8,
top_p=0.95,
repetition_penalty=1.1,
)
print(output[0]["generated_text"])
Llara is intended for:
Trained using Hugging Face Transformers Trainer on a single GPU.
Apache 2.0
<div> <blockquote><strong>Note:</strong> i am a AI hobbyist, not an AI engineer</blockquote> </div>