Downloads · 30 days
24
45% of all-time downloads
PeterRabbit/fineweb-100m-base
fineweb-100m-base is a text generation model from PeterRabbit. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A 97.5M-parameter decoder-only base language model trained from scratch on four billion FineWeb-Edu tokens.
Downloads · 30 days
24
45% of all-time downloads
All-time downloads
53
Public
Parameters
97.5M
390 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors390 MB · 100%
From the Hugging Face model README
A 97.5M-parameter decoder-only base language model trained from scratch on four billion FineWeb-Edu tokens.
This is a small educational and experimental model. It is not a production assistant and should not be relied on for factual, medical, legal, financial, safety-critical, or other consequential advice.
| Property | Value |
|---|---|
| Parameters | 97,536,768 |
| Layers | 12 |
| Hidden width | 768 |
| Attention heads | 12 |
| SwiGLU hidden width | 2048 |
| Vocabulary | 16,384 BPE tokens |
| Context window | 1,024 tokens |
| Architecture | Decoder-only Transformer |
| Position encoding | RoPE |
| Normalization | RMSNorm |
| Embeddings | Tied input/output embeddings |
The architecture was implemented from scratch in PyTorch. This repository includes custom Transformers-compatible configuration and modeling files.
The base model was trained from scratch on 4,000,055,296 packed tokens selected from the FineWeb-Edu sample-10BT corpus. Documents were selected deterministically by hashing document IDs rather than taking the first source shards. Text was tokenized with a locally trained byte-level BPE tokenizer, an <eos> token was appended after every document, and tokens were packed into fixed 1,024-token blocks.
Base training used AdamW, cosine learning-rate decay, gradient clipping, bfloat16 autocast, and a global batch size of 65,536 tokens on an NVIDIA RTX 3090.
Best validation loss: 2.8886 (next-token validation).
These losses are internal held-out validation measurements. Base and SFT losses use different objectives, and losses from the 8K and 16K tokenizers are not directly comparable. No broad academic benchmark suite or human preference evaluation was run.
Install the dependencies:
pip install -r requirements.txt
Run the included example after cloning this repository:
python generate.py --model . --prompt "The purpose of education is"
Or load it through Transformers:
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
".",
trust_remote_code=True,
)
trust_remote_code=True is required because this is a custom architecture. Review configuration_fineweb.py and modeling_fineweb.py before loading code from the Hub.
The source datasets are not redistributed in this model repository and remain subject to their respective licences and attribution requirements.
The model weights and original code in this repository are released under the Apache License 2.0. This does not replace or override the licences of the training datasets.