Downloads · 30 days
35
9% of all-time downloads
Serdar404/RecGPT-100M-Fixed
RecGPT-100M-Fixed is a text generation model from Serdar404. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
RecGPT-100M is a 124.03M-parameter recursive causal language model trained for the BabyLM 2026 Strict track. It was trained for 10 epochs on a custom 100M-word English corpus using a 32,768-token BPE vocabulary.
Downloads · 30 days
35
9% of all-time downloads
All-time downloads
369
Public
Parameters
124M
13.9 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors496 MB · 100%
From the Hugging Face model README
RecGPT-100M is a 124.03M-parameter recursive causal language model trained for the BabyLM 2026 Strict track. It was trained for 10 epochs on a custom 100M-word English corpus using a 32,768-token BPE vocabulary.
The model applies a shared Transformer block recursively for 24 iterations. Its hidden size is 1,408, embedding size is 768, and feed-forward intermediate size is 22,528. Training used Aurora for the recursive block and AdamW for the embedding-related parameters, with a token batch size of 32,768 and sequence length 512.
This repository contains custom Transformers code, so loading requires trust_remote_code=True:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Serdar404/RecGPT-100M-Fixed"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
The model is intended for scoring text as a causal language model. KV-cache generation is not currently implemented.
Final-checkpoint results before leaderboard submission:
| Evaluation | Score |
|---|---|
| BLiMP | 80.06 |
| BLiMP Supplement | 69.28 |
| EWoK | 59.05 |
| Entity Tracking | 19.50 |
| COMPS | 60.55 |
| GlobalPIQA | 43.15 |
| (Super)GLUE | 71.84 |
Intermediate Strict checkpoints are published as Hub revisions named chck_1M through chck_1000M using the official BabyLM checkpoint schedule.
This is a small research model trained under the BabyLM data constraint. It is not intended for production deployment, factual question answering, or safety-critical use. Its outputs may contain inaccuracies or undesirable content inherited from its training data.