Downloads · 30 days
65
8% of all-time downloads
augustinian-babylm/deberta-base-50k
deberta-base-50k is a fill-mask model from augustinian-babylm. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-4.0.
Random-initialization baseline for the 50k vocabulary. This is the control that vision-seeded models with the same vocabulary are compared against.
Downloads · 30 days
65
8% of all-time downloads
All-time downloads
836
Public
Parameters
124M
17.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors498 MB · 99%
From the Hugging Face model README
Random-initialization baseline for the 50k vocabulary. This is the control that vision-seeded models with the same vocabulary are compared against.
Seeded counterparts: augustinian-babylm/deberta-base-50k-sam, -dinov3, -ibot.
Part of the model set for Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model (paper, code).
from transformers import AutoModelForMaskedLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("augustinian-babylm/babylm-bpe-50k")
model = AutoModelForMaskedLM.from_pretrained("augustinian-babylm/deberta-base-50k")
DeBERTa-v3-base masked LM (768 hidden, 12 layers, 12 heads), trained for 10 epochs
on bb24.train, the ~9.9M-word corpus of Edman et al. (2024), which mixes
LLM-generated paraphrase data with portions of the official BabyLM corpus. This is
not the official 2026 strict-small distribution.
| Learning rate | 2e-4, cosine schedule, 4000 warmup steps |
| Optimizer | AdamW (beta1 0.9, beta2 0.95), weight decay 0.01 |
| Effective batch | 256 (gradient accumulation 4) |
| Context length | 64 to 128 from epoch 5, max 512 |
| Vocabulary | BPE, 50k merges (augustinian-babylm/babylm-bpe-50k) |
Only the input-embedding initialization differs between a vision-init model and its baseline. Every other hyperparameter, the data order, and the random seed are held fixed.
Intermediate checkpoints are stored as branches, saved Pythia-style by step:
from transformers import AutoModelForMaskedLM
AutoModelForMaskedLM.from_pretrained("augustinian-babylm/deberta-base-50k", revision="step256")
Call huggingface_hub.list_repo_refs("augustinian-babylm/deberta-base-50k") to list the available steps.
Vision seeding produces a consistent object-property gain (COMPS, +1.30 averaged over nine encoder-by-vocabulary configurations, positive in all nine) while grammar-focused benchmarks stay flat. On a corpus-tailored Visual-Property Swap probe the seeded model leads the baseline in all three training seeds; a pair-level re-analysis shows that the advantage cannot be attributed to the individual seeded words. Details: https://github.com/bylinina/augustinian_babylm
Trained on a non-standard corpus (see above), so results are not directly comparable to submissions trained on the official 2026 strict-small data. Single masked-LM architecture. Research artifact, not intended for deployment.
@inproceedings{bylinina2026augustinian,
title = {Augustinian BabyLM: What Ostensive Definition Can and Cannot
Teach a Small Language Model},
author = {Bylinina, Lisa},
booktitle = {Proceedings of the BabyLM Workshop},
year = {2026},
url = {https://openreview.net/forum?id=B4TD4XdlwF}
}