Downloads · 30 days
64
2% of all-time downloads
go76dof/wwm_curriculum_simplification_40k
wwm_curriculum_simplification_40k is a fill-mask model from go76dof. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as other.
A 34.7M-parameter DeBERTa-style masked language model trained on 10M words of FineWeb simplification pairs for the BabyLM 2026 strict-small track.
Downloads · 30 days
64
2% of all-time downloads
All-time downloads
3.5K
Public
Parameters
34.7M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors69.4 MB · 99%
From the Hugging Face model README
A 34.7M-parameter DeBERTa-style masked language model trained on 10M words of FineWeb simplification pairs for the BabyLM 2026 strict-small track.
This model is a BabyLM 2026 strict-small submission trained with meaning-preserving simplification pairs and a whole-word-masking curriculum.
The model is trained on original FineWeb sentences paired with simplified rewrites. During masked language model pretraining, these paired examples expose the model to two aligned ways of expressing similar content. The goal is to improve sample efficiency under the 10M-word BabyLM strict-small budget.
go76dof/wwm_curriculum_simplification_40kgo76dof/Fineweb_simplification_pairsThis is a pretrained masked language model. It can be used for:
This model is not intended as a general-purpose production language model. It is small and trained on only 10M words.
Example loading code:
from transformers import AutoModelForMaskedLM, AutoTokenizer
model_name = "go76dof/wwm_curriculum_simplification_40k"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForMaskedLM.from_pretrained(model_name)
The model was trained on FineWeb sentence-level simplification pairs:
go76dof/Fineweb_simplification_pairsFineWeb_simplification_pairs.trainEach pair is stored as an original sentence followed by its simplified rewrite, with blank lines separating pairs.
Example:
Enlightenment thinkers proposed that human reason coupled with empirical study of the physical world would lead to progress---namely, the advancement of science and the improvement of the human condition.
Enlightenment thinkers believed that using reason and studying the world would lead to scientific progress and better living conditions.
The tokenizer is a 40k SentencePiece BPE tokenizer trained for this data condition:
tokenizer/FineWeb_simplification_pairs_40k.modeltokenizer/FineWeb_simplification_pairs_40k.vocab| Hyperparameter | Value |
|---|---|
| Architecture | DeBERTa-v2 style encoder |
| Hidden size | 384 |
| Intermediate size | 1280 |
| Number of layers | 12 |
| Number of attention heads | 12 |
| Dropout | 0.1 |
| Vocabulary size | 40,000 |
| Objective | Masked language modeling |
| Optimizer | LAMB |
| Learning rate schedule | cosine |
| Max learning rate | 0.007 |
| Training epochs | 10 |
| Sequence length curriculum | 64 -> 256 |
| Masking curriculum | WWM7 -> Token3 |
The model uses a whole-word-masking curriculum:
The sequence length also follows a curriculum. Training starts with shorter sequences and later switches to longer sequences, while batch size is scaled inversely with sequence length.
The model has 34,677,952 parameters. The final checkpoint corresponds to the end of the 10-epoch training run on the 10M-word strict-small training budget.
This model was evaluated with the BabyLM 2026 strict-small evaluation pipeline. The reported submission scores are leaderboard scores for the strict-small track.
The evaluation includes:
| Metric | Score |
|---|---|
| Overall Average | 41.80 |
| NLP Average | 52.97 |
| Human-like Average | 2.71 |
| BLiMP | 67.20 |
| BLiMP Supplement | 56.01 |
| EWoK | 56.07 |
| Entity Tracking | 28.45 |
| COMPS | 53.57 |
| GlobalPIQA | 39.67 |
| (Super)GLUE | 69.79 |
| Reading | 5.42 |
| AoA | 0.00 |
The strongest gains in our experiments were observed on EWoK, Entity Tracking, and GLUE-style fine-tuning. In our BabyLM 2026 strict-small submission, this 10M-word model exceeded the official 100M-word GPT-2 baseline on EWoK, Entity Tracking, GlobalPIQA, and (Super)GLUE.
The model is a DeBERTa-v2 style encoder trained with the masked language modeling objective.
Important configuration values:
architectures: DebertaV2ForMaskedLMmodel_type: deberta-v2hidden_size: 384intermediate_size: 1280num_hidden_layers: 12num_attention_heads: 12vocab_size: 40000relative_attention: truepos_att_type: p2c, c2pThe model is compatible with Hugging Face Transformers and PyTorch.
This model is trained only on English text and only under a 10M-word pretraining budget. It is intended for research use rather than production deployment.
The training data is derived from web text and automatically generated simplifications. It may contain noise, omissions, or changes in nuance inherited from either the source text or the rewrite process.
The model card reports BabyLM evaluation results, but some hidden-task results, especially GlobalPIQA and AoA, can be unstable across seeds or evaluation settings.
If you use this model, please cite the corresponding BabyLM paper or repository once available.
Shaoxiang Wu