Downloads · 30 days
120
13% of all-time downloads
Boldt/Boldt-1B
Boldt-1B is a text generation model from Boldt. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
120
13% of all-time downloads
All-time downloads
936
Public
Parameters
1.2B
2.4 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors2.4 GB · 100%
From the Hugging Face model README
Boldt is a series of German Small Language Models (SLMs) trained from scratch. Our inital release includes four models:
The training philosophy behind Boldt is centered on a key finding from our research: repetition over diversity.
Standard pre-training paradigms typically balance quality filtering against the need for massive token volume and broad corpus diversity. In contrast, Boldt models are trained for multiple epochs on a highly filtered dataset: the German Dense-Core subset of FineWeb-2. We isolated this subset using a combination of three hierarchical filters:
We demonstrate that repeated exposure to this strict, high-quality subset is more sample-efficient than a single pass over less filtered and more diverse corpora. For a comprehensive look at our experiments, please refer to our preprint: Repetition over Diversity.
Boldt-1B builds upon the foundation of Boldt-DC-1B by adding 6B tokens of premium German news data, collected continuously since 2022 via the Fundus library. To support complex downstream tasks, it also features a doubled context window of 4096 tokens.
Note: This is a base language model, not an instruction-tuned model. It is not optimized for chat or instruction following. For best results, use standard text completion rather than chat templates.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "Boldt/Boldt-1B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
# Basic text completion
text = "Berlin ist eine Stadt, wo"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=64)

We evaluate Boldt-1B on our modernized German benchmark suite. See our paper (Aynetdinov et al., 2026) for details on the structural and translation corrections we performed.
Despite being trained on substantially fewer tokens, the Boldt-1B family outperforms other 1B-class models on German tasks and performs competitively with much larger multilingual models.
Note: Bold text indicates the best score in the 1B category.
| Model | Tokens | MMLU | ARC-C | ARC-E | H-Swag | LAMBADA | OBQA | Avg. |
|---|---|---|---|---|---|---|---|---|
| Boldt-DC-350M | 200B | 29.29 | 32.24 | 52.87 | 43.21 | 37.48 | 45.86 | 40.16 |
| Boldt-DC-1B | 200B | 31.06 | 35.99 | 57.30 | 48.69 | 42.80 | 48.48 | 44.05 |
| Boldt-1B (this model) | 230B | 31.42 | 34.11 | 55.78 | 48.77 | 44.70 | 52.32 | 44.52 |
| LLäMmlein-1B | 1T | 29.26 | 30.27 | 48.19 | 44.80 | 44.89 | 47.27 | 40.78 |
| Gemma-3-1B | 2T* | 30.01 | 30.55 | 47.89 | 43.43 | 41.71 | 45.05 | 39.77 |
| Llama-3.2-1B | 9T* | 28.58 | 29.90 | 40.51 | 40.07 | 44.31 | 44.04 | 37.90 |
| Qwen3.5-0.8B-Base | >36T* | 30.79 | 32.05 | 46.20 | 38.90 | 36.02 | 43.84 | 37.97 |
| Model | Tokens | MMLU | ARC-C | ARC-E | H-Swag | LAMBADA | OBQA | Avg. |
|---|---|---|---|---|---|---|---|---|
| EuroLLM-1.7B | 4T* | 31.04 | 31.58 | 54.68 | 45.30 | 44.52 | 50.50 | 42.94 |
| Qwen3-1.7B-Base | 36T* | 34.17 | 37.49 | 57.00 | 45.20 | 49.81 | 45.66 | 44.89 |
| BübleLM-2B | 2T* | 29.68 | 32.62 | 53.63 | 46.57 | 43.55 | 49.70 | 42.63 |
| Gemma-2-2B | 2T* | 33.99 | 37.11 | 57.47 | 49.62 | 52.64 | 48.89 | 46.62 |
| Gemma-4-E2B | N/A | 34.48 | 41.14 | 63.16 | 55.22 | 55.96 | 50.51 | 50.08 |
We have not conducted systematic model evaluations of toxicity, demographic biases, or harmful stereotypes. Quality filtering may reduce some risks relative to unfiltered web data, but cannot guarantee their absence, and repeated exposure during multi-epoch training could amplify rather than mitigate encoded biases. Users should exercise caution in sensitive use-cases without further evaluation.
@misc{boldt2026,
title={Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling},
author={Ansar Aynetdinov and Patrick Haller and Alan Akbik},
year={2026},
eprint={2604.28075},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.28075},
}