Downloads · 30 days
52
3% of all-time downloads
SII-GAIR-NLP/davinci-llm-model
davinci-llm-model is a text generation model from SII-GAIR-NLP. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
daVinci-LLM-3B is a 3B-parameter base language model presented in daVinci-LLM: Towards the Science of Pretraining. This project aims to make the pretraining process a transparent and reproducible scientific endeavor.
Downloads · 30 days
52
3% of all-time downloads
All-time downloads
1.7K
Public
Parameters
3.4B
673 GB on disk
Likes
30
Public
Click a slice to open those files.
.safetensors6.8 GB · 100%
From the Hugging Face model README
daVinci-LLM-3B is a 3B-parameter base language model presented in daVinci-LLM: Towards the Science of Pretraining. This project aims to make the pretraining process a transparent and reproducible scientific endeavor.
We release not only the final weights but also training trajectories, intermediate checkpoints, data processing decisions, and 200+ ablation studies covering data quality, mixture design, training dynamics, and evaluation validity.
The model follows a two-stage curriculum over ~8T tokens:
This is a base model and is not instruction- or safety-aligned. Additional safety evaluation and alignment are required for production deployment.
The training corpus spans general web text, code, science, and QA sources. Each dataset is annotated with a Data Darwinism level (L0–L9), and multiple sources receive L4/L5 generative refinement and cognitive completion.
Major categories:
The model reaches an overall average score of 51.72 across 19 benchmarks, matching or exceeding the performance of larger 7B-scale models like OLMo-3 7B.
| Capability Dimension | daVinci-3B | OLMo-3 7B | LLaMA-3.2-3B | Qwen-2.5-3B |
|---|---|---|---|---|
| Overall Performance | 51.72 | 51.65 | 37.58 | 51.44 |
| General Knowledge | 52.96 | 55.13 | 51.08 | 55.16 |
| Code Generation | 55.99 | 54.42 | 32.40 | 56.13 |
| Scientific Reasoning | 48.30 | 45.98 | 22.45 | 44.65 |
| MATH | 62.80 | 39.60 | 9.00 | 37.20 |
@misc{qin2026davincillmtowardssciencepretraining,
title={daVinci-LLM:Towards the Science of Pretraining},
author={Yiwei Qin and Yixiu Liu and Tiantian Mi and Muhang Xie and Zhen Huang and Weiye Si and Pengrui Lu and Siyuan Feng and Xia Wu and Liming Liu and Ye Luo and Jinlong Hou and Qipeng Guo and Yu Qiao and Pengfei Liu},
year={2026},
eprint={2603.27164},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2603.27164},
}