Downloads · 30 days
154
0% of all-time downloads
NucleusAI/nucleus-22B-token-500B
nucleus-22B-token-500B is a text generation model from NucleusAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Nucleus-22B-token-500B is a 22B parameters causal decoder-only model built by Nucleus.AI and trained on 500B tokens of RefinedWeb along with curated corpora. It is made available under the MIT license.
Downloads · 30 days
154
0% of all-time downloads
All-time downloads
46.1K
Public
Parameters
21.8B
87.3 GB on disk
Likes
27
Public
Click a slice to open those files.
.safetensors87.3 GB · 100%
From the Hugging Face model README
Nucleus-22B-token-500B is a 22B parameters causal decoder-only model built by Nucleus.AI and trained on 500B tokens of RefinedWeb along with curated corpora. It is made available under the MIT license.
1T-token model coming soon 😊.
⚠️ This is a raw, pretrained model, which should be further finetuned for most usecases.
Research on large language models; as a foundation for further specialization and finetuning for specific usecases (e.g., summarization, text generation, chatbot, etc.)
Production use without adequate assessment of risks and mitigation; any use cases which may be considered irresponsible or harmful.
Nucleus-22B-token-500B is trained on English data only, and will not generalize appropriately to other languages. Furthermore, as it is trained on a large-scale corpora representative of the web, it will carry the stereotypes and biases commonly encountered online.
We recommend users of Nucleus-22B-token-500B to consider finetuning it for the specific set of tasks of interest, and for guardrails and appropriate precautions to be taken for any production use.
Nucleus-22B-token-500B was trained on 500B tokens of RefinedWeb, along with other corpora.
| Data source | Fraction | Tokens | Sources |
|---|---|---|---|
| RefinedWeb-English | 75% | 200B | massive web crawl |
| Books | 7% | 21B | |
| Code | 7% | 21B | Big Code, CodeNet |
| Technical | 6% | 19B | arXiv |
| Math | 5% | 17B | Mathematica, Khan Academy |
The data was tokenized with the tokenizer similar to Llama-7B.
Nucleus-22B-token-500B was trained on 256 A100 80GB GPUs, using a FSDP
| Hyperparameter | Value | Comment |
|---|---|---|
| Precision | bfloat16 | |
| Optimizer | AdamW | |
| Learning rate | 2e-4 | 8B tokens warm-up, cosine decay to 1.e-5 |
| Weight decay | 1e-1 | |
| Batch size | 2048 | constant |
| Context length | 2048 | constant |
Training happened in early August 2023 and took about two weeks.