Downloads · 30 days
69
7% of all-time downloads
lugman/Droplet-v2
Droplet-v2 is a text generation model from lugman. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Droplet-v2 is a small, pretrained causal language model with 1,405,008 parameters. It uses the LlamaForCausalLM architecture with six transformer layers, a hidden size of 144, and a vocabulary of 1,536 tokens.
Downloads · 30 days
69
7% of all-time downloads
All-time downloads
968
Public
Parameters
1.4M
299 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt181 MB · 65%
From the Hugging Face model README
Droplet-v2 is a small, pretrained causal language model with 1,405,008 parameters. It uses the LlamaForCausalLM architecture with six transformer layers, a hidden size of 144, and a vocabulary of 1,536 tokens.
The model is a base checkpoint and has not been instruction-tuned. Its maximum context length is 1,024 tokens.
The following results were measured with zero-shot evaluation. The general benchmarks used lm-eval 0.4.12. ArithMark-3 was evaluated separately using its benchmark script.
| Benchmark | Metric | Score |
|---|---|---|
| ARC-Easy | acc_norm | 29.80% |
| ARC-Challenge | acc_norm | 22.78% |
| HellaSwag | acc_norm | 27.42% |
| PIQA | acc_norm | 53.54% |
| BoolQ | acc | 40.64% |
| SciQ | acc_norm | 53.30% |
| ArithMark-3 | acc_norm | 31.80% |
Evaluation was performed in float32 with a maximum context length of 1,024 tokens.
The main pretraining mixture contains approximately 6 billion tokens.
| Source | Tokens | Share | Role |
|---|---|---|---|
fineweb-edu (sample-10BT) | 2.70B | 45% | Educational web text |
Ultra-FineWeb-L3-en-Multi-Style-Synthetic | 1.80B | 30% | Synthetic multi-style text |
cosmopedia-v2 | 1.50B | 25% | Synthetic educational text |
An additional approximately 30 million tokens from lugman/add-sub-pre-training were used for arithmetic-focused pretraining.