Downloads · 30 days
405
100% of all-time downloads
basically-experimental/Pebble-50M-Chat-beta
Pebble-50M-Chat-beta is a text generation model from basically-experimental. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Pebble-50M-Chat-beta is the chat-tuned version of Pebble-50M-beta, an experimental 50M-parameter language model designed to test how a larger Pebble architecture performs with a 16,384-token vocabulary and 16,384-toke…
Downloads · 30 days
405
100% of all-time downloads
All-time downloads
405
Public
Parameters
49.3M
2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors98.7 MB · 98%
How the weights are stored.
BF1649.3M · 100%
From the Hugging Face model README
Pebble-50M-Chat-beta is the chat-tuned version of Pebble-50M-beta, an experimental 50M-parameter language model designed to test how a larger Pebble architecture performs with a 16,384-token vocabulary and 16,384-token context window.
The base model underperformed Pebble-25M and, on some evaluations, Pebble-10M. Pebble-50M-Chat-beta was subsequently fine-tuned to improve its ability to follow instructions and engage in conversational interactions.
The base model was trained on a 25B-token subset of the following datasets:
| Dataset | Token Allocation | Share |
|---|---|---|
| FineWeb-Edu | 7.50 billion | 30% |
| DCLM | 5.00 billion | 20% |
| Cosmopedia-v2 | 3.75 billion | 15% |
| FineMath-4+ | 3.75 billion | 15% |
| FinePhrase | 3.00 billion | 12% |
| NPset | 2.00 billion | 8% |
| Total | 25.00 billion | 100% |
Pebble-50M-Chat-beta was fine-tuned on approximately 250M tokens from Smol-SmolTalk to improve conversational ability and instruction following.
The original benchmark logs for the base model were lost, so exact evaluation results are unavailable.
The chat model is primarily intended for conversational use and should not be directly compared with the base model on benchmarks without considering the effects of fine-tuning.
Pebble-50M-Chat-beta does not require the mamba-ssm library and is intended to be usable with standard PyTorch-based inference implementations.
It may run on CUDA GPUs, AMD GPUs, Intel GPUs, and CPUs depending on the inference framework and available hardware acceleration.
This is a beta/experimental model. It is primarily intended for research, experimentation, and conversational use.
Apache 2.0