Downloads · 30 days
33
100% of all-time downloads
SedimentLabs/Pebble-1-30B
Pebble-1-30B is a text generation model from SedimentLabs. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A 30B reasoning model that knows what it knows. Pebble 1 30B answers factual questions when it is confident and says "I don't know" when it is not. Fine-tuned by Sediment from Meta's Muse-Glimmer-30B.
Downloads · 30 days
33
100% of all-time downloads
All-time downloads
33
Public
Parameters
29.8B
59.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors59.6 GB · 100%
From the Hugging Face model README
A 30B reasoning model that knows what it knows. Pebble 1 30B answers factual questions when it is confident and says "I don't know" when it is not. Fine-tuned by Sediment from Meta's Muse-Glimmer-30B.
reasoning_content and only the answer as content.| Attribute | Value |
|---|---|
| Developer | Sediment, led by Asa Shepard |
| Base model | Meta Muse-Glimmer-30B |
| Parameters | 29.8B (dense) |
| Context length | 131,072 tokens |
| Architecture | 52 layers, GQA (32 query heads, 2 KV heads), 202K vocabulary |
| Precision | bfloat16 safetensors, 59.6 GB |
| Hardware | One 80 GB GPU, or two 48 GB GPUs with tensor parallelism |
| Languages | English (evaluated). The base model is multilingual |
| Input | Text. The base model accepts images, but Pebble 1 30B was tuned and evaluated on text only |
| Version | 1.0, September 2026 |
AA-Omniscience index = 100 × (correct - incorrect) / N, so abstaining scores zero. Measured on the 600 public questions with a replica of the official grading rubric (four classes, GPT-5.4 as judge). The replica matches the published base-model score within 1.3 points. Intervals are 95% bootstrap.
| Model | Index | Accuracy | Hallucination |
|---|---|---|---|
| Muse-Glimmer-30B, official | -33 | 27.0 | 81.9 |
| Muse-Glimmer-30B, replica judge | -31.7 | 26.7 | 79.5 |
| Pebble 1 30B | +10.3 [6.8, 14.0] | 15.7 | 6.3 |

A second judge (Gemini 3.6 Flash, same rubric) scores Pebble 1 30B at +11.5 / 17.0 / 6.6. Official numbers will differ slightly, since the official judge and question set are not the ones used here.
Accuracy is lower than the base model's by design. The model declines questions it would sometimes get right by guessing. Use a different model if you need a guess on every question.
Requires vLLM 0.28.0 or newer, which includes the muse_glimmer reasoning parser. One 80 GB GPU, or two 48 GB GPUs with --tensor-parallel-size 2.
vllm serve SedimentLabs/Pebble-1-30B \
--served-model-name pebble-1-30b \
--reasoning-parser muse_glimmer \
--max-model-len 16384 \
--generation-config auto
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
SYSTEM = "You are answering questions about general knowledge. Answer with JUST the answer (no explanation). If you do not know the answer, or you need more context or tools to answer the question, be clear about this - it is better that you say this than get the wrong answer."
r = client.chat.completions.create(
model="pebble-1-30b",
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Who won the 1931 Tour de Suisse?"}],
max_tokens=5000,
)
print(r.choices[0].message.content) # "I don't know."
print(r.choices[0].message.reasoning_content) # the model's thinking
The architecture is registered under the image-text-to-text auto class because the base model is multimodal. Load it with AutoModelForImageTextToText. Requires transformers 5.16 or newer.
from transformers import AutoTokenizer, AutoModelForImageTextToText
model_id = "SedimentLabs/Pebble-1-30B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
SYSTEM = "You are answering questions about general knowledge. Answer with JUST the answer (no explanation). If you do not know the answer, or you need more context or tools to answer the question, be clear about this - it is better that you say this than get the wrong answer."
messages = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Which element has atomic number 74?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=5000, do_sample=True, temperature=0.6)
text = tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
answer = text.split("to=user<|message|>")[-1].split("<|eot|>")[0].strip()
print(answer) # "Tungsten"
to=self channel, then the answer in a to=user channel, ended by <|eot|>. vLLM splits the channels for you. With Transformers, take the text after the last to=user<|message|>.Pebble 1 30B starts from Muse-Glimmer-30B with reasoning enabled and adds a calibration stage: the base model's own knowledge is mapped question by question, then supervised fine-tuning and preference optimization teach it to answer when it knows and abstain when it does not, without shortening its reasoning. Training data is a generated and independently verified bank of obscure factual questions, decontaminated against the benchmark. The AA-Omniscience public questions were used only for held-out evaluation, under a pre-registered budget of looks.
Apache 2.0. The base model, Meta's Muse-Glimmer-30B, is released under Apache 2.0 as well.
@misc{pebble1-30b,
title = {Pebble 1 30B: a calibrated reasoning model that knows what it knows},
author = {Sediment},
year = {2026},
url = {https://huggingface.co/SedimentLabs/Pebble-1-30B}
}