Downloads · 30 days
20
5% of all-time downloads
camgeodesic/olmo3-7b-instruct-only
olmo3-7b-instruct-only is a text generation model from camgeodesic. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Fine-tuned from allenai/OLMo-3-1025-7B using GRPO (Group Relative Policy Optimization) on instruction-following tasks.
Downloads · 30 days
20
5% of all-time downloads
All-time downloads
421
Public
Parameters
528K
248 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors14.6 GB · 100%
From the Hugging Face model README
Fine-tuned from allenai/OLMo-3-1025-7B using GRPO (Group Relative Policy Optimization) on instruction-following tasks.
if_valley_thinker — valley length penalty (512–4096 token sweet spot) + think token reward shaping<think> tag for chain-of-thought reasoning)| Component | Description |
|---|---|
| IFEval verifiable reward | Binary per-constraint score for instruction-following |
| Valley length penalty | Penalizes responses <512 or >4096 tokens (coeff: -0.001) |
| Think tag reward | +0.125 for correct </think> closure |
| Think length penalty | -0.1 if thinking block <10 words |
| Metric | Value |
|---|---|
| IFEval correct rate | 0.88 |
| Training reward | 6.36 |
| Think word count | ~886 words |
| Sequence length | ~1353 tokens |
Each training checkpoint is available as a separate branch/revision:
main — step 3800 (latest)step_600 through step_3600 — intermediate checkpoints (every 200 steps)from transformers import AutoModelForCausalLM, AutoTokenizer
# Load latest
model = AutoModelForCausalLM.from_pretrained("camgeodesic/olmo3-7b-instruct-only")
# Load specific checkpoint
model = AutoModelForCausalLM.from_pretrained("camgeodesic/olmo3-7b-instruct-only", revision="step_2000")