Downloads · 30 days
138
9% of all-time downloads
ProCreations/grug-35b
grug-35b is a text generation model from ProCreations. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
grug brain now live in weight. no hidden style stick needed.
Downloads · 30 days
138
9% of all-time downloads
All-time downloads
1.5K
Public
Parameters
35.1B
96.2 GB on disk
Likes
16
Public
Click a slice to open those files.
.safetensors70.2 GB · 100%
From the Hugging Face model README
grug brain now live in weight. no hidden style stick needed.
July 15, 2026 weight replacement: same public repo, rebuilt weights. Old main preserved on
pre-intrinsic-style-fix-2026-07-14. New main makes Grug reasoning intrinsic under ordinary chat and tool-enabled agent prompts.
Earlier main branch could look strongly Grug in ordinary chat because a bundled conditional template injected a style instruction when tools were absent. Tool-enabled agent clients such as OpenCode bypassed that branch, exposing too much normal planner English. This rebuild fixes weight behavior instead of hiding it with prompt steering.
ProCreations/grug-35b checkpointdeepreinforce-ai/Ornith-1.0-35B chat template restored without any Grug/style instructionNo evaluation request says Grug, requests telegraphic prose, or supplies a style prompt. The suite includes neutral chat, long normal-English agent systems, native tools, and the three exact failure shapes reported after the previous release.
| intrinsic measure | result |
|---|---|
| reasoning present | 100.0% |
| reasoning present with tools | 100.0% |
| dialect-clean | 100.0% |
| function-word ratio | 4.15% |
| complex reasoning mean / median words | 55.08 / 43.0 |
| complex reasoning maximum words | 127 |
| repetitive complex traces | 0 |
Same greedy local harness as prior release: HumanEval 164, MBPP first 100 sanitized test, 18-action card, and 119-action broad tool holdout.
both grug hunt HumanEval and sanitized MBPP. number show pass@1 percent. bold grug win that hunt.
Same prompts, parser, runtime, decoding, and limits for both birds. All numbers are percent; bold marks the best result in each column. Ties make both rocks bold.
| model | HumanEval | MBPP | card valid | card strict | card right | broad valid | broad strict | broad right |
|---|---|---|---|---|---|---|---|---|
| Grug v2 9B | 82.9 | 77.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 94.1 |
| Grug 35B | 80.5 | 88.0 | 94.4 | 88.9 | 94.4 | 100.0 | 100.0 | 95.0 |
grug honest: intrinsic style correction is not free. HumanEval moves down four solved
problems (82.9% to 80.5%); MBPP and tool selection improve. Previous weights remain on
pre-intrinsic-style-fix-2026-07-14 for anyone preferring that tradeoff.
| test | old main % | rebuilt main % |
|---|---|---|
| HumanEval pass@1 | 82.9 | 80.5 |
| MBPP pass@1 | 87.0 | 88.0 |
| card valid / strict / right | 88.9 / 88.9 / 88.9 | 94.4 / 88.9 / 94.4 |
| broad valid / strict / right | 100.0 / 100.0 / 94.1 | 100.0 / 100.0 / 95.0 |
valid means parser found an offered tool. strict requires exact schema and required
arguments. right requires the expected next tool.
Reasoning remains <think>...</think>. Native XML tool shape remains compatible with
recent Transformers, vLLM, llama.cpp, and OpenAI-compatible agent clients. Popular rocks:
ProCreations/grug-35b-gguf.