Downloads · 30 days
49
14% of all-time downloads
Lasimeri/GrugCap-27B
GrugCap-27B is a text generation model from Lasimeri. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A layer splice, not a finetune. Decoder blocks 10–44 of ProCreations/grug-27b were copied byte-for-byte into bottlecapai/ThinkingCap-Qwen3.6-27B, replacing that range wholesale. Everything outside the range is Thinkin…
Downloads · 30 days
49
14% of all-time downloads
All-time downloads
342
Public
Parameters
27.8B
56.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors56.4 GB · 100%
From the Hugging Face model README
A layer splice, not a finetune. Decoder blocks 10–44 of ProCreations/grug-27b were copied byte-for-byte into bottlecapai/ThinkingCap-Qwen3.6-27B, replacing that range wholesale. Everything outside the range is ThinkingCap's.
Both parents are Apache-2.0 finetunes of Qwen/Qwen3.6-27B, which is what makes a raw weight-space splice coherent: their residual streams share a basis, so transplanted layers land in a space the surrounding layers still understand. Splicing across unrelated pretraining runs does not work this way and produces noise.
The premise being tested is that the middle depth band of a transformer carries most of the feature-composition work, so swapping it swaps "how the model thinks" while the outer layers keep the host's tokenization behaviour, output formatting and answer style. Blocks 10–44 are 15–70% of depth, a heuristic choice, not a discovered boundary.
| Component | Source |
|---|---|
embed_tokens, lm_head, final norm | ThinkingCap (identical in both parents) |
| Decoder blocks 0–9 | ThinkingCap |
| Decoder blocks 10–44 | grug-27b |
| Decoder blocks 45–63 | ThinkingCap |
| Vision tower, MTP head | ThinkingCap |
| Tokenizer, chat template, config | ThinkingCap |
463 of 1199 tensors come from grug: 24.8 GiB of 51.7 GiB, or 47.9% by weight. The 35 spliced blocks are 26 linear-attention and 9 full-attention layers, the hybrid pattern placing full attention every fourth layer.
Provenance was verified after the build by hashing tensors in the output against
the same tensors fetched from both parent repos. The boundary is exact:
blk.44.mlp.down_proj matches grug and not ThinkingCap, blk.45.mlp.down_proj
matches ThinkingCap and not grug, with no unexplained tensors.
A useful incidental finding: the two parents have byte-identical
embed_tokens, lm_head, every layernorm and every conv1d weight. They differ
only in projection matrices, in all 64 blocks. That is the expected signature of
grug's LoRA-on-linears method, and it means a provenance check that samples only
layernorms cannot distinguish the parents at all.
On a Q8_0 GGUF under llama.cpp on 2×RTX 3090:
Thinking traces are terse and enumerative, closer to grug's register than to a verbose reasoner. Mean trace length across the battery was 155 characters. On the 24-game prompt it enumerated ten dead ends and verified the winner in 438 characters.
Be skeptical of this model until someone runs real evals. In particular:
BF16 safetensors in 11 shards, plus model-base-aux.safetensors holding
ThinkingCap's MTP draft head. That aux file is not in the weight index; move it
out of the directory before running convert_hf_to_gguf.py or the converter will
trip over it.
python splice_remote.py \
--recipient bottlecapai/ThinkingCap-Qwen3.6-27B \
--donor ProCreations/grug-27b \
--start 10 --end 45 \
--out ./GrugCap-27B
The splice is IO-bound, so re-cutting at different boundaries is cheap relative
to any training. If you want a different band, change --start/--end; the
tooling streams tensors by HTTP range request and never stores either parent
locally.
Neither parent's authors were involved in or endorse this merge.