Downloads · 30 days
15
37% of all-time downloads
FINAL-Bench/Darwin-Chimera-4B-Gen1
Darwin-Chimera-4B-Gen1 is a text generation model from FINAL-Bench. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
⚠️ Generation-1 backbone — a research checkpoint, not a product. Private repo. This is a Qwen3-4B derivative, not a from-scratch model. We state this explicitly.
Downloads · 30 days
15
37% of all-time downloads
All-time downloads
41
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
⚠️ Generation-1 backbone — a research checkpoint, not a product. Private repo. This is a Qwen3-4B derivative, not a from-scratch model. We state this explicitly.
Darwin-Chimera-4B-Gen1 is the first-generation adapter backbone of the Darwin-Chimera
line. We take Qwen/Qwen3-4B and re-wire only its attention via VIDRAFT
attention-healing, while freezing the FFN, embeddings, and lm_head, and convert the
attention to a sliding-window configuration.
The purpose is to verify that a VIDRAFT-healed attention circuit can sit on a frozen knowledge core — the foundation for Generation-2 (FFN cross-breeding with other models).
Measured relative change ||A−B|| / ||A|| against the original Qwen3-4B:
| Component | Relative change | Note |
|---|---|---|
| FFN (mlp) | 0.000% | frozen — identical to Qwen3-4B |
| embed / lm_head | 0.000% | frozen — identical |
| attention (self_attn) | 3.0% mean (7.5% max) | healed |
| layernorm | 0.04% | minimal |
| config (hidden/inter/layers/vocab) | identical | only sliding_window=4096 added |
→ At the weight level this checkpoint is clearly a Qwen3-4B derivative. We make no claim of independence or from-scratch training. Knowledge/FFN is 100% Qwen3-4B.
Apache 2.0, inherited from Qwen/Qwen3-4B. Built on Qwen/Qwen3-4B.