Downloads · 30 days
9
47% of all-time downloads
DeathGodlike/Retreatcost_Evertide-RX-12B_EXL3
Retreatcost_Evertide-RX-12B_EXL3 is a text generation model from DeathGodlike. Use it when you need the model to write or continue text. It is set up for safetensors.
Downloads · 30 days
9
47% of all-time downloads
All-time downloads
19
Public
Repo size
30.6 GB
Likes
0
Public
Click a slice to open those files.
.md7.9 KB · 84%
From the Hugging Face model README
Evertide-RX-12B by Retreatcost
| Type | Size | CLI |
|---|---|---|
| H8-4.0BPW | 7.49 GB | Copy-paste the lines / Download the batch file |
| H8-6.0BPW | 10.22 GB | Copy-paste the lines / Download the batch file |
| H8-8.0BPW | 12.94 GB | Copy-paste the lines / Download the batch file |
Requirements: A python installation with huggingface-hub module to use CLI.
License detected: apache-2.0
The license for the provided quantized models is inherited from the source model (which incorporates the license of its original base model). For definitive licensing information, please refer first to the page of the source or base models. File and page backups of the source model are provided below.
Date: 15.05.2026
<details> <summary>Source page (click to expand)</summary>
A generalist model, with some reasoning capabilities and multi-lang support.
Supported languages:
This model is trained in FFT based on unreleased cowriter model merge (uses same models as Retreatcost/KansenSakura-Erosion-RP-12b, credits to all original model authors.), using in-progress dateset, that I am creating for another project.
Training stats can be found in "Training metrics" tab.
Reasoning should work out of the box most of the times with occasional replies without it. For absolute consistency you can prefill model responses with "< think >\n" (think tag without spaces, line break is preferred).
I haven't really tested or trained the model for long context, so it will probably break earlier than regular models. You can set a higher context, for example 16K, 24K or 32K, but I don't guarantee how it will behave.
I trained 2 variants of the model:
Unrolled turns teach local attention much better and train faster, but generalize worse for multi-turn (Evertide-LA-12B, Local attention). Regular turns have much better multi-turn generalisation, but they tend to memorize instead of training new capabilities. (Evertide-GA-12B, Global attention).
I also trained these with changed RoPE theta - 10K for GA, 10M for LA. My reasoning behind this is that during merging I "unrotate" the changes in config, effectively creating a distribution that I haven't trained in.
LA becomes shrinked to be even more specialized in short context, while GA gets stretched to cover longer context.
Then I merged these training runs using passthrough in a pattern 4:1, similar to how Gemma 4 models have layered SWA and GA.

The following YAML configuration was used to produce this model:
merge_method: passthrough
slices:
- sources:
- model: Evertide-LA-12B
layer_range: [0, 4]
- sources:
- model: Evertide-GA-12B
layer_range: [4, 5]
- sources:
- model: Evertide-LA-12B
layer_range: [5, 9]
- sources:
- model: Evertide-GA-12B
layer_range: [9, 10]
- sources:
- model: Evertide-LA-12B
layer_range: [10, 14]
- sources:
- model: Evertide-GA-12B
layer_range: [14, 15]
- sources:
- model: Evertide-LA-12B
layer_range: [15, 19]
- sources:
- model: Evertide-GA-12B
layer_range: [19, 20]
- sources:
- model: Evertide-LA-12B
layer_range: [20, 24]
- sources:
- model: Evertide-GA-12B
layer_range: [24, 25]
- sources:
- model: Evertide-LA-12B
layer_range: [25, 29]
- sources:
- model: Evertide-GA-12B
layer_range: [29, 30]
- sources:
- model: Evertide-LA-12B
layer_range: [30, 34]
- sources:
- model: Evertide-GA-12B
layer_range: [34, 35]
- sources:
- model: Evertide-LA-12B
layer_range: [35, 39]
- sources:
- model: Evertide-GA-12B
layer_range: [39, 40]
dtype: bfloat16
</details>
Probably not.
Not exactly. With some prompting it is definitely capable to output something, but it's not designed to be an ERP model in the first place. I would rate it 4/10 in this department, it's by design.
The same as above, it will absolutely refuse some of your more unhinged prompts. You can try to abliterate it, tho.
For this model achieving ERP capabilities wasn't the goal, so I'm happy with current state.
Soon™.
No, not yet, but that's one of future plans.
It's hard to tell exactly, it definitely has some elements of it, but it also was trainded with some specific constraints, that force causality between thinking blocks and answer. So I would say that it's at least a hybrid. Any further improvements require RL training.
Only 451 sample, but they are all manually crafted and refined using score-samples script.
</details>