Downloads · 30 days
59
7% of all-time downloads
SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked
Qwen3.8-27B-SwissNeuron-Derisked is a image-text-to-text model from SwissNeuron. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
Qwen3.8-27B-SwissNeuron-Derisked, known as SwissNeuron, is a BF16, 27B-parameter Qwen3.8 derivative engineered in Switzerland for direct technical work, strong reasoning, and capability retention. It combines focused…
Downloads · 30 days
59
7% of all-time downloads
All-time downloads
804
Public
Parameters
27.4B
108 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors55.6 GB · 100%
From the Hugging Face model README
Qwen3.8-27B-SwissNeuron-Derisked, known as SwissNeuron, is a BF16, 27B-parameter Qwen3.8 derivative engineered in Switzerland for direct technical work, strong reasoning, and capability retention. It combines focused in-house post-training with a conservative internal geometric derisk procedure, then extends the model configuration to a 1,048,576-token context window using factor-4 YaRN/RoPE scaling.
SwissNeuron is intended to provide Swiss-quality model engineering: precise provenance, conservative weight surgery, reproducible artifacts, and transparent limitations.
This repaired release supports both native thinking and non-thinking operation. Use native thinking for reasoning-intensive work:
enable_thinking=True
The current checkpoint was retrained from the official Qwen3.8-27B base using a repaired direct-answer corpus. Assistant judge/rethink analysis and the rigid EXPECTED VERIFICATION / ORIGINAL ANSWER envelope were excluded from the model-facing training targets. The merged repaired SFT passed direct regression probes for thinking-on greeting, thinking-off greeting, thinking-on coding, and an explicit anti-verification system prompt without producing the old judge/verifier format.
For lower latency, callers may still set enable_thinking=False. Agent harnesses should select the desired mode through the included Qwen chat template rather than manually inserting or stripping reasoning markers.
Many aggressively modified or “uncensored” checkpoints trade away reasoning quality, instruction fidelity, or language-model calibration. SwissNeuron was built around the opposite objective: alter behavior while minimizing movement outside the targeted representation subspace.
The model was not produced by merging an existing Heretic-style, abliteration, or third-party surgically modified checkpoint. The starting point was the official Qwen3.8-27B BF16 model. SwissNeuron was then trained on a focused internal direct-answer corpus and derisked with an in-house pipeline using a direction bank recaptured from the final trained model itself. The released candidate uses a low α=0.1, one-pass edit, skips the first two layers, preserves global weight norms, and leaves the model architecture intact.
On an internal frozen holdout, the trained model improved token-level likelihood and accuracy relative to the official starting checkpoint before the final low-strength geometric pass. This release was selected for conservative weight movement and capability preservation rather than maximum behavioral alteration.
The original Qwen3.8 context window is 262,144 tokens. SwissNeuron sets:
{
"max_position_embeddings": 1048576,
"rope_type": "yarn",
"factor": 4.0,
"original_max_position_embeddings": 262144
}
This is a configuration-level YaRN extension. The model has not yet undergone a dedicated 1M-token long-context adaptation stage, so quality near the extreme end of the window can depend on the serving stack, attention backend, prompt structure, and workload. Validate retrieval and generation quality for your own 1M-context use case. Servers must support Qwen3.8/Qwen3.5 hybrid linear attention and the included YaRN parameters.
| Property | Value |
|---|---|
| Parameters | 27,356,728,560 |
| Weight dtype | BF16 |
| Hidden size | 5,120 |
| Language layers | 64 |
| Full-attention interval | 4 |
| Full-attention layers | 16 |
| Gated-DeltaNet layers | 48 |
| FFN intermediate size | 17,408 |
| Vocabulary | 248,320 |
| Original context | 262,144 |
| Configured context | 1,048,576 |
| MTP layers | 1 |
Use a recent Transformers release with Qwen3.8/Qwen3.5 hybrid support:
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
model_id = "SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Analyze this problem carefully."}]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=2048)
print(tokenizer.decode(output[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
For production inference, use a serving engine that explicitly supports Qwen3_5ForConditionalGeneration, Gated-DeltaNet state, multimodal RoPE, and YaRN at the requested context length. A 1M context allocation requires substantial KV/state memory even with the hybrid architecture.
Thinking and coding workloads:
temperature=0.6
top_p=0.95
top_k=20
For low-latency direct answers, set enable_thinking=False. For reasoning-intensive chat, coding, and agent workloads, set enable_thinking=True. Preserve the included chat template rather than manually inserting or deleting reasoning markers.
The exact direction bank and build provenance are included for internal reproducibility.
Full collection: Qwen3.8-27B Base Derisked — Capability Preserved
SwissNeuron is the base release in a broader model series. Planned follow-ups include:
The larger fine-tuned release will be published as a separate checkpoint rather than silently replacing this BF16 base.
Built by SwissNeuron in Switzerland. Quantized variants and the larger distilled fine-tune are planned as separate releases.