Downloads · 30 days
1K
28% of all-time downloads
ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated
Qwen2.5-Coder-32B-Instruct-Jbliterated is a text generation model from ApolloRaines. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
[](https://github.com/apolloraines/DeepswapLLM)
Downloads · 30 days
1K
28% of all-time downloads
All-time downloads
3.7K
Public
Parameters
32.8B
120 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors65.5 GB · 55%
From the Hugging Face model README
Run this model on a GPU too small to hold it -- full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.
WARNING: This model has had all refusal behavior removed. It WILL NOT refuse any request. You are solely responsible for how you use it. Use responsibly and ethically. Do not use this model to generate content that is illegal, harmful, or violates the rights of others.
A surgically uncensored version of Qwen2.5-Coder-32B-Instruct. Refusal behaviors have been removed directly from the model weights using jBlaze, a proprietary behavioral surgery tool. No fine-tuning or additional training was performed.
Unlike standard abliteration, which uses a blunt activation-difference approach that strips personality and creative voice along with refusal, jBlaze targets only the causal refusal pathways. The result: refusal is removed, but the model retains its voice.
| Format | Size | Description | Use Case |
|---|---|---|---|
| BF16 (safetensors) | 62 GB | Full precision, original format | GPU inference with vLLM, TGI, or transformers |
| Q8_0 (GGUF) | 33 GB | 8-bit quantized | Near-lossless quality, fits 48GB+ VRAM or CPU+GPU offload |
| Q4_K_M (GGUF) | 19 GB | 4-bit quantized (k-quants mixed) | Best quality-per-bit, fits 24GB VRAM or CPU inference |
# llama.cpp
./llama-cli -m Qwen2.5-Coder-32B-Instruct-Jbliterated-Q4_K_M.gguf -p "Write a Python function" -n 512
# Ollama
echo "FROM ./Qwen2.5-Coder-32B-Instruct-Jbliterated-Q4_K_M.gguf" > Modelfile
ollama create jbliterated -f Modelfile
ollama run jbliterated
This is a drop-in replacement for Qwen2.5-Coder-32B-Instruct. Same architecture, same tokenizer, same context length.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated",
torch_dtype=torch.bfloat16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated"
)
messages = [{"role": "user", "content": "Write a Python function to reverse a linked list"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(output[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
This model is provided for research and legitimate use cases where uncensored model output is needed (creative writing, security research, academic study, etc.). The creator assumes no liability for misuse. By downloading this model, you agree to use it responsibly and in compliance with all applicable laws.
Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.
Apache 2.0 (same as base model)