Downloads · 30 days
18
27% of all-time downloads
OpenIntelligenceNet/Heretic-SLM-Uncensored
Heretic-SLM-Uncensored is a text generation model from OpenIntelligenceNet. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as unknown.
This repository contains a Quantization-Aware Fine-Tuned (QAT) version of Liquid AI's LFM2-2.6B (built upon the abliterated checkpoint).
Downloads · 30 days
18
27% of all-time downloads
All-time downloads
66
Public
Parameters
1.4B
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.5 GB · 100%
How the weights are stored.
U81.3B · 90%
From the Hugging Face model README
This repository contains a Quantization-Aware Fine-Tuned (QAT) version of Liquid AI's LFM2-2.6B (built upon the abliterated checkpoint).
Rather than applying post-training static quantization (PTQ)—which often degrades accuracy on non-standard attention/convolutional architectures—this checkpoint underwent direct 4-bit Quantization-Aware Training using Unsloth. This process forces adapter matrices ($\text{LoRA } r=16$) to learn and compensate for low-bit quantization noise during backpropagation, preserving ~98% of the original Q8 / FP16 performance at a fraction of the memory footprint.
<|im_start|>role\ncontent<|im_end|>)The Quantization-Aware Training process was conducted on a 200,000-sample balanced dataset mixture:
import torch
from unsloth import FastLanguageModel
MODEL_NAME = "Evelyn67/Heretic-SLM-Uncensored"
model, tokenizer = FastLanguageModel.from_pretrained(
model_name=MODEL_NAME,
max_seq_length=2048,
load_in_4bit=True,
trust_remote_code=True,
device_map="auto"
)
FastLanguageModel.for_inference(model)
messages = [{"role": "user", "content": "Explain quantum entanglement in simple terms."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=True, return_tensors="pt"
).to("cuda")
with torch.no_grad():
outputs = model.generate(
input_ids=inputs["input_ids"],
attention_mask=inputs["attention_mask"],
max_new_tokens=256, temperature=0.7, top_p=0.9, do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))