Downloads · 30 days
0
dadni/GPT2-Medium-2Bit-Demo
GPT2-Medium-2Bit-Demo is a machine learning model from dadni. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains a demonstration of a highly advanced, proprietary 2-bit quantization algorithm applied to the gpt2-medium model.
Downloads · 30 days
0
Access
Public
Updated Jun 26, 2026
Repo size
1.4 GB
Likes
0
Public
Click a slice to open those files.
.pt1.4 GB · 100%
From the Hugging Face model README
This repository contains a demonstration of a highly advanced, proprietary 2-bit quantization algorithm applied to the gpt2-medium model.
Despite being aggressively compressed to exactly 2 bits per weight, the model preserves strong grammatical structure, contextual awareness, and high-dimensional logic. This serves as a proof-of-concept for next-generation extreme compression technologies.
Note: The exact quantization method, clustering techniques, and matrix transformation parameters used to achieve this result are kept strictly confidential.
You can load and test this model using standard transformers and torch libraries. No custom or proprietary layers are required to perform inference.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from huggingface_hub import hf_hub_download
# 1. Download the quantized weights
weights_path = hf_hub_download(repo_id="dadni/GPT2-Medium-2Bit-Demo", filename="gpt2-2bit-matmulfree.pt")
# 2. Load the base architecture and tokenizer
model_id = "gpt2-medium"
model = AutoModelForCausalLM.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# 3. Load the 2-bit weights into the model
state_dict = torch.load(weights_path, map_location="cpu")
model.load_state_dict(state_dict)
model.eval()
# 4. Generate Text
prompt = "The true secret to human intelligence is"
input_ids = tokenizer.encode(prompt, return_tensors="pt")
with torch.no_grad():
output_ids = model.generate(
input_ids,
max_new_tokens=100,
do_sample=True,
top_p=0.9,
temperature=0.8,
repetition_penalty=1.3,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
For inquiries or demonstrations on larger models, please refer to the author.