Downloads · 30 days
18
20% of all-time downloads
mehta/CooperLM-354M-4bit
CooperLM-354M-4bit is a text generation model from mehta. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This is a 4-bit quantized version of CooperLM-354M, a 354M parameter GPT-2 style language model trained from scratch on a subset of Wikipedia, BookCorpus, and OpenWebText.
Downloads · 30 days
18
20% of all-time downloads
All-time downloads
88
Public
Parameters
364M
260 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors260 MB · 98%
How the weights are stored.
U8311M · 86%
From the Hugging Face model README
This is a 4-bit quantized version of CooperLM-354M, a 354M parameter GPT-2 style language model trained from scratch on a subset of Wikipedia, BookCorpus, and OpenWebText.
The quantized model is intended for faster inference and smaller memory footprint, especially useful for CPU or limited-GPU setups.
AutoGPTQ (safetensors)from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("mehta/CooperLM-354M-4bit")
model = AutoModelForCausalLM.from_pretrained("mehta/CooperLM-354M-4bit")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
prompt = "In the distant future,"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
outputs = model.generate(
**inputs,
max_length=100,
temperature=0.8,
top_p=0.95,
do_sample=True
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))