Downloads · 30 days
306
14% of all-time downloads
Felladrin/Qwen2-96M
Qwen2-96M is a text generation model from Felladrin. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Qwen2-96M is a small language model based on the Qwen2 architecture, trained from scratch on English datasets with a context length of 8192 tokens. With only 96 million parameters, this model serves as a lightweight b…
Downloads · 30 days
306
14% of all-time downloads
All-time downloads
2.2K
Public
Parameters
96.2M
204 MB on disk
Likes
4
Trending 1
Click a slice to open those files.
.safetensors192 MB · 94%
From the Hugging Face model README
Qwen2-96M is a small language model based on the Qwen2 architecture, trained from scratch on English datasets with a context length of 8192 tokens. With only 96 million parameters, this model serves as a lightweight base model that can be fine-tuned for specific tasks.
Due to its compact size, the model has significant limitations in reasoning, factual knowledge, and general capabilities compared to larger models. It may produce incorrect, irrelevant, or nonsensical outputs. Additionally, as it was trained on internet text data, it may contain biases and potentially generate inappropriate content.
pip install transformers==4.49.0 torch==2.6.0
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
import torch
model_path = "Felladrin/Qwen2-96M"
prompt = "I've been thinking about"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path).to(device)
streamer = TextStreamer(tokenizer)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
model.generate(
inputs.input_ids,
attention_mask=inputs.attention_mask,
max_length=tokenizer.model_max_length,
streamer=streamer,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
repetition_penalty=1.1,
do_sample=True,
temperature=0.7,
top_p=0.9,
top_k=0,
min_p=0.1,
)