Downloads · 30 days
18
3% of all-time downloads
Fu01978/SmollerLM2-360M-Instruct-Pruned
SmollerLM2-360M-Instruct-Pruned is a text generation model from Fu01978. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A structurally pruned version of the SmolLM2-360M-Instruct model.
Downloads · 30 days
18
3% of all-time downloads
All-time downloads
544
Public
Parameters
350M
700 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors700 MB · 100%
From the Hugging Face model README
A structurally pruned version of the SmolLM2-360M-Instruct model.
This model was created as an experiment in pruning.
The model underwent Structured Every-Nth Neuron Pruning. Unlike random dropout or unstructured pruning, this method maintains the dense matrix format required by standard hardware accelerators.
While the original model is distributed in FP32, this model provides an optimization that makes it significantly more accessible:
Total Savings: 51.6% smaller than the original version.
TECHNICAL NOTE: On CPUs without native FP16 support, this model may experience a 'tax' resulting in slower tokens-per-second than the original. This model is RAM-optimized, not necessarily CPU-Latency optimized in its raw FP16 state.
To get the best performance, it is recommended to use it on a GPU or via 4-bit/8-bit quantization to bypass CPU floating-point limitations.
You will first need to install
the bitsandbytes library
in Python.
(pip install bitsandbytes)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "Fu01978/SmollerLM2-360M-Instruct-Pruned"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
quantization_config={"load_in_8bit": True},
device_map="auto"
)
If running on lower-end CPUs, load the model in 4-bit to ensure the weights fit in the L1/L2 cache:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "Fu01978/SmollerLM2-360M-Instruct-Pruned"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config={"load_in_4bit": True},
device_map="auto"
)
Use the following snippet to chat with the model. This uses the chat template.
# Define your message(s)
messages = [
{"role": "user", "content": "Explain the concept of gravity."}
]
input_text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=150,
temperature=0.2,
do_sample=True,
repetition_penalty=1.1
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
Gravity is indeed one of the most fundamental concepts in physics and mathematics. It's essentially the "attraction" between two bodies or masses. According to Einstein's theory of general relativity, mass warps space-time around it, creating a gravitational field that attracts other objects with mass. This means that anything having mass has a gravitational pull on other matter, making them feel heavy. For example, when you drop an object, you're not really feeling its weight; rather, you're feeling the gravitational force exerted by the Earth. The actual weight of an object depends on how massive the object itself is, which can be calculated using formulas like F=G x m/r^2 where G is the gravitational constant, [hit 150 token limit]
As a pruned version of SmolLM2, this model inherits the biases of its parent.
While the pruning was found to be stable, users may encounter slight regressions in mathematical reasoning compared to the full model.