Downloads · 30 days
17
45% of all-time downloads
igidn/SmolLM3-Chat-v1-Adapter
SmolLM3-Chat-v1-Adapter is a machine learning model from igidn. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
This repository contains the LoRA (Low-Rank Adaptation) weights for SmolLM3-Chat-v1.
Downloads · 30 days
17
45% of all-time downloads
All-time downloads
38
Public
Repo size
2.9 GB
Likes
1
Public
Click a slice to open those files.
.safetensors2.4 GB · 81%
From the Hugging Face model README
This repository contains the LoRA (Low-Rank Adaptation) weights for SmolLM3-Chat-v1.
This adapter was trained to give the SmolLM3-3B-Base model a casual, witty, and "internet-native" personality. It moves away from robotic assistant responses in favor of a more human-like vibe.
Less is more.
This model relies on a specific "vibe" learned during training. Over-prompting it with complex system instructions (e.g., "You are a helpful assistant who is polite, follows rules X, Y, Z...") will degrade the output quality.
Recommended System Prompt: (simply leave it empty for the most raw, casual experience)
This script demonstrates how to load the base model in 4-bit and attach the adapter.
import torch
from threading import Thread
from peft import PeftModel
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
BitsAndBytesConfig,
TextIteratorStreamer
)
# 1. Define IDs
ADAPTER_ID = "igidn/SmolLM3-Chat-v1-adapter"
BASE_MODEL_ID = "HuggingFaceTB/SmolLM3-3B-Base"
# 2. Quantization Config (4-bit)
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True,
)
# 3. Load Base Model
tokenizer = AutoTokenizer.from_pretrained(ADAPTER_ID) # Load tokenizer from adapter to get special tokens
model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_ID,
quantization_config=bnb_config,
device_map="auto",
trust_remote_code=True
)
# 4. Attach Adapter
model = PeftModel.from_pretrained(model, ADAPTER_ID)
# 5. Define Conversation
messages = [
{"role": "user", "content": "Haiiii"}
]
# 6. Apply Chat Template
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer([prompt], return_tensors="pt").to(model.device)
# 7. Streamer & Generation
streamer = TextIteratorStreamer(tokenizer, timeout=10.0, skip_prompt=True, skip_special_tokens=True)
# --- CRITICAL GENERATION CONFIG ---
generate_kwargs = dict(
**inputs,
streamer=streamer,
max_new_tokens=512,
do_sample=True,
# Core Vibe Parameters
temperature=0.8,
top_p=0.85,
# Stability Parameters (Prevents looping)
repetition_penalty=1.15,
no_repeat_ngram_size=3,
pad_token_id=tokenizer.eos_token_id
)
thread = Thread(target=model.generate, kwargs=generate_kwargs)
thread.start()
print("Assistant: ", end="")
for new_text in streamer:
print(new_text, end="", flush=True)
The model was trained for 2 epochs using SFTTrainer.
| Metric | Value |
|---|---|
| Final Loss | 1.41 |
| Final Token Accuracy | ~65.9% |
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj, embed_tokens, lm_head)Created with <3 by me