Downloads · 30 days
690
100% of all-time downloads
brucoder/winter-frost-2
winter-frost-2 is a text generation model from brucoder. Use it when you need the model to write or continue text. It is set up for transformers.
A fine-tune of Qwen/Qwen2.5-7B-Instruct, created by INEZA AIME BRUNO (brucoder). This is a standalone, merged checkpoint — the LoRA adapter has already been folded into the base model's weights, so it loads and runs l…
Downloads · 30 days
690
100% of all-time downloads
All-time downloads
690
Public
Parameters
7.6B
15.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
A fine-tune of Qwen/Qwen2.5-7B-Instruct, created by INEZA AIME BRUNO (brucoder). This is a standalone, merged checkpoint — the LoRA adapter has already been folded into the base model's weights, so it loads and runs like any other model, no adapter-loading step required.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained("brucoder/winter-frost-2")
model = AutoModelForCausalLM.from_pretrained(
"brucoder/winter-frost-2",
torch_dtype=torch.float16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who made you?"},
]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_tensors="pt", return_dict=True
).to(model.device)
output = model.generate(**inputs, max_new_tokens=200, do_sample=True, temperature=0.7, top_p=0.9, pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Paste into one cell:
!pip install -q -U transformers accelerate bitsandbytes
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True)
tokenizer = AutoTokenizer.from_pretrained("brucoder/winter-frost-2")
model = AutoModelForCausalLM.from_pretrained("brucoder/winter-frost-2",
quantization_config=bnb_config, device_map="auto")
conversation = [{"role": "system", "content": "You are a helpful assistant."}]
print("Model loaded. Chat with winter-frost-2 (type 'quit' to stop).\n")
while True:
user_input = input("You: ")
if user_input.strip().lower() in ("quit", "exit"):
break
conversation.append({"role": "user", "content": user_input})
inputs = tokenizer.apply_chat_template(conversation, tokenize=True, add_generation_prompt=True,
return_tensors="pt", return_dict=True).to(model.device)
output = model.generate(**inputs, max_new_tokens=300, do_sample=True, temperature=0.7,
top_p=0.9, pad_token_id=tokenizer.eos_token_id)
response = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(f"winter-frost-2: {response}\n")
conversation.append({"role": "assistant", "content": response})
Requires Python 3.9+. A GPU with 16GB+ VRAM keeps this fast; it will also run on CPU, just slowly. First run downloads ~15GB of weights (cached after that):
pip install -q -U transformers accelerate torch
python3 -c "
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained('brucoder/winter-frost-2')
model = AutoModelForCausalLM.from_pretrained('brucoder/winter-frost-2', torch_dtype=torch.float16, device_map='auto')
conversation = [{'role': 'system', 'content': 'You are a helpful assistant.'}]
print('Model loaded. Chat with winter-frost-2 (type quit to stop).')
while True:
user_input = input('You: ')
if user_input.strip().lower() in ('quit', 'exit'):
break
conversation.append({'role': 'user', 'content': user_input})
inputs = tokenizer.apply_chat_template(conversation, tokenize=True, add_generation_prompt=True, return_tensors='pt', return_dict=True).to(model.device)
output = model.generate(**inputs, max_new_tokens=300, do_sample=True, temperature=0.7, top_p=0.9, pad_token_id=tokenizer.eos_token_id)
response = tokenizer.decode(output[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print('winter-frost-2:', response)
conversation.append({'role': 'assistant', 'content': response})
"
Not a new architecture and not trained from scratch. With ~690 training examples, this reliably shifts small, well-defined behaviors (like stating who created it) but should not be expected to meaningfully change general reasoning, coding ability, or knowledge compared to the base model.
If you'd rather load the LoRA adapter separately on top of the base model instead of this merged version, it's available at brucoder/winter-frost-2-adapter.
INEZA AIME BRUNO (brucoder) Instagram: iabru.ceo