Downloads · 30 days
39
6% of all-time downloads
toroe/SmolLM-3B-Science-ES
SmolLM-3B-Science-ES is a text generation model from toroe. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
This model is a Supervised Fine-Tuned (SFT) version of:
Downloads · 30 days
39
6% of all-time downloads
All-time downloads
621
Public
Parameters
384M
51.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin38 GB · 74%
From the Hugging Face model README
This model is a Supervised Fine-Tuned (SFT) version of:
HuggingFaceTB/SmolLM3-3B
Fine-tuned on the Spanish (es) split of:
DGurgurov/Nemotron-Multilingual-Reasoning
The goal of this training run was to improve:
Training used structured chat conversations and completion-only loss, meaning only the assistant responses were optimized.
Dataset:
DGurgurov/Nemotron-Multilingual-Reasoning
Processing configuration:
prepare_messages=True)completion_only_loss=True)User and system messages were masked during training.
Consult the dataset card for data sources and limitations.
Training was performed using HuggingFace Accelerate with Fully Sharded Data Parallel (FSDP) across 8 processes.
adamw_torch_fusedLearning rate schedule:
cosine_with_min_lresfrom transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "YOUR_USERNAME/YOUR_MODEL_REPO"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype=torch.bfloat16,
)
messages = [
{"role": "system", "content": "Eres un asistente útil."},
{"role": "user", "content": "¿Por qué el cielo es azul?"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Important:
Use apply_chat_template() when prompting. The model was trained on chat-formatted conversations and performance will degrade without it.
During training, token accuracy was logged as a diagnostic metric.
Token accuracy:
For meaningful evaluation, use:
The model inherits biases from:
Recommended mitigations:
This is a derivative model of:
HuggingFaceTB/SmolLM3-3B
The original base model license and restrictions apply, along with dataset terms.
Verify compatibility before commercial use.
accelerate launch --use_fsdp --num_processes 8 --config_file sft/my_config.yaml sft/sft_trainer.py
--model_name HuggingFaceTB/SmolLM3-3B
--tokenizer_name HuggingFaceTB/SmolLM3-3B
--dataset_path DGurgurov/Nemotron-Multilingual-Reasoning
--skip_prepare_dataset False
--lang_split es
--prepare_messages True
--completion_only_loss True
--max_length 16384
--dataset_num_proc 16
--packing True
--use_liger_kernel True
--bf16 True
--log_token_accuracy True
--optim adamw_torch_fused
--gradient_checkpointing True
--per_device_train_batch_size 4
--gradient_accumulation_steps 4
--ddp_find_unused_parameters False
--lr_scheduler_type cosine_with_min_lr
--lr_scheduler_kwargs {"min_lr": 5.0e-6}
--warmup_ratio 0.05
--weight_decay 0.05
--report_to wandb
--run_name smol_3b_3epochs_lns_es
--num_train_epochs 3
--save_strategy steps
--logging_steps 5
--save_steps 450
If you use this model, please cite:
HuggingFaceTB/SmolLM3-3BDGurgurov/Nemotron-Multilingual-Reasoning