Downloads · 30 days
11
18% of all-time downloads
Trelis/TrelisSmolLM-instruct
TrelisSmolLM-instruct is a text generation model from Trelis. Use it when you need the model to write or continue text. It is set up for transformers.
This model is a fine-tuned version of TrelisSmolLM-base, optimized for instruction following and conversational tasks using the WebInstructSub dataset.
Downloads · 30 days
11
18% of all-time downloads
All-time downloads
62
Public
Parameters
107M
214 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors214 MB · 98%
From the Hugging Face model README
This model is a fine-tuned version of TrelisSmolLM-base, optimized for instruction following and conversational tasks using the WebInstructSub dataset.
To purchase the training scripts used for this model, visit: https://trelis.com/advanced-fine-tuning-scripts/
TrelisLM-80M-SFT is an 80 million parameter language model derived from SmolLM-360M through pruning and distillation, and then fine-tuned on the WebInstructSub dataset for improved instruction following capabilities.
This model is designed for instruction following and conversational tasks. It can be used for:
This model should not be used for:
The model was fine-tuned on the TIGER-Lab/WebInstructSub dataset, which consists of instruction-response pairs. The training process used:
The training used a custom learning rate scheduler with an initial constant phase followed by cosine annealing.
Evaluation was performed on a randomly selected subset of 10,000 rows from the WebInstructSub dataset.
[More Information Needed]
As this model is fine-tuned on the WebInstructSub dataset, it may inherit biases present in that dataset. Additionally, as a smaller language model, it may have limitations in handling complex or highly specialized tasks compared to larger models.
You can use this model with the Transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Trelis/80M-2percent-corpus-SFT")
tokenizer = AutoTokenizer.from_pretrained("Trelis/80M-2percent-corpus-SFT")
# Example usage
input_text = "What is the capital of France?"
input_ids = tokenizer.encode(input_text, return_tensors="pt")
output = model.generate(input_ids, max_length=50)
response = tokenizer.decode(output[0], skip_special_tokens=True)
print(response)