Downloads · 30 days
12
21% of all-time downloads
kaanino/tiny_dpo
tiny_dpo is a text generation model from kaanino. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
TinyLlama-1.1B fine-tuned using DPO for QA.
Downloads · 30 days
12
21% of all-time downloads
All-time downloads
58
Public
Parameters
1.1B
13.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README
TinyLlama-1.1B fine-tuned using DPO for QA.
This modelcard aims to be a base template for new models. It has been generated using this raw template.
TinyLlama-1.1B fine-tuned using Direct Preference Optimization (DPO) for Question Answering (QA) tasks, specifically, stem courses QA. The model leverages quantization and parameter-efficient fine-tuning (PEFT) techniques to optimize performance and efficiency.
This model can be used directly for question answering tasks without additional fine-tuning.
The model can be fine-tuned further for specific QA datasets or integrated into larger systems for enhanced performance in question answering applications.
The model is not suitable for tasks outside of question answering, such as generating creative content, providing medical or legal advice, or any use case requiring high levels of accuracy and reliability without proper validation.
The model may exhibit biases present in the training data and could potentially generate harmful content. Users should exercise caution and consider these limitations when deploying the model.
Users (both direct and downstream) should be made aware of the risks, biases, and limitations of the model. Continuous monitoring and evaluation are recommended to mitigate potential negative impacts.
Use the code below to get started with the model.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "kaanino/tiny_dpo"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
# Example usage
input_text = "What is the capital of France?"
inputs = tokenizer(input_text, return_tensors="pt")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0]))
We mainly used three sources of data :
Direct Preference Optimization
[More Information Needed]