Downloads · 30 days
1
8% of all-time downloads
peter520416/t5-small-custom
t5-small-custom is a machine learning model from peter520416. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
1
8% of all-time downloads
All-time downloads
12
Public
Parameters
60.5M
243 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors242 MB · 99%
From the Hugging Face model README
This model is based on the t5-small architecture, which is a small version of the T5 (Text-To-Text Transfer Transformer) model.
In this case, the model has been fine-tuned for summarization tasks,
specifically to generate summaries for long articles such as those from the CNN/DailyMail dataset.
The model was fine-tuned using the CNN/DailyMail dataset (version 3.0.0). This dataset contains news articles with associated highlights that serve as summaries.
The dataset consists of:
The dataset contains long articles as inputs and corresponding highlights as target summaries.
To use this model for summarization, load the model and tokenizer using the transformers library from Hugging Face. Here’s an example in Python:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("path/to/saved/model")
model = AutoModelForSeq2SeqLM.from_pretrained("path/to/saved/model")
# Summarization example
summarizer = pipeline("summarization", model=model, tokenizer=tokenizer)
summary = summarizer("Input text here", max_length=150, min_length=30, num_beams=4, do_sample=False)
print(summary)
The model was evaluated using both ROUGE and BLEU metrics. Here are the scores on the validation dataset:
ROUGE-1: 43.25 ROUGE-2: 20.85 ROUGE-L: 39.18 BLEU-1: 31.5 BLEU-2: 21.7 BLEU-4: 10.3 These scores indicate that the model performs well at summarizing long texts while maintaining key information.
The model is trained on news articles and may not generalize well to other domains such as legal, scientific, or conversational text. Since the model uses beam search, the generated summaries might occasionally be repetitive or overly verbose.
Bias in data: The model was trained on news articles, which may contain inherent biases based on the sources of the data. Hallucination: As with most generative models, the T5 model may sometimes generate inaccurate or misleading information in the summary, especially if the input text contains ambiguous or conflicting information. Data privacy: This model was trained on publicly available datasets, and no private data was used in the fine-tuning process.