Downloads · 30 days
0
Yeop9690/custom_summarization_dataset
custom_summarization_dataset is a summarization model from Yeop9690. Use it when you need a shorter version of a longer text. The card lists the license as cc-by-4.0.
- Custom Text Summarization Dataset (CNN/DailyMail Subset)
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2024
Repo size
348 KB
Likes
0
Public
Click a slice to open those files.
.arrow348 KB · 99%
From the Hugging Face model README
This dataset contains a subset of the CNN/DailyMail news dataset, which is used for training text summarization models. The dataset consists of articles paired with human-generated summaries. It is widely used in the development of natural language processing models for summarization tasks.
article: The news article texthighlights: The human-generated summary of the articleThe dataset was collected by scraping news articles from CNN and DailyMail websites. The articles were paired with manually written summaries to form training examples. This dataset was originally prepared for the task of abstractive text summarization.
from datasets import load_dataset
dataset = load_dataset("cnn_dailymail", "3.0.0", split="train[:1%]")