Downloads · 30 days
0
peter520416/custom_summarization_dataset
custom_summarization_dataset is a machine learning model from peter520416. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2024
Repo size
348 KB
Likes
0
Public
Click a slice to open those files.
.arrow348 KB · 98%
From the Hugging Face model README
Custom Text Dataset for Summarization
This dataset contains articles and their corresponding summaries, created specifically for text summarization tasks. It is designed to train and evaluate models that can generate concise summaries from longer pieces of text. The dataset is based on publicly available news articles from various sources.
Columns:
article: Full text of the news article.summary: A concise summary or highlights of the article.Data types:
Dataset split:
The dataset was curated by scraping news articles from publicly available sources. We selected a wide range of articles to cover various domains such as politics, technology, sports, and health. After collection, articles were manually paired with summaries to ensure accuracy.
The dataset was preprocessed to clean and normalize the text data:
nltk.To use the dataset for text summarization tasks, you can load it using popular data handling libraries such as pandas or datasets.
Here's an example of loading and using the dataset for a summarization task:
import pandas as pd
# Load the dataset from CSV
df = pd.read_csv("path/to/custom_text_dataset.csv")
# Display first few rows
print(df.head())
# Example usage for a text summarization model
article = df['article'][0]
summary = df['summary'][0]
print("Article:", article)
print("Summary:", summary)
The dataset was evaluated using automatic summarization metrics such as ROUGE and BLEU. Summarization models trained on this dataset were evaluated on a separate test set.
ROUGE-1: 43.0 ROUGE-2: 21.0 BLEU-4: 10.5 These scores indicate how well the generated summaries match the human-written summaries in the dataset.
The dataset contains only English articles, so it is not applicable to non-English text summarization tasks. Due to the focus on news articles, the model may not generalize well to other domains like legal or medical text. Summaries may occasionally omit nuanced information due to manual summarization.
Data bias: Articles were scraped from specific sources, which may introduce bias depending on the perspective of the original publisher. Content accuracy: Summaries are human-written but may still contain errors or subjective interpretations of the articles. Data privacy: All data was collected from publicly available sources, and no private or sensitive data was used in the dataset creation. Use of dataset: The dataset should not be used to create misleading or false summaries that could misinform the public.