Downloads · 30 days
7
9% of all-time downloads
anonymous12321/Primera-Summarization-Council-PT
Primera-Summarization-Council-PT is a machine learning model from anonymous12321. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-nd-4.0.
Primera-Summarization-Council-PT is an abstractive text summarization model based on primera, fine-tuned to produce concise and informative summaries of discussion subjects from Portuguese municipal meeting minutes. T…
Downloads · 30 days
7
9% of all-time downloads
All-time downloads
79
Public
Parameters
447M
1.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.8 GB · 100%
From the Hugging Face model README
Primera-Summarization-Council-PT is an abstractive text summarization model based on primera, fine-tuned to produce concise and informative summaries of discussion subjects from Portuguese municipal meeting minutes.
The model was trained on a curated and annotated corpus of official municipal meeting minutes covering a variety of administrative and political topics at the municipal level.
Try out the model: Hugging Face Space Demo
allenai/primera using supervised fine-tuning.The model receives a discussion subject of a municipal meeting and outputs a short, coherent summary highlighting:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_name = "anonymous12321/Primera-Summarization-Council-PT"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
text = """
17. PROCESSO DE OBRAS N.º ***** -- EDIFIC\nPelo Senhor Presidente foi presente a esta reunião a informação n.º ****** da Secção de Urbanismo e Fiscalização -- Serviço de Obras Particulares que se anexa à presente ata. \nPonderado e analisado o assunto o Executivo Municipal deliberou por unanimidade aprovar as especialidades relativas ao processo de obras n.º ***** -- EDIFIC.
"""
inputs = tokenizer(text, return_tensors="pt", max_length=1024, truncation=True)
summary_ids = model.generate(**inputs, max_length=128, num_beams=4, early_stopping=True)
print(tokenizer.decode(summary_ids[0], skip_special_tokens=True))
Output:
"O Executivo Municipal aprovou, por unanimidade, as especialidades relativas a um processo de obras particulares."
| Metric | Score | Description |
|---|---|---|
| ROUGE-1 | 0.632 | Unigram overlap between generated and reference summaries |
| ROUGE-2 | 0.500 | Bigram overlap |
| ROUGE-L | 0.577 | Longest common subsequence overlap |
| BERTScore (F1) | 0.846 | Semantic similarity between summary and reference |
allenai/primeraeval_steps=100)max_length=512 and stride=256 for hierarchical input segmentationThe model was trained on a specialized dataset of Portuguese municipal meeting minutes, consisting of:
Dataset sources include:
This model is released under the
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License (CC BY-NC-ND 4.0).