Downloads · 30 days
18
39% of all-time downloads
Yakk99/scihigh2026-subtask1-bart-titled
scihigh2026-subtask1-bart-titled is a summarization model from Yakk99. Use it when you need a shorter version of a longer text. The card lists the license as apache-2.0.
Downloads · 30 days
18
39% of all-time downloads
All-time downloads
46
Public
Parameters
406M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
Full fine-tune of bart-large-cnn.
Fine-tuned for the FIRE 2026 SciHigh shared task, Subtask 1: generating research highlights from scientific paper abstracts (MixSub corpus).
<paper title> | <abstract> (corpus punctuation-stripped
style; abstracts repaired via DOI-verified Semantic Scholar recovery)| file | purpose |
|---|---|
train_final.py | standalone training recipe that produced this model |
train_final.ipynb | notebook version of the same recipe |
final/ | full data pipeline: abstract recovery, title fetch, dataset build, submission validation |
~40% of corpus abstracts are truncated mid-sentence. Each row was matched to
its ScienceDirect PII (100% verified join), resolved to a DOI via Elsevier's
keyless API, and its complete abstract recovered from Semantic Scholar under
a label-free >=90%-token-overlap acceptance filter (~95% of truncated rows
repaired). Paper titles were fetched for 100% of rows and prepended to
inputs. See final/README.md for the full pipeline and negative-results
summary. Trained 2026-08-08.