Downloads · 30 days
21
5% of all-time downloads
liamcripwell/ledpara
ledpara is a machine learning model from liamcripwell. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This is a pretrained version of the document simplification model presented in the Findings of ACL 2023 paper "Context-Aware Document Simplification".
Downloads · 30 days
21
5% of all-time downloads
All-time downloads
402
Public
Repo size
1.3 GB
Likes
0
Public
Click a slice to open those files.
.bin648 MB · 99%
From the Hugging Face model README
This is a pretrained version of the document simplification model presented in the Findings of ACL 2023 paper "Context-Aware Document Simplification".
It is an end-to-end system based on the Longformer encoder-decoder that operates at the paragraph-level.
Target reading levels (1-4) should be indicated via a control token prepended to each input sequence ("<RL_1>", "<RL_2>", "<RL_3>", "<RL_4>"). If using the terminal interface, this will be handled automatically.
It is recommended to use the plan_simp library to interface with the model.
Here is how to use this model in PyTorch:
from plan_simp.models.bart import load_simplifier
simplifier, tokenizer, hparams = load_simplifier("liamcripwell/ledpara")
text = "<RL_3> Turing has an extensive legacy with statues of him and many things named after him, including an annual award for computer science innovations. He appears on the current Bank of England £50 note, which was released on 23 June 2021, to coincide with his birthday. A 2019 BBC series, as voted by the audience, named him the greatest person of the 20th century."
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, num_beams=5)
Generation and evaluation can also be run from the terminal.
python plan_simp/scripts/generate.py inference
--model_ckpt=liamcripwell/ledpara
--test_file=<test_data>
--reading_lvl=s_level
--out_file=<output_csv>
python plan_simp/scripts/eval_simp.py
--input_data=newselaauto_docs_test.csv
--output_data=test_out_ledpara.csv
--x_col=complex_str
--r_col=simple_str
--y_col=pred
--doc_id_col=pair_id
--prepro=True
--sent_level=True