Downloads · 30 days
17
27% of all-time downloads
jtlucas/pyds_sum
pyds_sum is a summarization model from jtlucas. Use it when you need a shorter version of a longer text. It is set up for transformers. The card lists the license as mit.
This model performs abstract summarization of python data science code to english natural language. It is finetuned from google/flan-t5-small with a subset of Meta Kaggle For Code labeled with a 43B model.
Downloads · 30 days
17
27% of all-time downloads
All-time downloads
64
Public
Parameters
77M
309 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors308 MB · 99%
From the Hugging Face model README
This model performs abstract summarization of python data science code to english natural language. It is finetuned from google/flan-t5-small with a subset of Meta Kaggle For Code labeled with a 43B model.
This model was finetuned from the google/flan-t5-small and shares its architecture and tokenizer.
Code cells were extracted from Jupyter Notebooks, chunked into ~500 tokens, and labelled by a 43B model with the prompt: "Think step by step and then provide a two or three sentence summary of what the code is doing for an audience who may not be familiar with machine learning. Focus on the problem the authors' are trying to solve."
All code was extracted from .ipynb files that are part of the Meta Kaggle for Code dataset.
The tokenizer was not modified from the standard google/flan-t5-small tokenizer.
The model is available for use in the transformers library, and can be used as a pre-trained checkpoint for inference or for fine-tuning on another dataset.
## Generating summaries with this model
```python
ipynb_string = "import pandas as pd\nimport numpy as np"
tokenizer = AutoTokenizer.from_pretrained(model_checkpoint)
model = AutoModelForSeq2SeqLM.from_pretrained(model_checkpoint)
chunk_ids = tokenizer.encode("summarize: ```" + ipynb_string + "```", return_tensors="pt", truncation=True, padding="max_length", max_length=512)
output_tokens = model.generate(chunk_ids, max_length=128)
output_text = tokenizer.decode(output_tokens[0], skip_special_tokens=True)
This model accepts 512 tokens from the associated tokenizer. Preface input data with summarize: and wrap input as a markdown code block "```".
This model provides short natural language summaries of python data science code.
The Flan-T5-Small architecture was chosen to maximize portability, but summaries may sometimes be repetitive, incomplete, or too abstract. Remember that the model was finetuned with Kaggle notebooks and will perform better for code in that distribution.