Downloads · 30 days
1.1K
22% of all-time downloads
QuantFactory/gemma-2b-aps-it-GGUF
gemma-2b-aps-it-GGUF is a text generation model from QuantFactory. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
Downloads · 30 days
1.1K
22% of all-time downloads
All-time downloads
5K
Public
Repo size
23.8 GB
Likes
1
Public
Click a slice to open those files.
.gguf23.8 GB · 100%
From the Hugging Face model README
This is quantized version of google/gemma-2b-aps-it created using llama.cpp
Model Page: Gemma
This model card corresponds to the 2B finetuned version of the Gemma-APS model. You can also visit the model card of the 7B finetuned model.
Resources and Technical Documentation:
Terms of Use: Terms
Authors: Mohammad Javad Hosseini, Yang Gao, Tim Baumgärtner, Alex Fabrikant, Reinald Kim Amplayo
Summary description and brief definition of inputs and outputs.
Gemma-APS is a generative model and a research tool for abstractive proposition segmentation (APS for short), a.k.a. claim extraction. Given a text passage, the model segments the content into the individual facts, statements, and ideas expressed in the text, and restates them in full sentences with small changes to the original text.
This model can be used for research where there is a need to break down text content into meaningful components. Applications include grounding, retrieval, fact-checking, and evaluation of generation tasks (such as summarization) where it can be useful to divide up individual propositions (claims) so that they can be processed independently. For more information, check out the research paper.
Models are trained on a context length of 8192 tokens.
Below we share some code snippets on how to get quickly started with running the model. First make sure to pip install -U transformers nltk,
then copy the snippet from the section that is relevant for your usecase.
For ease-of-use, we define two helper functions for pre-processing input and post-processing output of the model:
import nltk
import re
nltk.download('punkt')
start_marker = '<s>'
end_marker = '</s>'
separator = '\n'
def create_propositions_input(text: str) -> str:
input_sents = nltk.tokenize.sent_tokenize(text)
propositions_input = ''
for sent in input_sents:
propositions_input += f'{start_marker} ' + sent + f' {end_marker}{separator}'
propositions_input = propositions_input.strip(f'{separator}')
return propositions_input
def process_propositions_output(text):
pattern = re.compile(f'{re.escape(start_marker)}(.*?){re.escape(end_marker)}', re.DOTALL)
output_grouped_strs = re.findall(pattern, text)
predicted_grouped_propositions = []
for grouped_str in output_grouped_strs:
grouped_str = grouped_str.strip(separator)
props = [x[2:] for x in grouped_str.split(separator)]
predicted_grouped_propositions.append(props)
return predicted_grouped_propositions
pipeline APIfrom transformers import pipeline
import torch
generator = pipeline('text-generation', 'google/gemma-2b-aps-it', device_map='auto', torch_dtype=torch.bfloat16)
passage = 'Sarah Stage, 30, welcomed James Hunter into the world on Tuesday.\nThe baby boy weighed eight pounds seven ounces and was 22 inches long.'
messages = [{'role': 'user', 'content': create_propositions_input(passage)}]
output = generator(messages, max_new_tokens=4096, return_full_text=False)
result = process_propositions_output(output[0]['generated_text'])
print(result)
<details>
<summary>Example output</summary>
[
[
"Sarah Stage welcomed James Hunter into the world.",
"Sarah Stage welcomed James Hunter on Tuesday.",
"Sarah Stage is 30 years old."
],
[
"James Hunter weighed eight pounds seven ounces.",
"James Hunter was 22 inches long."
]
]
</details>
AutoModel and AutoTokenizer APIsfrom transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = 'google/gemma-2b-aps-it'
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map='auto',
torch_dtype=torch.bfloat16,
)
passage = "For more than 40 years, the lyrics of American Pie have been puzzled over. This week the handwritten lyrics sold for more than $1 million at auction. The verses contain hidden references to seminal events of the 50s and 60s. It includes nods to Buddy Holly, Charles Manson and Martin Luther King."
messages = [{'role': 'user', 'content': create_propositions_input(passage)}]
inputs = tokenizer.apply_chat_template(messages, return_tensors='pt', add_generation_prompt=True, return_dict=True).to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
generated_text = tokenizer.batch_decode(output[:, inputs['input_ids'].shape[1]:], skip_special_tokens=True)[0]
result = process_propositions_output(generated_text)
print(result)
<details>
<summary>Example output</summary>
[
[
"The lyrics of American Pie have been puzzled over.",
"The lyrics of American Pie have been puzzled for more than 40 years."
],
[
"This week the handwritten lyrics sold for more than $1 million.",
"This week the handwritten lyrics sold at auction."
],
[
"The verses contain hidden references to seminal events.",
"The verses contain hidden references to events of the 50s.",
"The verses contain hidden references to events of the 60s."
],
[
"The lyrics include nods to Buddy Holly.",
"The lyrics include nods to Charles Manson.",
"The lyrics include nods to Martin Luther King."
]
]
</details>
Data used for model training and how the data was processed.
See the research paper for all the details.
Details about the model internals.
Similar to Gemma, Gemma-APS was trained on TPUv5e.
Training large language models requires significant computational power. TPUs, designed specifically for matrix operations common in machine learning, offer several advantages in this domain:
Performance: TPUs are specifically designed to handle the massive computations involved in training LLMs. They can speed up training considerably compared to CPUs. Memory: TPUs often come with large amounts of high-bandwidth memory, allowing for the handling of large models and batch sizes during training. This can lead to better model quality. Scalability: TPU Pods (large clusters of TPUs) provide a scalable solution for handling the growing complexity of large foundation models. You can distribute training across multiple TPU devices for faster and more efficient processing. Cost-effectiveness: In many scenarios, TPUs can provide a more cost-effective solution for training large models compared to CPU-based infrastructure, especially when considering the time and resources saved due to faster training. These advantages are aligned with Google's commitments to operate sustainably.
Training was done using JAX.
JAX allows researchers to leverage the latest generation of hardware, including TPUs, for faster and more efficient training of large models.
Model evaluation metrics and results.
Evaluation was done on one existing in-domain dataset (development set of the ROSE dataset filtered by an entailment model) and two out-of-domain datasets introduced in the paper. Evaluation was performed based on our new metrics for the abstractive proposition segmentation task.
Ethics and safety evaluation approach and results.
These models are only suitable for abstractive proposition segmentation for English text, not any other task or language. While we have tested the models on three evaluation datasets and have obtained positive results compared to strong baselines, the model might still have errors on some examples.
These models have certain limitations that users should be aware of.
These models are only suitable for abstractive proposition segmentation for English text, not any other task or language. While we have tested it on three evaluation datasets and have obtained positive results compared to strong baselines, the models might still have errors on some examples.
These models have certain limitations that users should be aware of.
The development of large language models (LLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
Risks identified and mitigations:
These models are useful for academics working on abstractive proposition segmentation (claim extraction) research or other problems (e.g., grounding, retrieval, fact-checking) that could benefit from this task.