Downloads · 30 days
11
20% of all-time downloads
tursunali/bpt2
bpt2 is a text generation model from tursunali. Use it when you need the model to write or continue text. It is set up for transformers.
See the GPT2 model card for considerations on limitations and bias. See the GPT2 documentation for details on GPT2.
Downloads · 30 days
11
20% of all-time downloads
All-time downloads
54
Public
Repo size
6.8 GB
Likes
0
Public
Click a slice to open those files.
.bin3.4 GB · 100%
From the Hugging Face model README
See the GPT2 model card for considerations on limitations and bias. See the GPT2 documentation for details on GPT2.
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
tokenizer = AutoTokenizer.from_pretrained("tursunali/bpt2")
model = AutoModelForCausalLM.from_pretrained("tursunali/bpt2")
prompt = "<your prompt>"
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
print(pipe(prompt)[0]["generated_text"])
Also, two tricks might improve the generated text:
output = model.generate(
# during training an EOS token was used to mark the beginning of each text
# so it can help to insert it at the start
torch.tensor(
[tokenizer.eos_token_id] + tokenizer.encode(prompt)
).unsqueeze(0),
do_sample=True,
# try setting bad_words_ids=[[0]] to disallow generating an EOS token, without this the model is
# prone to ending generation early because a significant number of texts from the training corpus
# is quite short
bad_words_ids=[[0]],
max_length=max_length,
)[0]
print(tokenizer.decode(output))
Please cite BPT2 as follows:
@misc{Backpacker_Trail_German_large_2022,
author = {BackpackerTrail, Tursunali Kholdorov},
title = {{BPT2: Backpacker Trail German versions of BPT2}},
url = {https://github.com/Tursunali-Kholdorov/bptTrainer},
year = {2022}
}