Downloads · 30 days
215
35% of all-time downloads
Corianas/TinyTask-minipaca
TinyTask-minipaca is a text generation model from Corianas. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
A llama.c model based on Karpathy's Llama2.c project. https://github.com/karpathy/llama2.c
Downloads · 30 days
215
35% of all-time downloads
All-time downloads
620
Public
Parameters
88.1M
353 MB on disk
Likes
0
Public
Click a slice to open those files.
.bin176 MB · 50%
How the weights are stored.
F1688.1M · 100%
From the Hugging Face model README
A llama.c model based on Karpathy's Llama2.c project. https://github.com/karpathy/llama2.c
Vocab of 4096, trained on Tinystories, and my custom littlestories dataset (currently unreleased.)
This version was further trained on following instructions... somewhat... using https://github.com/mlabonne/llm-course/blob/main/Fine_tune_Llama_2_in_Google_Colab.ipynb
Model uses ↨ as a shift key, instead of using capial letters, this allowed simplification of the tokenizer to avoid duplicates that are uppercase.
To convert normal text to the right format I use:
def add_caseifer(text):
# Using list comprehension for more efficient concatenation
return ''.join(['↨' + char.lower() if char.isupper() else char for char in text])
To return the text to human format I use:
def remove_caseifer(text):
new_text = ""
i = 0
while i < len(text):
if text[i] == "↨":
if i+1 < len(text):
new_text += text[i+1].upper()
i += 1
else:
pass # skip this index
else:
new_text += text[i]
i += 1
return new_text