Downloads · 30 days
45
18% of all-time downloads
akhauriyash/DeepSeek-R1-Distill-Llama-8B-Butler
DeepSeek-R1-Distill-Llama-8B-Butler is a feature extraction model from akhauriyash. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
<div align="center" <img src="https://github.com/abdelfattah-lab/TokenButler/blob/main/figs/tokenbutlerlogo.png?raw=true" width="50%" alt="TokenButler" / </div <hr <div align="center" style="line-height: 1;" <a href="…
Downloads · 30 days
45
18% of all-time downloads
All-time downloads
249
Public
Parameters
8.1B
74.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin32.5 GB · 50%
From the Hugging Face model README
The collection of TokenButler models can be found here. To run the DeepSeek-R1-Distill-Llama-8B model, follow:
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline
question = "If millionaires have butlers, why don't million dollar language models have a butler too? I think its because "
model_name = "akhauriyash/DeepSeek-R1-Distill-Llama-8B-Butler"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True)
generator = pipeline("text-generation", model=model, tokenizer=tokenizer)
response = generator(question, max_new_tokens=200, do_sample=True, top_p=0.95, temperature=0.7)
print(response[0]['generated_text'][len(question):])
Note that the 'default' configured sparsity is 50%. Further, there is a 'sliding window' of 128 and 8 'anchor tokens'. To 'change' the sparsity, you can use the following function after loading the model. Please note that the 'fixed' is the only supported strategy at the moment, which 'fixes' the sparsity of each layer (except the first) at the 'pc' (percentage) mentioned. This can also be found at test_hf.py. Sliding window and anchor tokens can be changed in a similar manner.
def set_sparsity(model, sparsity):
for module in model.modules():
if module.__class__.__name__.__contains__("AttentionExperimental"):
module.token_sparse_method = sparsity
module.set_token_sparsity()
return model
model = set_sparsity(model, "fixed_60pc")