Downloads · 30 days
143
31% of all-time downloads
dinalt/walsh_instruct-1-7b
walsh_instruct-1-7b is a text generation model from dinalt. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
143
31% of all-time downloads
All-time downloads
458
Public
Parameters
1.7B
3.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.5 GB · 100%
From the Hugging Face model README
Walsh_Instruct-1.7b
This is an instruction tuned fork of my "dinalt/walsh-1-7b" model... mostly for fun.
Hadamard-Walsh 1.7B is an experimental model using a new positional encoder. The encoder represents absolute positions by using a combination of rows from the Hadamard-Walsh matrix (https://en.wikipedia.org/wiki/Hadamard_code). Each row corresponds to a binary digit is the positional code, where the presence of a row codes for a 1 and the absence, a zero. While training, the base offset in the sequence is randomly chosen for each batch. The result is that the model is very proficient at sequences much longer than those seen in training.
Aside from the unsual positional encoder, the most interesting aspect of this model is the application of DITTO training:
Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation https://arxiv.org/abs/2206.02369
As described in the paper, the procedure is very effective at eliminating sentence level repition. As described in the paper, it also reduces perplexity slightly.
I will see about posting the code for running the training and generating a DITTO dataset later, althogh the "ditto-loss" function is already in the model implementation.
This is a toy instruciton following model. It's occasionally reliable at following directions.
[More Information Needed]
This is an uncensored instruction following model. No attempt has been made to make the model "safe." It may offend your sensibilities. It will likely provide inaccurate information. Use at your own risk. Whatever you do, don't put it in charge of the global defense grid!
The easiest way to get started with the model is to use text-generation-webui, which needs to be started with the "--trust-remote-code" flag.
https://github.com/oobabooga/text-generation-webui
It appears to work best with the "Big O" and "Simple-1" generation presets.
As an instruction model, the model has been trained to use the ChatML instruction format:
<|im_start|>system
Provide some context and/or instructions to the model.
<|im_end|>
<|im_start|>user
The user’s message goes here
<|im_end|>
<|im_start|>assistant
For details, see: https://github.com/MicrosoftDocs/azure-docs/blob/main/articles/ai-services/openai/includes/chat-markup-language.md#chatml
The model implementation is all my own, so you will need to use "trust_remote_code" to load the model.
from transformers import (
AutoTokenizer,
AutoModelForCausalLM,
)
model_id = "dinalt/walsh-1-7b"
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
# flash_attention_2 requires bfloat16 or float16
torch_dtype=torch.bfloat16,
# One of ["flash_attention_2", "sdpa", "eager"]
attn_implementation="flash_attention_2",
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
For batch instruction generation, see my example code here: https://discuss.huggingface.co/t/implimentation-of-stopping-criteria-list/20040/16?u=dinalt
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
It keeps my house warm in the winter...
[More Information Needed]
[More Information Needed]
6 x RTX4090
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]