Downloads · 30 days
21
2% of all-time downloads
Drenel/Hippo-6B
Hippo-6B is a text generation model from Drenel. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Hippo-6B is a cutting-edge, transformer-based language model designed to provide state-of-the-art performance across a wide range of natural language processing tasks. With 6.2 billion parameters, Hippo-6B strikes a b…
Downloads · 30 days
21
2% of all-time downloads
All-time downloads
1.3K
Public
Parameters
6.2B
12.5 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors12.5 GB · 100%
From the Hugging Face model README
Hippo-6B is a cutting-edge, transformer-based language model designed to provide state-of-the-art performance across a wide range of natural language processing tasks. With 6.2 billion parameters, Hippo-6B strikes a balance between computational efficiency and high performance, making it a versatile model for various applications.
Context Length: Supports up to 4K context length
Publisher: Drenel
Paper: Model Paper
flash_attn_func and flash_attn_varlen_func), to efficiently compute attention scores. This reduces the computational overhead and memory usage, enabling the model to handle longer context lengths without performance degradation.RotaryEmbedding) to encode positional information in a more continuous and differentiable manner, enhancing the model's ability to capture long-range dependencies.SuScaledRotaryEmbedding and YarnScaledRotaryEmbedding adapt the rotary embeddings to different scaling factors, providing finer control over the embedding space.RMSNorm) to stabilize training and improve convergence. RMS normalization helps in maintaining consistent gradient flow across layers, leading to more efficient training dynamics.Attention, FlashAttention2, SdpaAttention). This modularity allows easy customization and scalability of the attention mechanisms based on specific use cases.MLP class includes techniques such as expert gating and intermediate projections for more sophisticated representations.Cache, DynamicCache) to optimize memory usage during inference, allowing for faster and more efficient processing of long sequences.Hippo-6B can be used for a variety of NLP tasks, including but not limited to:
You can provide the prompt as a question with a generic template as follow:
<|user|>\nQuestion<|end|>\n<|assistant|>
Here is a quick example of how to use Hippo-6B for text generation:
# Libraries installation
# pip install -q transformers accelerate flash-attn
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
torch.random.manual_seed(0)
modelName = "Drenel/Hippo-6B"
model = AutoModelForCausalLM.from_pretrained(modelName, device_map="cuda",torch_dtype="auto",trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(modelName)
messages = [
{"role": "user", "content": "What is the capital of France? <|end|><|assistant|>"},
]
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
generation_args = {"max_new_tokens": 50, "return_full_text": False, "temperature": 0.7, "do_sample": False, "top_k": 50, "top_p": 0.95}
output = pipe(messages, **generation_args)
print(output[0]['generated_text'])
Hippo-6B is distributed under the Apache-2.0.