Downloads · 30 days
24
6% of all-time downloads
RedHatAI/Llama-2-7b-evol-code-alpaca-pruned_50-quantized-deepsparse
Llama-2-7b-evol-code-alpaca-pruned_50-quantized-deepsparse is a text generation model from RedHatAI. Use it when you need the model to write or continue text. It is set up for transformers.
This repo contains a 50% sparse Llama 2 7B finetuned for code generation tasks using the Evolved CodeAlpaca dataset. It was then quantized to 8-bit weights + activations and exported to deploy with DeepSparse, a CPU i…
Downloads · 30 days
24
6% of all-time downloads
All-time downloads
398
Public
Repo size
14.6 GB
Likes
0
Public
Click a slice to open those files.
.data7.2 GB · 100%
From the Hugging Face model README
This repo contains a 50% sparse Llama 2 7B finetuned for code generation tasks using the Evolved CodeAlpaca dataset. It was then quantized to 8-bit weights + activations and exported to deploy with DeepSparse, a CPU inference runtime for sparse models.
Official model weights from Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment.
Authors: Neural Magic, Cerebras
Below we share some code snippets on how to get quickly started with running the model.
By leveraging a pre-sparsified model's structure, you can efficiently fine-tune on new data, leading to reduced hyperparameter tuning, training times, and computational costs. Learn about this process here.
For accelerated inference with sparsity on CPUs, deploy with deepsparse.
# pip install deepsparse[llm]
from deepsparse import TextGeneration
model = TextGeneration(model_path="hf:neuralmagic/Llama-2-7b-pruned50-retrained-evolcodealpaca-quant-ds")
input_text = "def fibonacci(n):\n"
outputs = model(input_text, max_new_tokens=100)
print(outputs.generations[0].text)
Model evaluation metrics and results.
| Benchmark | Metric | Llama-2-7b-evolcodealpaca | Llama-2-7b-pruned50-retrained-evolcodealpaca-quant-ds |
|---|---|---|---|
| HumanEval | pass@1 | 32.03 | 36.34 |
For further support, and discussions on these models and AI in general, join Neural Magic's Slack Community