Downloads · 30 days
26
5% of all-time downloads
TheStageAI/Elastic-Llama-3.2-1B-Instruct
Elastic-Llama-3.2-1B-Instruct is a text generation model from TheStageAI. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Elastic models are the models produced by TheStage AI ANNA: Automated Neural Networks Accelerator. ANNA allows you to control model size, latency and quality with a simple slider movement. For each model, ANNA produce…
Downloads · 30 days
26
5% of all-time downloads
All-time downloads
568
Public
Repo size
20.1 GB
Likes
3
Public
Click a slice to open those files.
.enc18.9 GB · 100%
From the Hugging Face model README
Elastic models are the models produced by TheStage AI ANNA: Automated Neural Networks Accelerator. ANNA allows you to control model size, latency and quality with a simple slider movement. For each model, ANNA produces a series of optimized models:
XL: Mathematically equivalent neural network, optimized with our DNN compiler.
L: Near lossless model, with less than 1% degradation obtained on corresponding benchmarks.
M: Faster model, with accuracy degradation less than 1.5%.
S: The fastest model, with accuracy degradation less than 2%.
Goals of elastic models:
It's important to note that specific quality degradation can vary from model to model. For instance, with an S model, you can have 0.5% degradation as well.

To infer our models, you just need to replace transformers import with elastic_models.transformers:
import torch
from transformers import AutoTokenizer
from elastic_models.transformers import AutoModelForCausalLM
# Currently we require to have your HF token
# as we use original weights for part of layers and
# model configuration as well
model_name = "meta-llama/Llama-3.2-1B-Instruct"
hf_token = ''
device = torch.device("cuda")
# Create mode
tokenizer = AutoTokenizer.from_pretrained(
model_name, token=hf_token
)
model = AutoModelForCausalLM.from_pretrained(
model_name,
token=hf_token,
torch_dtype=torch.bfloat16,
attn_implementation="sdpa",
mode='S'
).to(device)
model.generation_config.pad_token_id = tokenizer.eos_token_id
# Inference simple as transformers library
prompt = "Describe basics of DNNs quantization."
messages = [
{
"role": "system",
"content": "You are a search bot, answer on user text queries."
},
{
"role": "user",
"content": prompt
}
]
chat_prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, tokenize=False
)
inputs = tokenizer(chat_prompt, return_tensors="pt")
inputs.to(device)
with torch.inference_mode():
generate_ids = model.generate(**inputs, max_length=500)
input_len = inputs['input_ids'].shape[1]
generate_ids = generate_ids[:, input_len:]
output = tokenizer.batch_decode(
generate_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False
)[0]
# Validate answer
print(f"# Q:\n{prompt}\n")
print(f"# A:\n{output}\n")
System requirements:
To work with our models just run these lines in your terminal:
pip install thestage
pip install 'thestage-elastic-models[nvidia]' --extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
pip install flash_attn==2.7.3 --no-build-isolation
pip uninstall apex
Then go to app.thestage.ai, login and generate API token from your profile page. Set up API token as follows:
thestage config set --api-token <YOUR_API_TOKEN>
Congrats, now you can use accelerated models!
Benchmarking is one of the most important procedures during model acceleration. We aim to provide clear performance metrics for models using our algorithms. The W8A8, int8 column indicates that we applied W8A8 quantization with int8 data type to all linear layers and used the same calibration data as for ANNA. The S model achieves practically identical speed but much higher quality, as ANNA knows how to improve quantization quality on sensitive layers!
| Metric/Model | S | M | L | XL | Original | W8A8, int8 |
|---|---|---|---|---|---|---|
| MMLU | 45.5 | 45.9 | 45.9 | 46.2 | 46.2 | 24 |
| PIQA | 73.1 | 73.7 | 74.2 | 74.3 | 74.3 | 55.8 |
| Arc Challenge | 34.5 | 35.9 | 36.0 | 35.8 | 35.8 | 20.3 |
| Winogrande | 60.4 | 59.7 | 60.8 | 59.5 | 59.5 | 50.3 |
100 input/300 output; tok/s:
| GPU/Model | S | M | L | XL | Original | W8A8, int8 |
|---|---|---|---|---|---|---|
| H100 | 436 | 436 | 409 | 396 | 110 | 439 |
| L40s | 290 | 251 | 222 | 210 | 103 | 300 |