Downloads · 30 days
53
6% of all-time downloads
TheStageAI/Elastic-Mistral-7B-Instruct-v0.3
Elastic-Mistral-7B-Instruct-v0.3 is a text generation model from TheStageAI. Use it when you need the model to write or continue text. The card lists the license as other.
Downloads · 30 days
53
6% of all-time downloads
All-time downloads
879
Public
Repo size
281 GB
Likes
5
Public
Click a slice to open those files.
.enc281 GB · 100%
From the Hugging Face model README
ElasticModels are the models produced by TheStage AI ANNA: Automated Neural Networks Accelerator. ANNA allows you to control model size, latency and quality with a simple slider movement, routing different compression algorithms to different layers. For each model, we have produced a series of optimized models:
Models can be accessed via TheStage AI Python SDK: ElasticModels, or deployed as Docker containers with REST API endpoints (see Deploy section).
| Property | Value |
|---|---|
| GPU | H100, L40s, B200, RTX 5090 |
| Python Version | 3.10-3.12 |
| CPU | Intel/AMD x86_64 |
| CUDA Version | 12.9+ |
Install TheStage AI CLI and setup API token:
pip install thestage
thestage config set --access-token <YOUR_ACCESS_TOKEN>
Install TheStage Elastic Models package:
pip install 'thestage-elastic-models[nvidia,cudnn]' \
--extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
pip install --force-reinstall --no-deps nvidia-cudnn-frontend==1.18.0
Elastic Models provides the same interface as HuggingFace Transformers. Here is an example of how to use the Mistral-7B-Instruct-v0.3 model:
import torch
from transformers import AutoTokenizer
from elastic_models.transformers import AutoModelForCausalLM
# Currently we require to have your HF token
# as we use original weights for part of layers and
# model configuration as well
model_name = "mistralai/Mistral-7B-Instruct-v0.3"
hf_token = ''
device = torch.device("cuda")
# Create mode
tokenizer = AutoTokenizer.from_pretrained(
model_name, token=hf_token
)
model = AutoModelForCausalLM.from_pretrained(
model_name,
token=hf_token,
torch_dtype=torch.bfloat16,
attn_implementation="sdpa",
mode='S'
).to(device)
model.generation_config.pad_token_id = tokenizer.eos_token_id
# Inference simple as transformers library
prompt = "Describe basics of DNNs quantization."
messages = [
{
"role": "system",
"content": "You are a search bot, answer on user text queries."
},
{
"role": "user",
"content": prompt
}
]
chat_prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, tokenize=False
)
inputs = tokenizer(chat_prompt, return_tensors="pt")
inputs.to(device)
with torch.inference_mode():
generate_ids = model.generate(**inputs, max_length=500)
input_len = inputs['input_ids'].shape[1]
generate_ids = generate_ids[:, input_len:]
output = tokenizer.batch_decode(
generate_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False
)[0]
# Validate answer
print(f"# Q:\n{prompt}\n")
print(f"# A:\n{output}\n")
We have used the lm_eval library to validate the models. For each model size (S, M, L, XL), we have run the following tasks: MMLU, PIQA, Arc Challenge, Winogrande.

| Metric/Model Size | S | M | L | XL | Original |
|---|---|---|---|---|---|
| MMLU | 59.2 | 59.6 | 59.6 | 59.8 | 59.8 |
| PIQA | 81.3 | 81.3 | 81.9 | 81.9 | 82.0 |
| Arc Challenge | 59.6 | 60.4 | 59.5 | 60.3 | 59.7 |
| Winogrande | 75.2 | 76.1 | 75.3 | 74.8 | 74.8 |
We measured TPS (tokens per second) for each model size using 100 input tokens and 300 output tokens.

Tokens per second for different model sizes on various GPUs.
| GPU/Model Size | S | M | L | XL | Original |
|---|---|---|---|---|---|
| H100 | 213 | 197 | 180 | 149 | 60 |
| L40s | 78 | 68 | 61 | 48 | 39 |
| B200 | 289 | 287 | 276 | 256 | 104 |
| GeForce RTX 5090 | 162 | N/A | N/A | N/A | 74 |
The benchmarking was performed on a single GPU with a batch size of 1. Each model was run for 10 iterations, and the average latency was calculated.
Algorithm summary:
- Load the Mistral-7B-Instruct-v0.3 model with the specified size (S, M, L, XL, original).
- Move the model to the GPU.
- Prepare a sample prompt for text generation.
- Run the model for a number of iterations (e.g., 10) and measure the time taken for each iteration. On each iteration:
- Synchronize the GPU to flush any previous operations.
- Record the start time.
- Generate the text using the model.
- Synchronize the GPU again.
- Record the end time and calculate the TTFT and TPS for that iteration.
- Calculate the average TTFT and TPS over all iterations.
For serving with Nvidia GPUs, we provide ready-to-go Docker containers with OpenAI-compatible API endpoints. Using our containers you can set up an inference endpoint on any desired cloud/serverless providers as well as on-premise servers. You can also use this container to run inference through TheStage AI platform.
Pull docker image and start inference container:
docker pull public.ecr.aws/i3f7g5s7/thestage/elastic-models:0.2.1.post0-llm-24.09a
docker run --rm -ti \
--name serving_thestage_model \
-p 8000:80 \
-e AUTH_TOKEN=<AUTH_TOKEN> \
-e MODEL_REPO=mistralai/Mistral-7B-Instruct-v0.3 \
-e MODEL_SIZE=<MODEL_SIZE> \
-e MODEL_BATCH=<MAX_BATCH_SIZE> \
-e HUGGINGFACE_ACCESS_TOKEN=<HUGGINGFACE_ACCESS_TOKEN> \
-e THESTAGE_AUTH_TOKEN=<THESTAGE_ACCESS_TOKEN> \
-v /mnt/hf_cache:/root/.cache/huggingface \
public.ecr.aws/i3f7g5s7/thestage/elastic-models:0.2.1.post0-llm-24.09a
| Parameter | Description |
|---|---|
<MODEL_SIZE> | Available: S, M, L, XL. |
<MAX_BATCH_SIZE> | Maximum batch size to process in parallel. |
<HUGGINGFACE_ACCESS_TOKEN> | Hugging Face access token. |
<THESTAGE_ACCESS_TOKEN> | TheStage token generated on the platform (Profile -> Access tokens). |
<AUTH_TOKEN> | Token for endpoint authentication. You can set it to any random string; it must match the value used by the client. |
You can invoke the endpoint using CURL as follows:
curl -X POST 'http://127.0.0.1:8000/v1/chat/completions' \
-H 'Authorization: Bearer 123' \
-H 'Content-Type: application/json' \
-H "X-Model-Name: mistral-7b-instruct-v0-3-<MODEL_SIZE>-bs<MAX_BATCH_SIZE>-paged" \
-d '{
"messages":[{"role":"user","content":"Define AI"}]
}'
Or using OpenAI python client:
import os, base64, pathlib, json
from openai import OpenAI
BASE_URL = "http://<your_ip>/v1"
API_KEY = "123"
MODEL = "mistral-7b-instruct-v0-3-<MODEL_SIZE>-bs<MAX_BATCH_SIZE>-paged"
client = OpenAI(
api_key=API_KEY,
base_url=BASE_URL,
default_headers={"X-Model-Name": MODEL}
)
response = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "user", "content": "Define AI"}
]
)
print(response.choices[0].message.content)
POST
/v1/chat/completions
Authorization:stringBearer token for authentication. Should match the
AUTH_TOKENset during container startup.
Content-Type:stringMust be set to
application/json.
X-Model-Name:stringSpecifies the model to use for generation. Format:
mistral-7b-instruct-v0-3-<size>-bs<batch_size>, where<size>is one ofS,M,L,XL,originaland<batch_size>is the maximum batch size configured during container startup.
messages:stringThe input text prompt.
For more details please use the tutorial Modal deployment
git clone https://github.com/TheStageAI/ElasticModels.git
cd ElasticModels/examples/modal
Set your environment variables in modal_serving.py:
# modal_serving.py
ENVS = {
"MODEL_REPO": "mistralai/Mistral-7B-Instruct-v0.3",
"MODEL_BATCH": "4",
"THESTAGE_AUTH_TOKEN": "",
"HUGGINGFACE_ACCESS_TOKEN": "",
"PORT": "80",
"PORT_HEALTH": "80",
"HF_HOME": "/cache/huggingface",
}
Set your desired GPU type and autoscaling variables in modal_serving.py:
# modal_serving.py
@app.function(
image=image,
gpu="B200",
min_containers=8,
max_containers=8,
timeout=10000,
ephemeral_disk=600 * 1024,
volumes={"/opt/project/.cache": HF_CACHE},
startup_timeout=60*20
)
@modal.web_server(
80,
label="mistralai/Mistral-7B-Instruct-v0.3-test",
startup_timeout=60*20
)
def serve():
pass
modal serve modal_serving.py