Downloads · 30 days
21
4% of all-time downloads
llmat/Mistral-Small-Instruct-2409-NVFP4
Mistral-Small-Instruct-2409-NVFP4 is a text generation model from llmat. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
NVFP4-quantized version of mistralai/Mistral-Small-Instruct-2409 produced with llmcompressor.
Downloads · 30 days
21
4% of all-time downloads
All-time downloads
496
Public
Parameters
12.7B
13.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors13.1 GB · 100%
How the weights are stored.
U810.9B · 86%
From the Hugging Face model README
NVFP4-quantized version of mistralai/Mistral-Small-Instruct-2409 produced with llmcompressor.
lm_head excluded)This model can be deployed efficiently using the vLLM backend, as shown in the example below.
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
model_id = "llmat/Mistral-Small-Instruct-2409-NVFP4"
number_gpus = 1
sampling_params = SamplingParams(temperature=0.6, top_p=0.9, max_tokens=256)
tokenizer = AutoTokenizer.from_pretrained(model_id)
messages = [
{"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
{"role": "user", "content": "Who are you?"},
]
prompts = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
llm = LLM(model=model_id, tensor_parallel_size=number_gpus)
outputs = llm.generate(prompts, sampling_params)
generated_text = outputs[0].outputs[0].text
print(generated_text)
vLLM also supports OpenAI-compatible serving. See the documentation for more details.