Mistral Inference
Mistral Inference - Run Mistral AI Models Locally on Your Hardware
Quick facts
- Best for
- Mistral Inference - Run Mistral AI Models Locally on Your Hardware
- Pricing
- Free
- Editor rating
- 4.5 / 5
- Community saves
- 0
About Mistral Inference
Mistral Inference is the official Python inference library for Mistral AI models. It provides minimal, high-performance code to download, load, and run Mistral's open-weight models locally. The library supports the full Mistral model family including Mistral 7B, Mixtral 8x7B, Mixtral 8x22B, Codestral 22B, Codestral Mamba, Mathstral, Mistral Nemo, Mistral Large 2, Pixtral 12B, and Mistral Small 3.1. It offers CLI tools for quick testing and Python APIs for programmatic inference, with support for multi-GPU setups via torchrun. All models support function calling, and several offer custom licenses for research and non-commercial use.
Pros
- Download and run any Mistral AI open-weight model locally
- CLI demo tool for quick model testing (mistral-demo)
- Python API for programmatic inference
- Multi-GPU support via torchrun for large models
- Function calling support across all models
- Hugging Face Hub integration for model weights
- Optimized with xformers for efficient attention computation
- Support for instruction-tuned and base model variants
- Safetensors format for fast and safe model loading
- Extended 32K+ token vocabulary on newer model versions
Cons
Pricing
Open Source
$0
- • Full library access
- • All open-weight models
- • CLI and Python API
- • Apache 2.0 licensed code
- • Community support via Discord
