Downloads · 30 days
11
4% of all-time downloads
brandonbeiler/InternVL3_5-2B-FP8-Dynamic
InternVL3_5-2B-FP8-Dynamic is a image-text-to-text model from brandonbeiler. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a fp8 dynamic (w8a8) version of OpenGVLab/InternVL35-2B, optimized for high-performance inference with vLLM. The model utilizes fp8 dynamic (w8a8) for optimal performance and deployment.
Downloads · 30 days
11
4% of all-time downloads
All-time downloads
257
Public
Parameters
2.3B
3.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.3 GB · 100%
How the weights are stored.
F8_E4M31.4B · 60%
From the Hugging Face model README
This is a fp8 dynamic (w8a8) version of OpenGVLab/InternVL3_5-2B, optimized for high-performance inference with vLLM. The model utilizes fp8 dynamic (w8a8) for optimal performance and deployment.
You can serve the model using vLLM's OpenAI-compatible API server.
vllm serve brandonbeiler/InternVL3_5-2B-FP8-Dynamic \
--quantization compressed-tensors \
--served-model-name internvl3_5-2b \
--reasoning-parser qwen3 \
--trust-remote-code \
--max-model-len 32768 \
--tensor-parallel-size 1 # Adjust based on your GPU setup
Notes
This model was created using:
llmcompressor==0.7.1
compressed-tensors==0.10.2
transformers==4.55.0
torch==2.7.1
vllm==0.10.1.1
Quantized with ❤️ using LLM Compressor for the open-source community