Downloads · 30 days
49
5% of all-time downloads
brandonbeiler/InternVL3_5-4B-FP8-Dynamic
InternVL3_5-4B-FP8-Dynamic is a image-text-to-text model from brandonbeiler. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a fp8 dynamic (w8a8) version of OpenGVLab/InternVL35-4B, optimized for high-performance inference with vLLM. The model utilizes fp8 dynamic (w8a8) for optimal performance and deployment.
Downloads · 30 days
49
5% of all-time downloads
All-time downloads
928
Public
Parameters
4.7B
5.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.8 GB · 100%
How the weights are stored.
F8_E4M33.6B · 77%
From the Hugging Face model README
This is a fp8 dynamic (w8a8) version of OpenGVLab/InternVL3_5-4B, optimized for high-performance inference with vLLM. The model utilizes fp8 dynamic (w8a8) for optimal performance and deployment.
You can serve the model using vLLM's OpenAI-compatible API server.
vllm serve brandonbeiler/InternVL3_5-4B-FP8-Dynamic \
--quantization compressed-tensors \
--served-model-name internvl3_5-4b \
--reasoning-parser qwen3 \
--trust-remote-code \
--max-model-len 32768 \
--tensor-parallel-size 1 # Adjust based on your GPU setup
Notes
This model was created using:
llmcompressor==0.7.1
compressed-tensors==latest
transformers==4.55.0
torch==2.7.1
vllm==0.10.1.1
Quantized with ❤️ using LLM Compressor for the open-source community