Downloads · 30 days
251
0% of all-time downloads
BCCard/Qwen2.5-VL-32B-Instruct-FP8-Dynamic
Qwen2.5-VL-32B-Instruct-FP8-Dynamic is a image-text-to-text model from BCCard. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
- Model Architecture: Qwen2.5-VL-32B-Instruct - Input: Vision-Text - Output: Text - Model Optimizations: - Weight quantization: FP8 - Activation quantization: FP8 - Release Date: 5/3/2025 - Version: 1.0 - Model Develo…
Downloads · 30 days
251
0% of all-time downloads
All-time downloads
287K
Public
Parameters
33.5B
69.2 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors35 GB · 100%
How the weights are stored.
F8_E4M331.9B · 95%
From the Hugging Face model README
Quantized version of Qwen/Qwen2.5-VL-32B-Instruct.
This model was obtained by quantizing the weights of Qwen/Qwen2.5-VL-32B-Instruct to FP8 data type, ready for inference with vLLM >= 0.5.2.
This model can be deployed efficiently using the vLLM backend, as shown in the example below.
from vllm.assets.image import ImageAsset
from vllm import LLM, SamplingParams
# prepare model
llm = LLM(
model="BCCard/Qwen2.5-VL-32B-Instruct-FP8-Dynamic",
trust_remote_code=True,
max_model_len=4096,
max_num_seqs=2,
)
# prepare inputs
question = "What is the content of this image?"
inputs = {
"prompt": f"<|user|>\n<|image_1|>\n{question}<|end|>\n<|assistant|>\n",
"multi_modal_data": {
"image": ImageAsset("cherry_blossom").pil_image.convert("RGB")
},
}
# generate response
print("========== SAMPLE GENERATION ==============")
outputs = llm.generate(inputs, SamplingParams(temperature=0.2, max_tokens=64))
print(f"PROMPT : {outputs[0].prompt}")
print(f"RESPONSE: {outputs[0].outputs[0].text}")
print("==========================================")
vLLM also supports OpenAI-compatible serving. See the documentation for more details.