Downloads · 30 days
55
18% of all-time downloads
ucsbcit/InternVL3_5-38B-Flash-FP8
InternVL3_5-38B-Flash-FP8 is a image-text-to-text model from ucsbcit. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This repository contains an FP8 quantized version of the multi-modal model OpenGVLab/InternVL35-38B-Flash.
Downloads · 30 days
55
18% of all-time downloads
All-time downloads
311
Public
Parameters
39.6B
47.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors47.9 GB · 100%
How the weights are stored.
F8_E4M331.2B · 79%
From the Hugging Face model README
This repository contains an FP8 quantized version of the multi-modal model OpenGVLab/InternVL3_5-38B-Flash.
OpenGVLab/InternVL3_5-38B-FlashLinear layersllmcompressor (Post-Training Quantization / PTQ)ultrachat-200k (train_sft)The model was quantized using Neural Magic's llm-compressor framework.
Post-training quantization (PTQ) was applied directly to the language backbone (model.language_model). This strategy ensures that all heavy Linear projections in the main LLM layers are converted to FP8 for maximum speedup and reduced VRAM footprint, while preserving the vision architecture intact.
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8",
ignore=["lm_head"]
)
oneshot(
model=model.language_model,
tokenizer=tokenizer,
dataset="ultrachat-200k",
splits="train_sft[:512]",
recipe=recipe,
max_seq_length=2048,
num_calibration_samples=512,
)