Downloads · 30 days
9
0% of all-time downloads
bubblspace/Bubbl-P4-multimodal-instruct
Bubbl-P4-multimodal-instruct is a machine learning model from bubblspace. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains a 4-bit quantized version of the microsoft/Phi-4-multimodal-instruct model.
Downloads · 30 days
9
0% of all-time downloads
All-time downloads
5.9K
Public
Parameters
5.7B
4 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors4 GB · 99%
How the weights are stored.
U85B · 87%
From the Hugging Face model README
This repository contains a 4-bit quantized version of the microsoft/Phi-4-multimodal-instruct model.
Quantization was performed using the bitsandbytes library integrated with transformers.
bitsandbytes Post-Training Quantization (PTQ)load_in_4bit=Truebnb_4bit_quant_type="nf4" (NormalFloat 4-bit)bnb_4bit_compute_dtype=torch.bfloat16 (Computation performed in BF16 for compatible GPUs like A100)bnb_4bit_use_double_quant=True (Enables nested quantization for potentially more memory savings)This version was created to provide the capabilities of Phi-4-multimodal with a significantly reduced memory footprint, making it suitable for deployment on GPUs with lower VRAM.
This quantized model is primarily intended for scenarios where VRAM resources are constrained, but the advanced multimodal reasoning, language understanding, and instruction-following capabilities of Phi-4-multimodal-instruct are desired.
Refer to the original model card for the full range of intended uses and capabilities of the base model.
You can load this 4-bit quantized model directly using the transformers library. Ensure you have bitsandbytes and accelerate installed (pip install transformers bitsandbytes accelerate torch torchvision pillow soundfile scipy sentencepiece protobuf).
from transformers import AutoModelForCausalLM, AutoProcessor
import torch
model_id = "bubblspace/Bubbl-P4-multimodal-instruct"
# Load the processor (requires trust_remote_code)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
# Load the model with 4-bit quantization enabled
# The quantization config is loaded automatically from the model's config file
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True, # Essential for Phi-4 models
load_in_4bit=True, # Explicitly activate 4-bit loading (though config should handle it)
device_map="auto" # Automatically map model layers to available GPU(s)
# torch_dtype=torch.bfloat16 # Often not needed here as bnb_4bit_compute_dtype is handled
)
print("4-bit quantized model loaded successfully!")
# --- Example: Text Inference ---
prompt = "<|user|>\nExplain the benefits of model quantization.<|end|>\n<|assistant|>"
inputs = processor(text=prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=150)
response_text = processor.batch_decode(outputs)[0]
print(response_text)
# --- Example: Image Inference Placeholder ---
# from PIL import Image
# import requests
# url = "your_image_url.jpg"
# image = Image.open(requests.get(url, stream=True).raw)
# image_prompt = "<|user|>\n<|image_1|>\nDescribe this image.<|end|>\n<|assistant|>"
# inputs = processor(text=image_prompt, images=image, return_tensors="pt").to(model.device)
# outputs = model.generate(**inputs, max_new_tokens=100)
# response_text = processor.batch_decode(outputs, skip_special_tokens=True)[0]
# print(response_text)
# --- Example: Audio Inference Placeholder ---
# import soundfile as sf
# audio_path = "your_audio.wav"
# audio_array, sampling_rate = sf.read(audio_path)
# audio_prompt = "<|user|>\n<|audio_1|>\nTranscribe this audio.<|end|>\n<|assistant|>"
# inputs = processor(text=audio_prompt, audios=[(audio_array, sampling_rate)], return_tensors="pt").to(model.device)
# # ... generate and decode ...
Important: Remember to always pass trust_remote_code=True when loading both the processor and the model for Phi-4 architectures.
microsoft/Phi-4-multimodal-instruct model. Please refer to its model card for detailed information on responsible AI practices.The model is licensed under the MIT License, consistent with the original microsoft/Phi-4-multimodal-instruct model.
Please cite the original work if you use this model:
@misc{phi4multimodal2025,
title={Phi-4-multimodal: A Compact Multimodal Model for Recommendation, Recognition, and Reasoning},
author={Microsoft},
year={2025},
eprint={2503.01743},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Additionally, if you use this specific 4-bit quantized version, please acknowledge **Bubblspace** ([bubblspace.com](https://bubblspace.com)) and **AIEDX** ([aiedx.com](https://aiedx.com)) for providing this quantized model. You could add a note such as:
> *"We used the 4-bit quantized version of Phi-4-multimodal-instruct provided by Bubblspace/AIEDX, available at huggingface.co/bubblspace/Bubbl-P4-multimodal-instruct."*