Downloads · 30 days
7
22% of all-time downloads
KYAGABA/combined-multimodal-model
combined-multimodal-model is a machine learning model from KYAGABA. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model performs medical image classification and report generation using a custom architecture that combines a video model and a text generation model.
Downloads · 30 days
7
22% of all-time downloads
All-time downloads
32
Public
Repo size
801 MB
Likes
0
Public
Click a slice to open those files.
.bin801 MB · 99%
From the Hugging Face model README
This model performs medical image classification and report generation using a custom architecture that combines a video model and a text generation model.
r3d_18) and BioBART.import torch
from transformers import AutoTokenizer
from model import CombinedModel, ImageToTextProjector
from torchvision import models
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("YOUR_HF_USERNAME/combined-multimodal-model")
# Initialize models
video_model = models.video.r3d_18(pretrained=True)
video_model.fc = torch.nn.Linear(video_model.fc.in_features, 512)
report_generator = AutoModelForSeq2SeqLM.from_pretrained("GanjinZero/biobart-v2-base")
projector = ImageToTextProjector(512, report_generator.config.d_model)
num_classes = 4
combined_model = CombinedModel(video_model, report_generator, num_classes, projector)
# Load state dict
state_dict = torch.hub.load_state_dict_from_url(
"https://huggingface.co/YOUR_HF_USERNAME/combined-multimodal-model/resolve/main/pytorch_model.bin",
map_location=torch.device('cpu')
)
combined_model.load_state_dict(state_dict)
combined_model.eval()
# Now you can use combined_model for inference