Downloads · 30 days
8
42% of all-time downloads
JoVal26/ja-med-clef-model
ja-med-clef-model is a machine learning model from JoVal26. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
Downloads · 30 days
8
42% of all-time downloads
All-time downloads
19
Public
Repo size
507 MB
Likes
0
Public
Click a slice to open those files.
.safetensors169 MB · 90%
From the Hugging Face model README
LLaVA-LLaMA 3 8B Fine-tuned Adapter
This repository contains the LoRA adapter weights and processor files for fine-tuning of LLaVA LLaMA 3 8B on the ROCO_v2 dataset for dual tasks:
The adapter allows applying the fine-tuning over the existing base model for reproducibility and lightweight sharing of modifications.
tags:
This repository contains the LoRA fine-tuned adapter for LLaVA LLaMA 3 8B applied to our dataset.
The base model used is:
xtuner/llava-llama-3-8b-v1_1-transformersYou must load this base model and apply this adapter to reproduce our results.
from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel from huggingface_hub import snapshot_download
adapter_path = snapshot_download(repo_id="JoVal26/ja-med-clef-model")
model_base = AutoModelForCausalLM.from_pretrained("xtuner/llava-llama-3-8b-v1_1-transformers")
tokenizer = AutoTokenizer.from_pretrained(f"{adapter_path}/final_processor")
model = PeftModel.from_pretrained(model_base, f"{adapter_path}/final_lora_adapter_explicit")
This model is designed for use in clinical vision-language research and evaluation challenges with the following subtasks:
Identify presence and location of relevant clinical concepts from medical images.
It serves as a building block for scene understanding and supports downstream image retrieval and clinical decision support.
Evaluation metrics: precision, recall, F1 (set coverage metrics).
Generate coherent medical captions describing an image, focusing on the interplay of visible elements and detected concepts for full-scene interpretation.
Evaluation metrics: BLEU, CIDEr, METEOR, ROUGE, and clinical concept coverage.
inputs = tokenizer("Describe the clinical findings in this radiology image:", return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This model is for research purposes only. It has been fine-tuned on an extended version of ROCO_v2 and may not generalize to other medical imaging datasets or modalities. Clinical use without additional validation is NOT recommended.
This fine-tuned adapter repo: JoVal26/ja-med-clef-model
This work is released under CC BY-NC 4.0 license. For non-commercial research use only.
Fine-tuned model from xtuner/llava-llama-3-8b-v1_1-transformers.
transformers' direct loading (no local weight download).License: CC BY-NC 4.0 (non-commercial use only)
BibTeX:
If you use this adapter in academic work or publications, please cite:
@misc{JoVal26-ja-med-clef-model,
title = {LLaVA-LLaMA 3 8B Fine-tuned Adapter},
author = {JoVal26},
year = {2025},
howpublished = {\url{https://huggingface.co/JoVal26/ja-med-clef-model}},
note = {Fine-tuned adapter for LLaVA-LLaMA 3 8B for medical vision-language tasks.}
}