Downloads · 30 days
0
SobanHM/LLaVA-via
LLaVA-via is a image-to-text model from SobanHM. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.
LLaVA-via is a PEFT LoRA adapter fine-tuned on top of LLaVA-1.5-7B to improve scene narration and navigation guidance for assistive vision-language applications.
Downloads · 30 days
0
Access
Public
Updated Jun 29, 2026
Repo size
40.4 MB
Likes
1
Public
Click a slice to open those files.
.safetensors39.9 MB · 91%
From the Hugging Face model README
LLaVA-via is a PEFT LoRA adapter fine-tuned on top of LLaVA-1.5-7B to improve scene narration and navigation guidance for assistive vision-language applications.
The objective of this work is to generate contextually grounded, coherent, and navigation-aware responses that can support visually impaired assistance scenarios. Rather than answering general visual questions, the adapter is optimized to describe surrounding environments and provide guidance-oriented information from input images.
This repository contains only the LoRA adapter and associated tokenizer/processor configuration files. It does not contain the full LLaVA model weights.
This adapter must be loaded on top of the following Hugging Face model:
Base Model
llava-hf/llava-1.5-7b-hf
The base model is not included in this repository.
This repository includes:
adapter_model.safetensors — Fine-tuned LoRA adapter weightsadapter_config.json — PEFT configurationTraining checkpoints, optimizer states, and scheduler files are intentionally excluded.
LLaVA-via is intended for research and development involving assistive vision-language systems, including:
Example applications include:
q_proj, v_projThe fine-tuned adapter was evaluated using multiple automatic metrics covering lexical similarity, semantic similarity, vision-language alignment, and hallucination analysis.
| Metric | Score |
|---|---|
| BLEU-1 | 0.3125 |
| BLEU-2 | 0.1842 |
| BLEU-3 | 0.1228 |
| BLEU-4 | 0.0860 |
| METEOR | 0.3165 |
| ROUGE-L | 0.2755 |
| BERTScore-F1 | 0.8872 |
| SBERTScore | 0.7116 |
| CLIPScore (ViT-B/32) | 0.2839 |
| CLIPScore (ViT-B/16) | 0.2329 |
| CLIPScore (ViT-L/14) | 0.1828 |
| PickScore (CLIP-ViT-H-14) | 0.2030 |
| SPECS-VIA Score | 3.2770 / 5 |
| PickScore-VIA | 0.8132 |
| CHAIR-i (Semantically Refined) | 0.410 |
The evaluation includes lexical overlap metrics (BLEU, METEOR, ROUGE-L), semantic similarity metrics (BERTScore and SBERTScore), vision-language alignment metrics (CLIPScore and PickScore), a task-specific evaluation (SPECS-VIA and PickScore-VIA), and hallucination analysis using the semantically refined CHAIR-i metric.
The adapter can be loaded using the Hugging Face Transformers and PEFT libraries.
from transformers import AutoProcessor, LlavaForConditionalGeneration
from peft import PeftModel
base_model = LlavaForConditionalGeneration.from_pretrained(
"llava-hf/llava-1.5-7b-hf"
)
processor = AutoProcessor.from_pretrained(
"llava-hf/llava-1.5-7b-hf"
)
model = PeftModel.from_pretrained(
base_model,
"SobanHM/LLaVA-via"
)
The adapter is designed for prompts such as:
This repository provides a fine-tuned PEFT LoRA adapter evaluated using offline automatic metrics.
The reported evaluation does not include:
Accordingly, inference latency, energy consumption, real-time responsiveness, and human-subject usability remain outside the scope of the reported experiments.
If you use this repository in your research, please cite this work and the original LLaVA and PEFT publications.
A BibTeX entry for this work may be added once the associated thesis or publication becomes available.
Soban Hussain Department of Computer Science Sukkur IBA University
Praih Alias Faiza Department of Computer Science Sukkur IBA University
Tasmia Department of Computer Science Sukkur IBA University
This work builds upon the following open-source projects:
We gratefully acknowledge the developers and research community whose contributions made this work possible.