Downloads · 30 days
8
4% of all-time downloads
Narrator5000/llavanext-finetuned-stackoverflow-vqa
llavanext-finetuned-stackoverflow-vqa is a machine learning model from Narrator5000. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
Finetuned LLaVA-Next model for Visual QA on Stack Overflow questions with images.
Downloads · 30 days
8
4% of all-time downloads
All-time downloads
209
Public
Repo size
797 MB
Likes
1
Public
Click a slice to open those files.
.safetensors625 MB · 78%
From the Hugging Face model README
Finetuned LLaVA-Next model for Visual QA on Stack Overflow questions with images.
This model is a finetuned version of LLaVA-Next (llava-hf/llava-v1.6-mistral-7b-hf) specifically for visual question answering (VQA) on Stack Overflow questions containing images. The model was finetuned using QLoRA with 4-bit quantization, optimized to handle both text and image inputs.
The training dataset was filtered from the mirzaei2114/stackoverflowVQA-filtered-small dataset. Only samples with a maximum input length of 1024 (for both question and answer combined) were used. Images were kept to size to capture detail needed for methods such as optical character recognition.
Drag a snipping rectangle for a screenshot around the exact focus/context for a question related to software development(usually front end) and accompany it with the question for inference.
Visual Question Answering (VQA) on technical Stack Overflow (software-adjacent) questions with accompanying images.
General-purpose VQA tasks, though performance on non-technical domains may vary.
Model Capacity: The model was trained using 4-bit QLoRA. Dataset Size: The training dataset is relatively small, and this may impact generalization to other VQA datasets or domains outside of Stack Overflow.
To use this model, ensure you have the following dependencies installed: torch==2.4.1+cu121 transformers==4.45.1
Do inference according to this multi-image inference llava-next example: https://huggingface.co/docs/transformers/en/model_doc/llava_next#:~:text=skip_special_tokens%3DTrue))-,Multi%20image%20inference,-LLaVa%2DNext%20can
mirzaei2114/stackoverflowVQA-filtered-small
TrainingArguments( per_device_train_batch_size=4, per_device_eval_batch_size=4, max_grad_norm=0.1, evaluation_strategy="steps", eval_steps=15, group_by_length=True, logging_steps=15, gradient_checkpointing=True, gradient_accumulation_steps=2, num_train_epochs=3, weight_decay=0.1, warmup_steps=10, lr_scheduler_type="cosine", learning_rate=1e-5, save_steps=15, save_total_limit=5, bf16=True, remove_unused_columns=False )
checkpoint-240
Evaluation Loss (Pre-finetuning): 2.93 Validation Loss (Post-finetuning): 1.78
mirzaei2114/stackoverflowVQA-filtered-small
L4 GPU
Google Colab