Downloads · 30 days
40
0% of all-time downloads
bczhou/TinyLLaVA-2.0B
TinyLLaVA-2.0B is a image-text-to-text model from bczhou. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<h2 align="center" <a href="https://arxiv.org/abs/2402.14289"TinyLLaVA: A Framework of Small-scale Large Multimodal Models</a
Downloads · 30 days
40
0% of all-time downloads
All-time downloads
8.8K
Public
Parameters
2B
4.1 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
From the Hugging Face model README
We recommend the requirements as follows.
git clone https://github.com/DLCV-BUAA/TinyLLaVABench.git
cd TinyLLaVABench
conda create -n tinyllava python=3.10 -y
conda activate tinyllava
pip install --upgrade pip # enable PEP 660 support
pip install -e .
pip install -e ".[train]"
pip install flash-attn --no-build-isolation
git pull
pip install -e .
# if you see some import errors when you upgrade, please try running the command below (without #)
# pip install flash-attn --no-build-isolation --no-cache-dir
| Name | LLM | Checkpoint | LLaVA-Bench-Wild | MME | MMBench | MM-Vet | SQA-image | VQA-v2 | GQA | TextVQA |
|---|---|---|---|---|---|---|---|---|---|---|
| TinyLLaVA-3.1B | Phi-2 | TinyLLaVA-3.1B | 75.8 | 1464.9 | 66.9 | 32.0 | 69.1 | 79.9 | 62.0 | 59.1 |
| TinyLLaVA-2.0B | StableLM-2-1.6B | TinyLLaVA-2.0B | 66.4 | 1433.8 | 63.3 | 32.6 | 64.7 | 78.9 | 61.9 | 56.4 |
| TinyLLaVA-1.5B | TinyLlama | TinyLLaVA-1.5B | 60.8 | 1276.5 | 55.2 | 25.8 | 60.3 | 76.9 | 60.3 | 51.7 |
Launch a local web demo by running:
python tinyllava/serve/app.py --model-path bczhou/TinyLLaVA-3.1B --model-name TinyLLaVA-3.1B
We also support running inference with CLI. To use our model, run:
python -m tinyllava.serve.cli \
--model-path bczhou/TinyLLaVA-3.1B \
--image-file "./tinyllava/serve/examples/extreme_ironing.jpg"
from tinyllava.model.builder import load_pretrained_model
from tinyllava.mm_utils import get_model_name_from_path
from tinyllava.eval.run_tiny_llava import eval_model
model_path = "bczhou/TinyLLaVA-3.1B"
tokenizer, model, image_processor, context_len = load_pretrained_model(
model_path=model_path,
model_base=None,
model_name=get_model_name_from_path(model_path)
)
</details>
Here's an example of running inference with TinyLLaVA-3.1B
<details> <summary>Run Inference</summary>from tinyllava.model.builder import load_pretrained_model
from tinyllava.mm_utils import get_model_name_from_path
from tinyllava.eval.run_tiny_llava import eval_model
model_path = "bczhou/TinyLLaVA-3.1B"
prompt = "What are the things I should be cautious about when I visit here?"
image_file = "https://llava-vl.github.io/static/images/view.jpg"
args = type('Args', (), {
"model_path": model_path,
"model_base": None,
"model_name": get_model_name_from_path(model_path),
"query": prompt,
"conv_mode": "phi",
"image_file": image_file,
"sep": ",",
"temperature": 0,
"top_p": None,
"num_beams": 1,
"max_new_tokens": 512
})()
eval_model(args)
</details>
We use different conv_mode for different models. Replace the conv_mode in args according to this table:
| model | conv_mode |
|---|---|
| TinyLLaVA-3.1B | phi |
| TinyLLaVA-2.0B | phi |
| TinyLLaVA-1.5B | v1 |
To ensure the reproducibility, we evaluate the models with greedy decoding.
See Evaluation.md
In our paper, we used two different datasets: the LLaVA dataset and the ShareGPT4V dataset, and compared their differences. In this section, we provide information on data preparation.
The majority of the two SFT datasets are the same, with the exception that the 23K detailed description data in LLaVA-1.5-SFT being replaced with detailed captions randomly sampled from the 100K ShareGPT4V data.
.jpgOrganize the image files and annotation files as follows in path/to/your/data:
data
├── llava
│ ├── llava_pretrain
│ │ ├── images
│ │ ├── blip_laion_cc_sbu_558k.json
├── coco
│ ├── train2017
├── sam
│ ├── images
├── gqa
│ ├── images
├── ocr_vqa
│ ├── images
├── textvqa
│ ├── train_images
├── vg
│ ├── VG_100K
│ ├── VG_100K_2
├── share_textvqa
│ ├── images
├── web-celebrity
│ ├── images
├── web-landmark
│ ├── images
├── wikiart
│ ├── images
├── text_files
│ ├── llava_v1_5_mix665k.json
│ ├── share-captioner_coco_lcs_sam_1246k_1107.json
│ ├── sharegpt4v_mix665k_cap23k_coco-ap9k_lcs3k_sam9k_div2k.json
This section we describe the base recipe.
Both hyperparameters used in pretraining and finetuning are provided below.
| Hyperparameter | Global Batch Size | Learning rate | Epochs | Max length | Weight decay |
|---|---|---|---|---|---|
| TinyLLaVA-3.1B | 256 | 1e-3 | 1 | 3072 | 0 |
| Hyperparameter | Global Batch Size | Learning rate | Epochs | Max length | Weight decay |
|---|---|---|---|---|---|
| TinyLLaVA-3.1B | 128 | 2e-5 | 1 | 3072 | 0 |
Replace paths to your paths
Training script with DeepSpeed ZeRO-2: pretrain.sh.
Replace paths to your paths
Training script with DeepSpeed ZeRO-3: finetune.sh.
Check out our custom finetune using LoRA here.
If you find our paper and code useful in your research, please consider giving a star :star: and citation :pencil:.
@misc{zhou2024tinyllava,
title={TinyLLaVA: A Framework of Small-scale Large Multimodal Models},
author={Baichuan Zhou and Ying Hu and Xi Weng and Junlong Jia and Jie Luo and Xien Liu and Ji Wu and Lei Huang},
year={2024},
eprint={2402.14289},
archivePrefix={arXiv},
primaryClass={cs.LG}
}