Downloads · 30 days
16
3% of all-time downloads
LearnItAnyway/YOLO_LLaMa_7B_VisNav
YOLO_LLaMa_7B_VisNav is a text generation model from LearnItAnyway. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
This project aims to support visually impaired individuals in their daily navigation.
Downloads · 30 days
16
3% of all-time downloads
All-time downloads
504
Public
Repo size
13.5 GB
Likes
1
Public
Click a slice to open those files.
.bin13.5 GB · 100%
From the Hugging Face model README
This project aims to support visually impaired individuals in their daily navigation.
This project combines the YOLO model and LLaMa 2 7b for the navigation.
YOLO is trained on the bounding box data from the AI Hub,
Output of YOLO (bbox data) is converted as lists like [[class_of_obj_1, xmin, xmax, ymin, ymax, size], [class_of...] ...] then added to the input of question.
The LLM is trained to navigate using LearnItAnyway/Visual-Navigation-21k multi-turn dataset
We show how to use the model in yolo_llama_visnav_test.ipynb