Downloads · 30 days
8
14% of all-time downloads
SE6446/Untitled7-colab_checkpoint
Untitled7-colab_checkpoint is a image-to-text model from SE6446. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as unlicense.
This model was lovingly named after the Google Colab notebook that made it. It is a finetune of Microsoft's git-large-coco model on the 1k subset of poloclub/diffusiondb.
Downloads · 30 days
8
14% of all-time downloads
All-time downloads
59
Public
Parameters
394M
3.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin1.6 GB · 50%
From the Hugging Face model README
This model was lovingly named after the Google Colab notebook that made it. It is a finetune of Microsoft's git-large-coco model on the 1k subset of poloclub/diffusiondb.
It is supposed to read images and extract a stable diffusion prompt from it but, it might not do a good job at it. I wouldn't know I haven't extensivly tested it.
As the title suggests this is a checkpoint as I formerly intended to do it on the entire dataset but, I'm unsure if I want to now...
This is my first public model so please be nice!
Fun!
# Load model directly
from transformers import AutoProcessor, AutoModelForCausalLM
processor = AutoProcessor.from_pretrained("SE6446/Untitled7-colab_checkpoint")
model = AutoModelForCausalLM.from_pretrained("SE6446/Untitled7-colab_checkpoint")
#################################################################
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("image-to-text", model="SE6446/Untitled7-colab_checkpoint")
Don't use this model to discriminate, alienate or in any other way harm/harass individuals. You guys know the drill...
This model does not produce accurate prompts, this is merely a bit of fun (and waste of funds). However it can suffer from bias present in the orginal git-large-coco model.
I.e boring stuff
If you want to further finetune it then you should freeze the embedding and vision tranformer layers