Downloads · 30 days
9
2% of all-time downloads
pdich2085/new-blip
new-blip is a image-to-text model from pdich2085. Use it when you need a caption or text from an image. It is set up for generic. The card lists the license as bsd-3-clause.
This repository implements a custom task for image-captioning for 🤗 Inference Endpoints. The code for the customized pipeline is in the pipeline.py. To use deploy this model a an Inference Endpoint you have to select…
Downloads · 30 days
9
2% of all-time downloads
All-time downloads
451
Public
Repo size
5.6 GB
Likes
0
Public
Click a slice to open those files.
.h51.9 GB · 50%
From the Hugging Face model README
image-captioning task on 🤗Inference endpoint.This repository implements a custom task for image-captioning for 🤗 Inference Endpoints. The code for the customized pipeline is in the pipeline.py.
To use deploy this model a an Inference Endpoint you have to select Custom as task to use the pipeline.py file. -> double check if it is selected
{
"image": "/9j/4AAQSkZJRgA.....", #encoded image
"text": "a photography of a"
}
below is an example on how to run a request using Python and requests.
!wget https://storage.googleapis.com/sfr-vision-language-research/BLIP/demo.jpg
2.run request
import json
from typing import List
import requests as r
import base64
with open("/content/demo.jpg", "rb") as image_file:
encoded_string = base64.b64encode(image_file.read()).decode()
ENDPOINT_URL = ""
HF_TOKEN = ""
def query(payload):
response = requests.post(API_URL, headers=headers, json=payload)
return response.json()
output = query({
"inputs": {
"images": [encoded_string], # using the base64 encoded string
"texts": ["a photography of"] # Optional, based on your current class logic
}
})
print(output)
Example parameters depending on the decoding strategy:
"parameters": {
"num_beams":5,
"max_length":20
}
"parameters": {
"num_beams":1,
"max_length":20,
"do_sample": True,
"top_k":50,
"top_p":0.95
}
"parameters": {
"penalty_alpha":0.6,
"top_k":4
"max_length":512
}
See generate() doc for additional detail
expected output
{'captions': ['a photography of a woman and her dog on the beach']}