Downloads · 30 days
11
3% of all-time downloads
createveai/Florence-2-base-PromptGen-v1.5
Florence-2-base-PromptGen-v1.5 is a machine learning model from createveai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is MiaoshouAI/Florence-2-base-PromptGen-v1.5 which retains its existing features, but with changes to supporting configuration and code to ensure drop-in replacement for Microsoft Florence-2 Model Base the when u…
Downloads · 30 days
11
3% of all-time downloads
All-time downloads
343
Public
Parameters
271M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
From the Hugging Face model README
This is MiaoshouAI/Florence-2-base-PromptGen-v1.5 which retains its existing features, but with changes to supporting configuration and code to ensure drop-in replacement for Microsoft Florence-2 Model Base the when using Transformers library in Python.
Florence-2-base-PromptGen is a model trained by MiaoshouAI that specializes in generating highly descriptive prompts and tags that assist with training image generation models like FLUX.1-dev and creating descriptive prompts for image generation.
Supported prompts include standard prompts from Florence2-base such as <0D> for identifying object locations and enhanced prompts by MiaoshouAI including <CAPTION>, <DETAILED_CAPTION>, <MORE_DETAILED_CAPTION> and additional prompts included <GENERATE_TAGS> and <MIXED_CAPTION>. See the original repo for more details.
To use this model, you can load it directly from the Hugging Face Model Hub.
To run it as an API Server, either on Windows or Linux, with command line clients (including fast captioning of all images in folders) you can use Florence2 Vision API Server .
First, install dependancies (in a virtual environent if you prefer), for example:
pip3 install torch torchvision --index-url https://download.pytorch.org/whl/cu124
pip3 install transformers pillow einops timm
The following code is based on the microsoft/Florence2-base example but with updated prompt and model, and correct imports.
import requests
import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForCausalLM
device = "cuda:0" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32
model = AutoModelForCausalLM.from_pretrained("createveai/Florence-2-base-PromptGen-v1.5", torch_dtype=torch_dtype, trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("createveai/Florence-2-base-PromptGen-v1.5", trust_remote_code=True)
# Examples include CAPTION>, <DETAILED_CAPTION>, <MORE_DETAILED_CAPTION>,<GENERATE_TAGS>, <MIXED_CAPTION>, <0D>
prompt = "<CAPTION>"
url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg?download=true"
image = Image.open(requests.get(url, stream=True).raw)
inputs = processor(text=prompt, images=image, return_tensors="pt").to(device, torch_dtype)
generated_ids = model.generate(
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
num_beams=3,
do_sample=False
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]
parsed_answer = processor.post_process_generation(generated_text, task=prompt, image_size=(image.width, image.height))
print(parsed_answer[prompt])