Downloads · 30 days
20
37% of all-time downloads
AI-Anon/MINGLE-1.0
MINGLE-1.0 is a machine learning model from AI-Anon. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains the fine-tuned vision-language model used in our AAAI submission "MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes." The model builds upon Qwen2-VL-7B with LoRA (rank=64)…
Downloads · 30 days
20
37% of all-time downloads
All-time downloads
54
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
This repository contains the fine-tuned vision-language model used in our AAAI submission "MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes." The model builds upon Qwen2-VL-7B with LoRA (rank=64), and is specifically trained to classify social affiliation between pairs of individuals based on visual and depth cues.
For data files (CSV) and example scripts, please visit the companion GitHub repository:
👉 https://github.com/ResearchDataset/AAAI2025-Dataset
Given a pair of individuals in a street-view scene, the model determines whether they are socially interacting. This task requires reasoning over:
Unlike traditional object detection, this task targets semantically complex visual regions defined by interpersonal interaction—not discrete physical entities.
'Yes', 'No', or 'Not sure' (text-based classification)The model is used with structured prompts combining bounding box and depth values.
Example User Prompt:
In the first image, two individuals are highlighted with bounding boxes (each given as [x1, y1, x2, y2]):
Box1: <|box_start|> 135,117,194,263 <|box_end|>
Box2: <|box_start|> 286,103,328,220 <|box_end|>
In the second image, which is the depth view of the same scene, the same two individuals are highlighted again:
Box1’: <|box_start|> 135,117,194,263 <|box_end|>
Box2’: <|box_start|> 286,103,328,220 <|box_end|>
Note: Box1 and Box1’ represent the same person, and Box2 and Box2’ represent the same person.
Depth Values (from 0-255, where 0 means far and 255 means close; these values are critical for determining proximity): Box1 and Box1’ = 179, Box2 and Box2’ = 142, so the Depth Difference is 37
Carefully analyze both images by considering all visual cues. In particular, pay attention to the following cues:
• Body orientation
• Facial expressions
• Gestures
• Depth distance
• Relative positioning
Based on these details, determine whether the individuals are actively interacting (e.g., engaged in conversation or displaying clear interactive behavior) or if they are merely near each other without meaningful interaction.
Your output must contain exactly one choice: ‘Yes’, ‘No’, or ‘Not sure’ with no additional commentary.
On a manually validated test set:
See full details in the accompanying paper.
This repository includes:
config.json, generation_config.json, etc.)To cite the accompanying paper, please refer to the official publication link once available.
This model was fine-tuned using the open-source framework lmms-finetune.
This model is released for research purposes only. It is not intended for surveillance or any harmful use case. The annotations and model outputs may reflect biases present in the data sources.