Downloads · 30 days
0
boudiafA/AgriScope
AgriScope is a image-text-to-text model from boudiafA. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for sam2. The card lists the license as apache-2.0.
<p align="center" <strongPixel-Grounded Multimodal Understanding for Agriculture Images</strong </p
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2026
Repo size
7.7 GB
Likes
0
Public
Click a slice to open those files.
.gz7.7 GB · 100%
From the Hugging Face model README
Release status: AgriGround train/test annotations, task definitions, project figures, and reported manuscript results are available here. The AgriScope project repository is available on GitHub. Model weights, configuration files, processors, and model training/inference code are coming soon.
AgriScope is a pixel-grounded multimodal model for agricultural image understanding. It supports image-level, region-level, and pixel-level interaction within one framework, connecting generated agricultural concepts to segmentation masks and spatial annotations.
The model is designed to identify, describe, localize, and segment agricultural entities including plant diseases, lesions, pests, crops, weeds, botanical species, plant organs, and structural features. It supports both text-only responses and responses grounded in masks or normalized bounding boxes.
AgriScope is trained with AgriGround, a large-scale pixel-grounded agricultural instruction-tuning dataset containing 503,919 images and 11,421,148 samples across 14 tasks.
| Property | Description |
|---|---|
| Model type | Pixel-grounded multimodal language model |
| Domain | Agriculture and plant imagery |
| Languages | English |
| Parameter count | 2B total parameters, as reported in the manuscript |
| Visual representation | Biological semantic and dense spatial encoders |
| Grounding | Language-conditioned [SEG] states with a SAM2-driven decoder |
| Adaptation | Projection alignment followed by LoRA instruction tuning |
| License | Apache-2.0 for repository materials and future code release |
| Model weights | Coming soon |
AgriScope combines complementary semantic and spatial pathways:
[SEG] tokens for grounded concepts.[SEG] hidden state conditions the SAM2-driven mask decoder to produce a corresponding pixel-level mask.The manuscript implementation initializes its visual and grounding components from BioCLIP, DINOv3, and SAM2. The large pretrained encoders remain frozen while alignment modules and LoRA parameters are optimized.
AgriGround is produced using a four-stage annotation and task-generation pipeline:
| Split | Images | Samples | Average samples/image |
|---|---|---|---|
| Train | 401,234 | 9,095,320 | 22.66 |
| Test | 102,685 | 2,325,828 | 22.66 |
| Total | 503,919 | 11,421,148 | 22.66 |
| Source group | Images | Share |
|---|---|---|
| Classification datasets | 232,923 | 46.22% |
| Detection datasets | 27,938 | 5.54% |
| iNatAg subset | 169,324 | 33.60% |
| Insects (IP102) | 73,734 | 14.63% |
| Family | Tasks |
|---|---|
| Captioning | Image-level captioning, region-level captioning, grounded caption generation |
| Segmentation | Referring expression segmentation, semantic segmentation, part segmentation |
| Detection and localization | Phrase grounding, grounded counting, grounded detection, reasoning detection |
| Conversation and QA | Region-level conversation, multi-turn grounded conversation, classification QA, negative absence QA |
See docs/TASKS.md for definitions and per-task sample counts.
The train/test annotations are available in annotations/, with file-level checksums and counts in annotations/manifest.json. Source agricultural images are not redistributed; image_path identifies the corresponding image in its source dataset, whose original license and terms remain applicable.
Training follows two stages described in the manuscript:
The joint objective combines autoregressive language modeling with binary cross-entropy and Dice losses for segmentation supervision. Full hyperparameters, preprocessing, and reproducibility scripts will be released with the code.
The following results are reported in the current manuscript draft.
| Task | Metrics | AgriScope |
|---|---|---|
| Image-level captioning | CIDEr / ASF | 146.4 / 86.7 |
| Region-level captioning | CIDEr / ASF | 132.5 / 84.8 |
| Classification QA | Accuracy / F1 | 82.4 / 80.7 |
| Grounded counting | Accuracy | 74.8 |
| Semantic segmentation | mIoU / Dice | 66.1 / 78.4 |
| Referring expression segmentation | J&F / cIoU | 67.30 / 72.65 |
| Grounded caption generation | METEOR / CIDEr | 27.9 / 118.6 |
| Grounded caption generation | AP50 / mIoU / Recall | 63.9 / 59.4 / 74.2 |
| Parameters | GFLOPs | GPU memory | Inference time |
|---|---|---|---|
| 2B | 177 | 6 GB | 480 ms/image |
Complete baseline comparisons, cross-dataset evaluation, ablations, and experimental settings will accompany the paper release.
.
|-- README.md
|-- CITATION.cff
|-- LICENSE
|-- annotations/
| |-- README.md
| |-- manifest.json
| |-- validation.json
| |-- train/
| `-- test/
|-- docs/
| `-- TASKS.md
`-- images/
|-- overview.png
|-- architecture.png
|-- annotation_pipeline.png
|-- task_examples_v3.png
|-- qualitative_gcg.png
|-- qualitative_tasks.png
`-- referring_segmentation_comparison.png
| Resource | Status |
|---|---|
| Model weights and configuration | Coming soon |
| Processor and inference example | Coming soon |
| Training and evaluation code | Coming soon |
| Project repository | Available on GitHub |
| AgriGround train/test annotations | Available |
| Paper and final citation | Coming soon |
Development updates and future code releases are tracked in the AgriScope GitHub repository.
The final paper link and citation will be added upon release. Until then, please use:
@misc{boudiaf2026agriscope,
title = {AgriScope: Pixel-Grounded Multimodal Understanding for Agriculture Images},
author = {Boudiaf, Abderrahmene and Alanssari, Mohamad and Hussain, Irfan and Javed, Sajid},
year = {2026},
note = {Manuscript under review}
}
Repository documentation and code are provided under the Apache License 2.0. The source images are not redistributed and remain subject to their original dataset licenses and terms.
This work was conducted at Khalifa University of Science and Technology, Abu Dhabi, UAE. We acknowledge the creators and maintainers of the agricultural datasets and open-source foundation models that support this research.
For questions and collaborations, use the AgriScope GitHub issue tracker or the Hugging Face Community tab.