Downloads · 30 days
42
19% of all-time downloads
AdaReasoner/AdaReasoner-VSP-7B
AdaReasoner-VSP-7B is a image-text-to-text model from AdaReasoner. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div align="center" <img src="logo.png" alt="Logo" width="300" <h1 align="center"Dynamic Tool Orchestration for Iterative Visual Reasoning</h1
Downloads · 30 days
42
19% of all-time downloads
All-time downloads
223
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
AdaReasoner-7B is a vision-language model trained with dynamic tool orchestration capabilities for iterative visual reasoning. This model is AdaReasoner-7B-Non-Randomized.
We provide three variants of AdaReasoner-7B, each optimized for different use cases:
| Model | Description | Hugging Face |
|---|---|---|
| AdaReasoner-7B-Randomized | Trained with the adaptive learning method, enabling strong generalization to unseen tools and tasks. Designed for open-ended and evolving tool environments where adaptability is required. | 🤗 Link |
| AdaReasoner-7B-Non-Randomized | Trained without adaptive learning, providing more stable and reliable performance on known tools and tasks, but limited generalization to unseen tools or task settings. | 🤗 Link |
| AdaReasoner-VSP-7B | Task-specialized model trained exclusively on the Visual Spatial Planning (VSP) task, achieving strong performance on VSP benchmarks but not intended for cross-task generalization. | 🤗 Link |
Key Differences:
AdaReasoner-7B can be deployed for single-turn inference using standard inference frameworks such as vLLM. However, AdaReasoner is a tool-planning model whose full capabilities require interaction with an external tool environment. To fully evaluate or utilize its tool-planning behavior, we recommend using AdaEval provided in our repository for batch inference and evaluation, or trying the Demo interface for interactive, single-instance GUI-based reasoning.
The model supports a diverse set of visual reasoning tasks, covering both structured reasoning and open-ended visual understanding:
For full tool-augmented inference capabilities, please refer to the AdaReasoner repository which includes:
Please refer to our paper for detailed benchmark results across multiple visual reasoning tasks.
If you use this model in your research, please cite:
@article{adareasoner2024,
title={Dynamic Tool Orchestration for Iterative Visual Reasoning},
author={AdaReasoner Team},
journal={arXiv preprint arXiv:XXXX.XXXXX},
year={2024}
}
Apache 2.0
This model is part of the AdaReasoner project. For more information, visit our GitHub repository.
For questions and feedback, please open an issue in our GitHub repository.