Downloads · 30 days
14
19% of all-time downloads
zhou777/LandAI-L1
LandAI-L1 is a image-text-to-text model from zhou777. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
14
19% of all-time downloads
All-time downloads
73
Public
Parameters
8.3B
16.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
[Paper (Under Review)] | [Dataset]
</div>LandAI-L1 is a multimodal large language model designed for verifiable land-use reasoning. Unlike traditional black-box classification models, LandAI-L1 enforces a strict cognitive path: "Visual Indexing、Geometric Localization and Language Reasoning".
By compelling the model to explicitly localize visual evidence (bounding boxes) before drawing semantic conclusions, we achieve state-of-the-art accuracy in land-use classification while significantly mitigating multimodal hallucinations.
This model is built upon the Qwen2.5-VL-7B-Instruct architecture and trained using the GRPO-L1 algorithm.
LandAI-L1 establishes a new benchmark on the independent CN-MSLU test set, outperforming both open-source baselines and commercial models.
| Model | Architecture | Training Samples | Accuracy (%) | Hallucination Resistance |
|---|---|---|---|---|
| LandAI-L1 (Ours) | Qwen2.5-VL-7B | ~20k | 86.41 | High |
| LandAI-L1-Zero (Baseline) | Qwen2.5-VL-7B | ~20k | 72.21 | Low |
| LandGPT | InternVL2 | ~80k | 82.5 (approx) | Low |
| Gemini 2.5 Pro | Closed | N/A | 52.21 | Medium |
Note: Hallucination resistance refers to the model's ability to reject misleading textual priors in favor of visual evidence (Visual-Linguistic Conflict Experiment).
Since LandAI-L1 strictly follows the Qwen2.5-VL architecture, you can load it directly using transformers without custom modeling code.
pip install git+https://github.com/huggingface/transformers
pip install qwen-vl-utils
The model was trained using ms-swift, a lightweight and extensible framework for LLM/MLLM fine-tuning.
To reproduce the training or fine-tune on your own geospatial data:
Clone ms-swift: git clone https://github.com/modelscope/swift.git
Prepare your dataset in the standard format.
Run the training ms-swift script.