Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing vision-language models can inspect local regions with zooming and cropping tools, but most exploration strategies follow either a one-shot focus or a single sequential trajectory. Such single-path exploration can lose global context, leave scattered regions unvisited, and revisit or count the same evidence multiple times.
To this end, we propose GeoVista, a planning-driven active perception framework for UHR remote sensing interpretation. Instead of committing to one zooming path, GeoVista first builds a global exploration plan, then verifies multiple candidate regions through branch-wise local inspection, while maintaining an explicit evidence state for cross-region aggregation and de-duplication.
To enable this behavior, we introduce APE-GRO, a cold-start supervised trajectory corpus that reformulates diverse UHR tasks as Global-Region-Object interactive reasoning processes with a unified, scale-invariant spatial representation. We further design an Observe-Plan-Track mechanism for global observation, adaptive region inspection, and evidence tracking, and align the model with a GRPO-based strategy using step-wise rewards for planning, localization, and final answer correctness. Experiments on RSHR-Bench, XLRS-Bench, and LRS-VQA show that GeoVista achieves state-of-the-art performance.
To provide a cold start for this behavior, we construct APE-GRO (Active Planning and Execution across Global-Region-Object levels), a supervised trajectory dataset for tool-interleaved UHR reasoning. Each trajectory in APE-GRO is represented as a sequence of reasoning states, tool actions, visual observations, evidence updates, and the final answer. APE-GRO is synthesized through a four-stage pipeline: First (Data Collection), we compile an initial foundation of 142,125 raw data instances collected from diverse RS data sources, including high-resolution imagery (GeoLLaVA-8K, FAIR1M, DOTA-v1.5) and low-resolution imagery (AID, UC-Merced, NWPU-RESISC45, WHU-RS19, and EarthVQA). Second (Context Engine), a context engine converts annotations into structured teacher inputs using a nested region checklist, while injecting constraints to hide the final answer, mask object details, and enforce progressive human-like search. Third (Multi-Turn Execution Loop), a teacher VLM iteratively generates structured reasoning, issues zoom_in calls with relative bounding boxes, receives cropped visual feedback, and updates the checklist. Finally (Quality Control), the generated traces are filtered through rejection sampling, leakage sanitization, and structure validation to distill 33,437 high-quality, multi-turn active perception reasoning trajectories.
APE-GRO organizes UHR active perception into a three-level Global-Region-Object hierarchy. At the Global level (L0), the model observes a downsampled global view and produces a coarse exploration plan. At the Region level (L1), it inspects candidate ROIs to verify structural patterns and narrow down task-relevant areas. At the Object level (L2), it performs fine-grained inspection for small targets, object attributes, and precise localization. Furthermore, to support consistent tool invocation across different crop sizes and image resolutions, we adopt a scale-invariant spatial representation using a discrete relative coordinate system (mapped to a 0-1000 space). This representation provides a nominal relative resolution of 0.1% within each ROI and decouples spatial outputs from the absolute resolution of the original image.
GeoVista follows the same action space and spatial syntax as APE-GRO during inference. Given an UHR image and a query, it avoids greedily committing to a single zooming path. Instead, it first constructs a global plan over multiple ROIs, and then verifies these ROIs with local observations while maintaining a structured evidence state. Because all branches correspond to different planned ROIs but share a global plan, we refer to this process as multi-branch exploration. The framework operates through four key steps:
1. Observe (Global Scene Observation): GeoVista first obtains a downsampled global view V0global of the UHR image. This view provides broad spatial coverage and supports reasoning about scene layout, potential target distribution, and task-relevant search regions. It does not resolve all fine-grained details but provides the necessary context for exploration planning.
2. Plan (Multi-ROI Exploration Planning): Conditioned on Q and V0global, the policy generates a structured plan. Each item in the plan records an ROI, a verification goal, and its current status. The plan specifies exactly which regions should be inspected and what evidence should be verified, enabling targeted exploration instead of unstructured zooming.
3. Track (Branch-wise Verification and Evidence Update): For each pending ROI, GeoVista extracts a high-resolution crop from the original image and decodes a branch trajectory. The evidence state E is updated after each branch, recording the inspected ROI, global coordinates, verification status, local observations, detected objects, and unresolved sub-goals. Local crop coordinates are mapped back to the global image frame before updating, which helps reduce redundant exploration and duplicated counting. If a crop remains ambiguous, GeoVista can append a sub-plan and inspect at a finer scale.
4. Evidence Aggregation: After branch-wise verification, GeoVista aggregates the evidence state into a final prediction. This step maps local findings back to their corresponding nodes in the global plan, forming a hierarchical evidence tree with verified observations, rejected hypotheses, and unresolved cases. Ultimately, the final answer is generated from organized spatial evidence rather than an unstructured long reasoning history.
As illustrated in the figure above, GeoVista is trained in two stages.
Stage I: Supervised Fine-Tuning. The first stage performs SFT on APE-GRO to initialize the model with tool syntax, scale-invariant spatial grounding, and cross-scale reasoning patterns. By optimizing the model to imitate interleaved reasoning trajectories, SFT teaches the model to produce parseable outputs encompassing reasoning text, plan items, tool calls, evidence updates, and final answers. This decomposes large-scale visual reasoning into global observation, regional inspection, and object-level verification.
Stage II: GRPO Alignment. The second stage applies GRPO-based alignment to improve long-horizon active perception on verifiable tasks. To comprehensively evaluate the trajectories, the multi-dimensional reward function combines the following score items:
We evaluate GeoVista on three UHR remote sensing benchmarks: XLRS-Bench, RSHR-Bench, and LRS-VQA. Compared against 19 diverse representative baselines, GeoVista achieves the best overall performance, scoring 41.38 on RSHR-Bench, 50.65 on XLRS-Bench, and 27.68 on LRS-VQA. It significantly outperforms the strongest sequential zoom-in baseline (GeoEyes), demonstrating that increasing visual input capacity alone is insufficient without adaptive multi-region evidence acquisition.
To analyze the role of SFT trajectory data before RL alignment, we compared Qwen2.5-VL-7B fine-tuned on APE-GRO against other trajectory datasets like LRS-GRO and UHR-CoZ. The results indicate that trajectory data with limited interaction diversity or overly sequential zoom-in patterns (like UHR-CoZ) can substantially degrade performance across benchmarks. APE-GRO provides a much more stable and transferable initialization, equipping the model with structured tool use, normalized spatial references, and multi-scale Global-Region-Object exploration patterns—creating a robust, RL-ready initialization.
While SFT teaches the model to imitate structured trajectories, GRPO alignment optimizes exploration decisions in unseen UHR scenes. Starting from the APE-GRO SFT checkpoint, GRPO alignment significantly improves XLRS-Bench performance from 40.53 to 50.65. Training dynamics reveal that as training progresses, tool usage and observation length increase, indicating a shift from supervised imitation to active reward-based multi-region inspection and explicit evidence aggregation.
We conduct an ablation study on the multi-dimensional reward design used in GRPO. The state-aware planning reward (rplan) plays the most critical role; removing it causes the largest degradation (dropping XLRS-Bench performance by 17.18 points). Explicit spatial overlap supervision (riou) and final-task correctness (racc) also remain essential for producing accurate inspected regions and aligning active perception with downstream prediction quality.
GeoVista consistently achieves the best results across perception sub-categories, including Counting (41.25), Scene Classification (54.00), Object Spatial Relationship (42.40), and Object Properties (52.47).
The overall tasks include visual perception, reasoning, and multi-turn dialogue. Notably, GeoVista demonstrates a commanding advantage in the Perception dimension, achieving the highest average score of 41.8. Specifically, our method secures the best performance across fine-grained sub-tasks, including Color Detection (68.0), Shape Recognition (50.0), and Object Classification (60.5). Additionally, it maintains highly competitive performance in multi-turn dialogue tasks, such as multi-turn anomaly detection (68.3) and multi-turn object state judgment (78.8).
The sub-tasks evaluate capabilities across Object Counting, Background, Category, Color, Shape, Status, Reasoning, and Rural/Urban classification. Notably, GeoVista achieves optimal performance in Shape Recognition (40.34) and Rural/Urban classification (62.22), yielding the highest overall average accuracy (27.68) across the entire dataset. This demonstrates that the learned active observation policy generalizes well to broader, open-ended high-resolution remote sensing scenarios.
@article{zhu2026geovista,
title={GeoVista: Visually Grounded Active Perception for Vision-Language Understanding of Ultra-High-Resolution Remote Sensing Images},
author={Zhu, Jiashun and Fu, Ronghao and Hu, Jiasen and Huang, Jing and Xing, Nachuan and Yang, Bo},
journal={arXiv preprint arXiv:2605.14475},
year={2026}
}