DiVE: Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes

Suyun Lee1,*, Minsoo Choi1,*, Jihwan Gong1, Unghui Nam1, Minji Bae1, Byonghyo Shim1
1Seoul National University
*Equal contribution
Conference on Robot Learning (CoRL) 2026
DiVE example rollout on a real shelf.

DiVE decides whether to look or to push, and reveals objects that no viewpoint alone can uncover.

Demo video

Abstract

Active exploration in cluttered environments is essential for reconstructing occluded 3D scenes, requiring both observation and physical interaction. Existing methods rely on belief maps that indicate where the scene is unresolved but not how each region can be resolved, estimating the gains of viewing and pushing only through costly lookahead or hand-tuned heuristics.

We propose DiVE (Decomposed Visibility Estimator), a learned model that predicts two decomposed visibility maps: a view-resolvable map and a push-resolvable map, making explicit what each action mode can reveal. By converting these modality-specific predictions into a common visibility-gain scale, DiVE supports a hierarchical decision structure: the robot first chooses between observation and interaction, then selects the specific action through a mode-specific policy.

In simulation, DiVE improves occupancy IoU by 11.3% on extreme-occlusion shelf scenes and runs 13.2× faster per step (19.8× faster in planning) than a state-of-the-art baseline. Deployed zero-shot on a real robot, it recovers 79.2% of initially occluded objects.

Key Contributions

01

Decomposed Visibility Estimator

Our main methodological contribution is the Decomposed Visibility Estimator (DiVE), a learned module that predicts view-resolvable and push-resolvable visibility maps from partial observations. By preserving which action modality can resolve each region, the representation exposes action-relevant structure that is absent from a unified belief or uncertainty map.

02

Efficient planning from the representation

We show how the decomposed maps support efficient planning: modality-specific predictions are converted to a common expected-visibility-change scale for mode selection, and are used by lightweight policies for within-mode candidate ranking. Unlike prior hierarchical view-and-push planners, the view-versus-push signal follows from the proposed representation rather than candidate-wise post-action belief prediction or a hand-tuned balancing constant.

03

Validated in simulation and the real world

We validate our framework in both simulation and real-world cluttered shelf scenes. In simulation, DiVE improves occupancy IoU by 11.3% under extreme occlusion and runs 13.2× faster per step than a state-of-the-art baseline. On a real robot, DiVE transfers zero-shot and recovers 79.2% of initially occluded objects.

Method

A unified belief or uncertainty map says only that a region is unresolved. It does not say whether the robot should move the camera or change the scene. We therefore center the representation on visibility—the quantity that actions directly change—and decompose it by action modality.

Overview of the DiVE framework.

Overview of the proposed framework. (a) DiVE ($f_{\text{DiVE}}$) predicts view-resolvable ($\phi_{\text{view}}$) and push-resolvable ($\phi_{\text{push}}$) visibility maps from the current partial observation. (b) The maps are converted to a common scale of expected visibility change to select the action mode. (c) The corresponding policy ($\pi_{\text{view}}$ or $\pi_{\text{push}}$) selects the best viewpoint or push. (d) The belief map estimator ($f_{\text{scene}}$) fuses the new observation and, after a push, the swept region to update the global scene belief $\phi_{\text{scene}}$.

Decomposed visibility estimation

DiVE is a recursive estimator with a shared backbone and two heads, updating both maps from the current observation and the previous maps:

$(\phi_{\text{view},t},\, \phi_{\text{push},t}) = f_{\text{DiVE}}(\phi_{\text{view},t-1},\, \phi_{\text{push},t-1},\, o_t)$

The two heads share an architecture but are supervised with different targets, reflecting the different mechanisms by which viewpoint motion and pushing change visibility.

  • View-resolvable visibility $\phi_{\text{view}}$ is a visibility state: the empirical frequency with which each voxel is visible across the viewpoints visited since the most recent push.
  • Push-resolvable visibility $\phi_{\text{push}}$ is a visibility change: the expected absolute change in marginal visibility (over the candidate viewpoint set) induced by a randomly sampled candidate push.

Mode and Action Selection

The two maps carry different units, so they cannot be compared directly. Using $\phi_{\text{view},t}(v)$ as a plug-in Bernoulli mean for the visibility of voxel $v$ from one additional viewpoint, the expected absolute update of the running average has a closed form, which puts both modes on the same expected per-voxel visibility change scale:

$\mathcal{G}_{\text{view}} = \sum_{v \in \mathcal{V}} \dfrac{2\,\phi_{\text{view},t}(v)\,(1 - \phi_{\text{view},t}(v))}{|\mathcal{K}_t| + 1} \qquad \mathcal{G}_{\text{push}} = \sum_{v \in \mathcal{V}} \phi_{\text{push},t}(v)$

The planner picks the mode with the larger expected gain, then ranks candidates within that mode with a lightweight policy ($\pi_{\text{view}}$ or $\pi_{\text{push}}$) in a single forward pass—no candidate-wise post-action belief prediction, and no hand-tuned cross-mode threshold.

Results

We evaluate in Isaac Sim (UR5 in front of a cluttered shelf with YCB objects) against CNABU [Marques et al., RSS 2025], a View-Only ablation, and Random. Models are trained only on low- and high-occlusion scenes; extreme occlusion is held out.

Main results, generalization, and efficiency.

Main results, generalization, and efficiency. (a) Reconstruction quality vs. action budget. As occlusion increases, View-Only and Random plateau early while push-enabled methods keep improving. (b) Zero-shot generalization to out-of-distribution object shapes, and unseen categories with lightweight adaptation of the reconstruction module only. (c) Per-step compute over 100 episodes: 19.8× faster planning and 13.2× faster total step time than CNABU.

Where the gain comes from

Action-selection analysis, robustness, and extensibility.

Action-selection analysis, robustness, and extensibility. (a) With the belief estimator fixed and only the selection strategy changed, decomposed-visibility selection beats belief-gain selection—the gap reflects action-selection quality, not the reconstruction module. (b) Varying one action-space parameter at a time changes final IoU by at most 1.75 points. (c) Fixed push schedules and longer warm-ups both hurt, supporting learned gain-based mode selection. (d) The same recipe extends to a grasp-resolvable head (representation-level feasibility only).

Real-world experiments

We transfer the policy zero-shot to a Piper AgileX arm with an end-effector-mounted RGB-D camera, on 10 cluttered shelf scenes built from physical YCB-category objects, with a 20-action budget. A full real episode, step by step, is in the demo video.

Real-world setup and objects.

Real-world setup. (a) Robot and shelf. (b) Physical objects corresponding to the YCB categories used in simulation, differing in shape, material, texture, and sensing noise.

Method Correctly found ↑ Misclassified but found Missed ↓ Occluded-object discovery ↑
View-Only 54.3 ± 10.2 0.0 ± 0.0 45.7 ± 10.2 54.2 ± 35.4
DiVE (ours) 65.5 ± 14.2 5.4 ± 6.9 29.1 ± 15.4 79.2 ± 23.3
Δ +11.2 +5.4 −16.6 +25.0

Zero-shot real-world results over 10 scenes. All values are percentages.

BibTeX

@inproceedings{lee2026dive,
  title     = {Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes},
  author    = {Lee, Suyun and Choi, Minsoo and Gong, Jihwan and Nam, Unghui and Bae, Minji and Shim, Byonghyo},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026},
}