Skip to content

Repository files navigation

Plan-Conditioned Imitation for Robust Object Retrieval under Self-Occlusion in Dense Clutter

Kowndinya Boyalakuntla1, Ajinkya Pawar2, Abdeslam Boularias1, Jingjin Yu1
1 Rutgers University
2 Indian Institute of Technology Bombay

Project page · Paper · Dataset

We introduce TRACE (Teacher Rollouts for Adaptive Closed-loop Execution), a plan-conditioned imitation learning framework for object retrieval under self-occlusion in dense clutter. TRACE uses non-prehensile pushes to expose a target object for grasping. A privileged teacher solves the scene once inside a digital twin built from a single RGB-D view; a student then executes on the real robot from partial observations, conditioned on that nominal plan and free to deviate from it.

This repository provides our reference implementation, comparison baselines, hardware pipeline, and scripts for reproducing the simulation results in our paper. The project page includes method walkthroughs and videos of our hardware trials.

trace/sim/         task, policies, digital twin, evaluation protocol, baselines
trace/hardware/    UR5e pipeline: calibration, perception, the four controllers, recording
trace/common/      path configuration shared by both
scripts/           data download and one-command reproduction
docs/              installation, reproduction, hardware build

Quick start (simulation)

git clone https://github.com/Kowndinya2000/trace-code.git && cd trace-code
conda env create -f environment.yml && conda activate trace   # see docs/INSTALL.md for Isaac Gym
python scripts/download_data.py                               # assets, scenes, checkpoints (~336 MB)
python -m trace.common.paths                                  # confirm the layout resolves
python scripts/reproduce_sim.py main --gpu 0                  # main table, ~1 h on one RTX 3090

scripts/reproduce_sim.py writes RESULTS.md next to the raw per-scene records under $TRACE_RUNS, so every reported number can be traced back to the episodes behind it.

The assets, scene sets, checkpoints and training labels live in a companion dataset, https://huggingface.co/datasets/Kowndi/trace. download_data.py fetches them and checks each archive against a recorded SHA-256, so a truncated or altered download fails rather than quietly changing a result; --only collections adds the label sets needed to retrain a student. --list prints the payload table with sizes and checksums. Terms differ per payload and are stated on the dataset page: the grasp-quality network is not trained in this work, and NOTICE records the same attributions for the code.

What the method does

  1. Perceive once. One top-down RGB-D frame is segmented into the eleven blocks, and each silhouette is matched to a known shape to recover its pose.
  2. Solve in a twin. The scene is rebuilt in simulation and a privileged reinforcement-learning teacher, which sees the complete state, pushes until the target becomes graspable. Its end-effector path is the nominal plan.
  3. Execute closed loop. The student sees only what the camera can see past the arm, keeps a recurrent memory of objects it can no longer observe, and is conditioned on a short look-ahead of the nominal plan. It re-perceives before every 4 cm push and may leave the plan.
  4. Grasp. A grasp-quality network scores sixteen orientations; the robot grasps when the target clears the threshold with enough clearance for the jaws.

Execution is bounded by a teacher-relative budget: the plan's length plus twenty decisions, and its nominal travel plus 15 cm. Exhausting either is a failure, not a reset.

Reproducing the paper

Table Command Cost
Plan-conditioned policies and references python scripts/reproduce_sim.py main ~1 GPU-hour
Planning and heuristic baselines python scripts/reproduce_sim.py baselines ~6 GPU-hours
Plan-window and recurrent-memory ablations python scripts/reproduce_sim.py ablations ~1 GPU-hour

These evaluations use the same 511 development scenes, with arm occlusion from the robot's own links, 10% detection dropout, a five-step observation blackout, and the teacher-relative budget. Confidence intervals are tier-stratified scene bootstraps over 2,000 resamples. See docs/REPRODUCTION.md for the expected numbers and per-table runtimes.

Running on a robot

The hardware pipeline is included for anyone with the same setup: a UR5e with a Robotiq 2F-85, one RealSense D455 looking down at the workspace, and an optional third-person webcam for recordings. docs/HARDWARE.md covers the bill of materials, the ChArUco calibration board and procedure, the workspace geometry, how to launch each controller, and how the annotated review videos are produced.

The simulation evaluations, including the baselines and ablations, do not require a robot.

Layout and configuration

trace/common/paths.py configures paths using the repository root and these environment variables:

Variable Meaning Default
TRACE_ROOT repository root directory of this file
TRACE_DATA downloaded assets, scenes, checkpoints $TRACE_ROOT/data
TRACE_RUNS where runs write results $TRACE_ROOT/runs
TRACE_PYTHON interpreter for child processes the current one
TRACE_ROBOT_IP, TRACE_D455_SERIAL, TRACE_D415_SERIAL, TRACE_CALIB hardware only see docs/HARDWARE.md

License and third-party code

The simulation environment builds on Isaac Gym Preview 4 and its IsaacGymEnvs examples; the parallel-search baseline derives from the authors' published implementation, retained under trace/sim/pmbs_baseline with its original attribution. See NOTICE.

Citation

@misc{boyalakuntla2026trace,
  title         = {Plan-Conditioned Imitation for Robust Object Retrieval under Self-Occlusion in Dense Clutter},
  author        = {Kowndinya Boyalakuntla and Ajinkya Pawar and Abdeslam Boularias and Jingjin Yu},
  year          = {2026},
  eprint        = {2609.38857},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.38857}
}

About

Plan-Conditioned Imitation for Robust Object Retrieval under Self-Occlusion in Dense Clutter

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages