CueNav

Seeing what matters

Visual cue guided video planning for generalizable robot navigation.

Hojin Lee1, Sizhe Lester Li2, Maximilian Hilger1, Susie Lu2, Achim J. Lilienthal1, Vincent Sitzmann2, Daniel A. Duecker1

1TUM, MIRMI  ·  2MIT CSAIL

How do people drive robots?

Try it yourself.

“Navigate through the maze and reach the red box.”

Track
Bird's-eye map
90.0 s contacts 0/3
observation

Ready when you are

W A S D drive  ·  R F camera pitch  ·  M map  ·  Esc abort

Intro

Just as people rely on visual context when driving, a video planner can benefit from task and embodiment context expressed through visual observations. We illustrate this with a bird’s-eye view map serving as global context for navigation, and a partial view of the robot body as embodiment context for precise, embodiment-aware control.

Left: a driver's view, with the dashboard map marked as task context and the car's hood as embodiment context. Right: the same two cues in a robot's onboard view — an inset bird's-eye map, and the robot's own body visible at the bottom of frame — shown for a quadruped and for a wheeled robot.

Overview

Given recent visual observations with embodiment or global visual context and a text prompt, the video planner predicts short-horizon future observations. The resulting flow fields from the predicted video are mapped to continuous robot actions by the IDM, followed by closed-loop replanning from new observations.

Acknowledgment

This work was supported by the Bavarian State Ministry of Science and the Arts (StMWK) through the MIT–TUM Collaboration (grant no. 151223031). Vincent Sitzmann and Sizhe Lester Li were supported by the National Science Foundation under Grant No. EEC 2330040 and the CAREER program under 2543631. The authors gratefully acknowledge funding and computational resources provided by AMD through the MIT AI hardware program and the AMD University Program’s AI & HPC Cluster.

BibTeX

@article{lee2026cuenav,
  title   = {Seeing What Matters: Visual Cue Guided Video Planning
             for Generalizable Robot Navigation},
  author  = {Lee, Hojin and Li, Sizhe Lester and Hilger, Maximilian and
             Lu, Susie and Lilienthal, Achim J. and Sitzmann, Vincent and
             Duecker, Daniel A.},
  journal = {arXiv preprint arXiv:2609.16737},
  year    = {2026}
}