Lossy terrain reconstruction
Dense decoders smooth action-critical boundaries such as stair edges and sparse footholds.
One depth camera. No external localization. A continuous 1.5-km outdoor traversal.
1USTC AGI Institute · 2X-Humanoid · 3HKUST(GZ) · 4HKU · 5ANU · 6SJTU · 7CUHK · 8Tsinghua
Video 1. Continuous 1.5-km outdoor traversal over natural stairs, slopes, grass transitions, and uneven ground.
Humans can traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies can become fragile as local perception and control errors accumulate during continuous deployment. We attribute this long-horizon fragility to two compounding failure modes in standard teacher–student pipelines: dense terrain reconstructors smooth out action-critical spatial details, while pointwise imitation objectives provide no explicit credit assignment for actions that lead to subsequent teacher–student divergence.
We present SOLO, a unified framework that combines a Query Reconstructor with Trajectory-Aware MSE Distillation. SOLO achieves 97.5% mean traversal success on stress-test terrains and completes a continuous 1.5-km outdoor route using only a single chest-mounted depth camera and proprioception.
Perceptive locomotion policies can clear isolated obstacles yet fail during continuous deployment as small perception and control errors compound. SOLO addresses two coupled failure modes in the standard teacher–student pipeline.
Dense decoders smooth action-critical boundaries such as stair edges and sparse footholds.
Matching the teacher only at the current state does not assign credit to actions that cause future trajectory drift.
Query reconstruction and trajectory-aware distillation improve both the geometry passed to the actor and the training signal used to match the privileged teacher.
A Fourier-encoded query for every height-map cell retrieves spatially specific evidence from a shared depth–proprioception token memory.
Next-state teacher–student disagreement enters the PPO reward, allowing GAE to propagate future disagreement penalties to preceding actions.
The same policy is deployed zero-shot at 50 Hz, without motion capture, external odometry, or external mapping.
Stairs, sparse footholds, a gap, descent, and a movable obstacle in a continuous traversal.
A long, uninterrupted ascent over successive flights of fire stairs.
The perceptive locomotion technology developed in this work was further evaluated in timed competition and demonstrated publicly on challenging terrain.
Tiangong Omni robots powered by the locomotion technology used in SOLO won gold and bronze in the 400 m obstacle course, together with silver in the 100 m obstacle course.
At WRC 2026, Tiangong Omni demonstrated perceptive locomotion on stepping stones and stairs in front of a public audience.
Quantitative and qualitative evaluation at the highest curriculum difficulty shows that QR improves both reconstruction fidelity and the traversal outcomes that depend on precise geometry.