SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

One depth camera. No external localization. A continuous 1.5-km outdoor traversal.

Pihai Sun1,2,* Gang Han2,* Jingkai Sun2,4,* Jiahao Ma2,5 Zeran Su1,2 Zelin Tao2 Peiran Liu2,3 Shuai Shi2 Wei Cui2 Zifan Wang3 Jialin Yu2 Wen Zhao2 Kangning Yin6 Jiaxu Wang7 Jiahang Cao4 Lingfeng Zhang8 Hao Cheng3 Jian Tang2 Yijie Guo2 Qiang Zhang1,†

1USTC AGI Institute · 2X-Humanoid · 3HKUST(GZ) · 4HKU · 5ANU · 6SJTU · 7CUHK · 8Tsinghua

* Equal contribution · Corresponding author

Video 1. Continuous 1.5-km outdoor traversal over natural stairs, slopes, grass transitions, and uneven ground.

Abstract

Humans can traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies can become fragile as local perception and control errors accumulate during continuous deployment. We attribute this long-horizon fragility to two compounding failure modes in standard teacher–student pipelines: dense terrain reconstructors smooth out action-critical spatial details, while pointwise imitation objectives provide no explicit credit assignment for actions that lead to subsequent teacher–student divergence.

We present SOLO, a unified framework that combines a Query Reconstructor with Trajectory-Aware MSE Distillation. SOLO achieves 97.5% mean traversal success on stress-test terrains and completes a continuous 1.5-km outdoor route using only a single chest-mounted depth camera and proprioception.

Problem Setting

Perceptive locomotion policies can clear isolated obstacles yet fail during continuous deployment as small perception and control errors compound. SOLO addresses two coupled failure modes in the standard teacher–student pipeline.

Lossy terrain reconstruction

Dense decoders smooth action-critical boundaries such as stair edges and sparse footholds.

Myopic pointwise imitation

Matching the teacher only at the current state does not assign credit to actions that cause future trajectory drift.

Method

Query reconstruction and trajectory-aware distillation improve both the geometry passed to the actor and the training signal used to match the privileged teacher.

SOLO training and deployment overview, Query Reconstructor, and trajectory-aware MSE distillation
Figure 1. SOLO framework. The student reconstructs action-critical terrain geometry from depth and proprioception, while trajectory-aware distillation penalizes future policy disagreement.

Query Reconstructor

A Fourier-encoded query for every height-map cell retrieves spatially specific evidence from a shared depth–proprioception token memory.

  • Preserves sharp terrain boundaries
  • Predicts a 16 × 32 height map and base velocity
  • Uses one noisy chest-mounted depth camera

Trajectory-Aware Distillation

Next-state teacher–student disagreement enters the PPO reward, allowing GAE to propagate future disagreement penalties to preceding actions.

Real-World Experiments

The same policy is deployed zero-shot at 50 Hz, without motion capture, external odometry, or external mapping.

Indoor · Mixed terrain

One course, multiple terrain interactions

Stairs, sparse footholds, a gap, descent, and a movable obstacle in a continuous traversal.

Fire stairs · Continuous ascent

Continuous multi-flight stair climbing

A long, uninterrupted ascent over successive flights of fire stairs.

1depth camera
50 Hzpolicy execution
25controlled joints
0external state systems

Competition & Public Demonstrations

The perceptive locomotion technology developed in this work was further evaluated in timed competition and demonstrated publicly on challenging terrain.

World Humanoid Robot Games · Beijing · 2026

Obstacle-course competition

Tiangong Omni robots powered by the locomotion technology used in SOLO won gold and bronze in the 400 m obstacle course, together with silver in the 100 m obstacle course.

  • Gold400 m obstacle course
  • Bronze400 m obstacle course
  • Silver100 m obstacle course

Additional field videos

Open-world traversal, filmed in one take

Autonomous locomotion across complex outdoor terrain · Bilibili ↗

Perceptive locomotion across complex terrain

Real-world obstacle traversal and long-horizon stair ascent · Bilibili ↗

Results

Quantitative and qualitative evaluation at the highest curriculum difficulty shows that QR improves both reconstruction fidelity and the traversal outcomes that depend on precise geometry.

3.3–4.0× lower average height-map L1 error
97.5% mean stress-test traversal success
96% stepping-stone success, up from 0–3%
Qualitative terrain reconstruction comparison
Figure 2a · Action-critical geometry. Per-cell queries preserve stair edges and sparse foothold boundaries. Open the figure for the full-resolution comparison.
Traversal success at the highest terrain difficulty
Figure 2b · Traversal success. SOLO remains reliable on extreme terrains.
Normalized foot stumble and foothold gradient penalties
Figure 2c · Foot placement quality. Cleaner contacts across difficult stair settings.

BibTeX

BibTeX will be added with the public paper release.