World SLAM Model Joint World Modeling for SLAM and Navigation

Minghui Qin1,3*, Yijun Yuan2*†, Weicheng Zheng2, Kenan Li2, Weibang Wang2, Chang Sun2, Junhao Huang2, Anmin Liu2, Yicheng Yao2, Hang Zhao1,2†
1Shanghai Qi Zhi Institute 2IIIS, Tsinghua University 3Shanghai Jiao Tong University
*Equal contribution †Corresponding author

TL;DR

WSM brings the SLAM mechanism itself, rather than only its outputs, into navigation: the generation expert dreams goal-conditioned futures, and the SLAM expert estimates world states from both observed and dreamed frames, maintaining one persistent world state from RGB input alone. Incorporating SLAM into the learning process benefits both training and inference, rather than merely providing its outputs to the model at inference time. WSM thereby closes the perception–action loop within a single world model.

Overview of the World SLAM Model.

A SLAM expert and a generation expert in one autoregressive model, supporting navigation, localization, and mapping.

Abstract

We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend refinement of accumulated errors, to maintain a consistent world state during interaction. Given the current observation and a navigation goal, WSM predicts future visual states and jointly estimates their camera motion and dense geometry, grounding visual prediction in an evolving spatial world state. This spatial state is continuously updated as new observations arrive and provides the basis for action generation and closed-loop navigation. WSM is trained end-to-end with a joint navigation–SLAM objective, enabling downstream navigation to benefit directly from SLAM-style state maintenance and refinement while preserving accurate geometric estimation. Experiments demonstrate improved navigation performance together with strong SLAM accuracy, highlighting the potential of SLAM as an intrinsic mechanism for long-horizon world modeling and embodied interaction.

Framework overview of the World SLAM Model.

Our model unifies visual dreaming and SLAM within an autoregressive Mixture-of-Transformers (MoT) for start–goal navigation.

Training sequence and attention mask.

All chunks of an episode are trained in one forward pass. No prediction token attends to its own target.

On the InternVLA-N1 start–goal benchmark, WSM uses RGB only and achieves the best SR and SPL in both Home and Commercial scenes among the evaluated methods.

Method RGB Depth Home Commercial
SR↑SPL↑ SR↑SPL↑
Reported by LoGoPlanner
DD-PPO✓✓0.40.45.35.2
iPlanner–✓43.040.654.652.8
ViPlanner✓✓45.043.263.761.9
LoGoPlanner✓✓57.352.467.163.9
Our evaluation
LoGoPlanner†✓✓59.653.665.862.0
World SLAM Model (Ours)✓–70.366.373.971.8

SR and SPL in %. †Re-evaluated with the latest official code.

Visualization of navigation results in simulation.

(a) Navigation result. (b) Observation and dreamed frames at Q1–Q3. (c) Estimated vs. GT trajectory. (d) GT vs. predicted depth.

BibTeX

@article{qin2026wsm,
  title   = {World SLAM Model: Joint World Modeling for SLAM and Navigation},
  author  = {Minghui Qin and Yijun Yuan and Weicheng Zheng and Kenan Li and Weibang Wang and
             Chang Sun and Junhao Huang and Anmin Liu and Yicheng Yao and Hang Zhao},
  journal = {arXiv preprint arXiv:2609.32626},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.32626}
}