On the InternVLA-N1 start–goal benchmark, WSM uses RGB only and achieves the best SR and SPL in both Home and Commercial scenes among the evaluated methods.
TL;DR
WSM brings the SLAM mechanism itself, rather than only its outputs, into navigation: the generation expert dreams goal-conditioned futures, and the SLAM expert estimates world states from both observed and dreamed frames, maintaining one persistent world state from RGB input alone. Incorporating SLAM into the learning process benefits both training and inference, rather than merely providing its outputs to the model at inference time. WSM thereby closes the perception–action loop within a single world model.
A SLAM expert and a generation expert in one autoregressive model, supporting navigation, localization, and mapping.
We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend refinement of accumulated errors, to maintain a consistent world state during interaction. Given the current observation and a navigation goal, WSM predicts future visual states and jointly estimates their camera motion and dense geometry, grounding visual prediction in an evolving spatial world state. This spatial state is continuously updated as new observations arrive and provides the basis for action generation and closed-loop navigation. WSM is trained end-to-end with a joint navigation–SLAM objective, enabling downstream navigation to benefit directly from SLAM-style state maintenance and refinement while preserving accurate geometric estimation. Experiments demonstrate improved navigation performance together with strong SLAM accuracy, highlighting the potential of SLAM as an intrinsic mechanism for long-horizon world modeling and embodied interaction.
Our model unifies visual dreaming and SLAM within an autoregressive Mixture-of-Transformers (MoT) for start–goal navigation.
All chunks of an episode are trained in one forward pass. No prediction token attends to its own target.
On the InternVLA-N1 start–goal benchmark, WSM uses RGB only and achieves the best SR and SPL in both Home and Commercial scenes among the evaluated methods.
| Method | RGB | Depth | Home | Commercial | ||
|---|---|---|---|---|---|---|
| SR↑ | SPL↑ | SR↑ | SPL↑ | |||
| Reported by LoGoPlanner | ||||||
| DD-PPO | ✓ | ✓ | 0.4 | 0.4 | 5.3 | 5.2 |
| iPlanner | – | ✓ | 43.0 | 40.6 | 54.6 | 52.8 |
| ViPlanner | ✓ | ✓ | 45.0 | 43.2 | 63.7 | 61.9 |
| LoGoPlanner | ✓ | ✓ | 57.3 | 52.4 | 67.1 | 63.9 |
| Our evaluation | ||||||
| LoGoPlanner† | ✓ | ✓ | 59.6 | 53.6 | 65.8 | 62.0 |
| World SLAM Model (Ours) | ✓ | – | 70.3 | 66.3 | 73.9 | 71.8 |
SR and SPL in %. †Re-evaluated with the latest official code.
(a) Navigation result. (b) Observation and dreamed frames at Q1–Q3. (c) Estimated vs. GT trajectory. (d) GT vs. predicted depth.
@article{qin2026wsm,
title = {World SLAM Model: Joint World Modeling for SLAM and Navigation},
author = {Minghui Qin and Yijun Yuan and Weicheng Zheng and Kenan Li and Weibang Wang and
Chang Sun and Junhao Huang and Anmin Liu and Yicheng Yao and Hang Zhao},
journal = {arXiv preprint arXiv:2609.32626},
year = {2026},
url = {https://arxiv.org/abs/2609.32626}
}