WorldWeaveGrowing Persistent Geometric Worlds
for Video Generation

Yifan Huang1 Lifan Jiang1 Qingyue Hao1 Cheng Chen1 Boxi Wu2 Xiaoxue Ren1 Xiaofei He1 Dehai Zhao1,†

1 Zhejiang University2 Daerwen AI

† Corresponding author

WorldWeave overview: terrain completion, hierarchical agent planning, and read-only rendering progressively extend a persistent geometric world.
Grow the world. Preserve its structure. Render new observations.

Abstract

Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.

Video demonstration

Explore. Turn away. Revisit.

Shared geometry across changing observations.

Four video sequences show farmstead, river crossing, town and winter scenes remaining coherent during camera motion and revisitation.
Each row follows a viewing trajectory from left to right. Insets show depth guidance; numbered markers identify the same scene elements across views.

From persistent world to video

Terrain completion · Agent planning · Read-only rendering

The WorldWeave pipeline connects metric terrain generation, hierarchical scene planning and validation, and depth-guided video synthesis.

Quantitative results

Visual quality, memory, camera control and structural consistency.

Main quantitative results
MethodVisual qualityMemory & camera controlStructural consistency
Imaging
quality ↑
Aesthetic
quality ↑
Structural
memory ↑
Camera
compliance ↑
MN-MS ↑MC-GeCo ↓MC-MEt3R ↓MC-GeoCon ↓
SANA-WM0.7390.6190.32855.50.9730.1130.2030.189
Zing0.7630.6850.98955.80.9740.06980.1290.142
SolarWM-5B0.7550.6040.37054.70.9560.07860.1660.217
AlayaWorld0.7840.6411.5156.40.8960.9260.2390.295
EVOKE-Turbo0.7840.7181.1261.60.9610.1050.1820.249
Echo-WM0.7910.6552.2163.50.9680.06040.1350.135
Seedance 2.00.7170.6172.1469.80.9800.1010.1700.307
Wan3.00.7860.7021.6359.10.9380.1460.1840.131
Kling 3.00.7390.5882.0959.30.9780.07450.1550.163
MiniMax-H3 (Base)0.7250.6170.43850.80.9790.06330.1400.178
WorldWeave0.7810.7042.30 (+426.7%)72.3 (+42.2%)0.9810.05730.1220.121 (-32.2%)

The first six baselines are open-source world models; the remaining baselines are video models. Bold and underline indicate best and second-best scores. Highlighted gains are relative to MiniMax-H3 (Base). Evaluation covers 120 tasks, with structural memory measured on the common 48 applicable tasks.

Download CSVView JSON