Python

WorldCrafter

[Arxiv 2026] WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

T

TencentARC

Dernière activité 28 sept. 2026
TencentARC/WorldCrafter

414

étoiles

12

forks

2

issues ouvertes

video-generationworld-models

Ce README est souvent en anglais.

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Wangbo Yu1*, Kunhao Liu1*, Wenbo Hu1†, Shenghai Yuan2, Chaoran Feng2, Haiyang Zhou2
Yukun Huang1, Yiran Wang1, Wang Zhao1, Yingmin Luo1, Ying Shan1

1ARC Lab, Tencent IEG   2Peking University

arXiv Paper   Project Page   YouTube Video   Hugging Face Demo

🤗 If you find WorldCrafter useful, please consider giving this repo a ⭐. Your support helps us share and improve the project. Thank you!

🔆 Introduction

WorldCrafter enables consistent, camera-controlled scene exploration from an image or text prompt. Its camera-queryable implicit 3D-aware memory preserves scene information across viewpoints and over long horizons.

We provide WorldCrafter-Base and WorldCrafter-Fast, a distilled model for faster inference.

🎮 Our interactive demo code and serving infrastructure are fully open source, enabling the community to run, customize, and build on WorldCrafter. See Interactive Demo to get started.

WorldCrafter.mp4

⚙️ Setup

1. Clone WorldCrafter

git clone https://github.com/TencentARC/WorldCrafter.git
cd WorldCrafter

2. Environment

Set up the environment with uv or conda + pip. Both methods use Python 3.11 on Linux and require an NVIDIA GPU with a compatible driver.

A: uv (recommended)

Install uv, then run from the repository root:

# Ubuntu / Debian
sudo apt-get update
sudo apt-get install -y ffmpeg

uv sync --project uvenv --frozen --extra demo
source uvenv/.venv/bin/activate

For other Linux distributions, install FFmpeg using your system package manager.

This installs the locked PyTorch 2.10 / CUDA 12.8 environment and its acceleration dependencies.

B: conda + pip

Create an environment and install PyTorch for your machine. For CUDA 12.8:

conda create -n worldcrafter -c conda-forge python=3.11 pip ffmpeg -y
conda activate worldcrafter
python -m pip install torch==2.10.0 torchvision==0.25.0 \
  --index-url https://download.pytorch.org/whl/cu128
python -m pip install -e ".[demo,xformers]" flash-attn-3==3.0.0 \
  --extra-index-url https://download.pytorch.org/whl/cu128

Choose the appropriate CUDA build from the PyTorch installation commands.

3. Model weights

Models Download Link Notes
WorldCrafter-Base 🤗 Hugging Face Base model
WorldCrafter-Fast 🤗 Hugging Face Distilled high- and low-noise models for faster inference

Download weights with the Hugging Face CLI:

hf download TencentARC/WorldCrafter-Fast --local-dir weights/WorldCrafter-Fast

# Optional: also download Base to run the base model
hf download TencentARC/WorldCrafter-Base --local-dir weights/WorldCrafter-Base

# Optional: generate prompts automatically from input images
hf download Qwen/Qwen3-VL-4B-Instruct --local-dir weights/Qwen3-VL-4B-Instruct

Base model uses shared components from WorldCrafter-Fast, so keep both folders when using base model.

💫 Inference

See the inference guide for camera controls, prompt writing, examples and custom inputs.

1. Image-to-video

Run with the Base or distilled Fast model:

# Base
python inference.py --model-type base --mode i2v \
  --image-path test/I2V/00_cat_robot_vacuum/image.png \
  --prompt test/I2V/00_cat_robot_vacuum/prompt.txt \
  --camera-path test/I2V/00_cat_robot_vacuum/camera.npy

# Fast
python inference.py --model-type fast --mode i2v \
  --image-path test/I2V/03_waterfall/image.png \
  --prompt test/I2V/03_waterfall/prompt.txt \
  --camera-path test/I2V/03_waterfall/camera.npy

--prompt accepts text or a .txt file. For a custom input image, use --prompt auto-first-person for a first-person view scene description or --prompt auto-third-person for a third-person view scene description. This uses Qwen3-VL-4B-Instruct to automatically write the prompt. For example:

# Base
python inference.py --model-type base --mode i2v \
  --image-path test/I2V/00_cat_robot_vacuum/image.png \
  --prompt auto-third-person \
  --camera-path test/I2V/00_cat_robot_vacuum/camera.npy

# Fast
python inference.py --model-type fast --mode i2v \
  --image-path test/I2V/03_waterfall/image.png \
  --prompt auto-first-person \
  --camera-path test/I2V/03_waterfall/camera.npy

2. Text-to-video

# Base
python inference.py --model-type base --mode t2v \
  --prompt test/T2V/00_red_balloon/prompt.txt \
  --camera-path test/T2V/00_red_balloon/camera.npy

# Fast
python inference.py --model-type fast --mode t2v \
  --prompt test/T2V/02_tokyo_street/prompt.txt \
  --camera-path test/T2V/02_tokyo_street/camera.npy

Compilation is off by default. Add --enable-compile to enable it; the first run takes longer to start.

🎮 Interactive Demo

Explore a scene from an image with keyboard camera controls, powered by WorldCrafter-Fast. Run from the repository root with your environment activated:

python -m demo --model-path weights/WorldCrafter-Fast

Open http://localhost:8080 in your browser. Compilation is enabled by default, so the first generation takes longer. Add --devices 0,1 to run on two GPUs.

See the demo guide for camera controls, automatic prompt generation, and deployment options.

📝 Citation

If you find WorldCrafter useful in your research, please cite:

@article{yu2026worldcrafter,
  title={WorldCrafter: Consistent Video World Model with Implicit {3D}-aware Memory},
  author={Yu, Wangbo and Liu, Kunhao and Hu, Wenbo and Yuan, Shenghai and Feng, Chaoran and Zhou, Haiyang and Huang, Yukun and Wang, Yiran and Zhao, Wang and Luo, Yingmin and Shan, Ying},
  journal={arXiv preprint arXiv:2609.24984},
  year={2026}
}

📄 License

See LICENSE.txt for the terms of use and third-party attributions.

Helios, LagerNVS, DreamX-World, EVOKE, HY-WorldPlay, Lyra 2.0, Echo-WM, LingBot-World 2, Matrix-Game 3.5, SANA-WM.

Projets similaires

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

Pythonworld-model
TTencentARC
440 étoiles36

Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Pythonvideo-generationworld-modelsworldmodel
HH-EmbodVis
282 étoiles14

🌍 WorldGen - Generate Any 3D Scene in Seconds

Python3d-generation3d-reconstructiongenerative-ai
ZZiYang-xie
2,1 k étoiles201