A World Action Model enabling local real-time deployment with 85ms low latency.
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.
conda create -n gigaworld-policy python=3.11 -y
conda activate gigaworld-policypython -m pip install --upgrade pip setuptools wheel
python -m pip install \
torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 \
--index-url https://download.pytorch.org/whl/cu126
python -m pip install -r requirements.txtpython -m pip install --no-deps -e ./third_party/giga-train
python -m pip install --no-deps -e ./third_party/giga-models
python -m pip install --no-deps -e ./third_party/giga-datasetsBefore training, compute the normalization statistics and pre-compute the T5 text embeddings for your LeRobot v3.0 dataset.
python scripts/compute_wam_task_norm.py \
--data-root <lerobot_v3_dataset_root> \
--output <norm_stats_json> \
--model-dim 32 \
--action-horizon 48python scripts/compute_t5_embedding.py \
--root <lerobot_v3_dataset_root> \
--wan_path <wan_t5_model_root> \
--device cuda \
--t5_folder_name t5_embeddingDownload the open-sourced pretrained transformer weights from Hugging Face:
# Hugging Face CLI
huggingface-cli download open-gigaai/Giga-World-Policy-0.5 --local-dir ./Giga-World-Policy-0.5
# or Git LFS
git lfs install
git clone https://huggingface.co/open-gigaai/Giga-World-Policy-0.5from huggingface_hub import snapshot_download
snapshot_download(
repo_id="open-gigaai/Giga-World-Policy-0.5",
local_dir="./Giga-World-Policy-0.5",
)Set the dataset, normalization, WAN base model, and initialization checkpoint paths before launching post-training.
export GWP_AGILEX_DATA_PATHS=<lerobot_v3_dataset_root>
export GWP_NORM_STATS_PATH=<norm_stats_json>
export GWP_WAN_PRETRAINED=<wan_diffusers_model_or_path>
export GWP_T5_LOAD_FROM=dir
export GWP_T5_EMBEDDING_DIR=<lerobot_v3_dataset_root>/t5_embedding
export GWP05_TRANSFORMER_PRETRAINED=<gwp05_init_transformer_checkpoint>
export GWP05_PROJECT_DIR=<output_dir_for_gwp05>
python scripts/train.py --config giga_world_policy_0_5_agilex_finetuneexport GWP_AGILEX_DATA_PATHS=<lerobot_v3_dataset_root>
export GWP_NORM_STATS_PATH=<norm_stats_json>
export GWP_WAN_PRETRAINED=<wan_diffusers_model_or_path>
export GWP_T5_LOAD_FROM=dir
export GWP_T5_EMBEDDING_DIR=<lerobot_v3_dataset_root>/t5_embedding
export GWP0_TRANSFORMER_PRETRAINED=<gwp0_init_transformer_checkpoint>
export GWP0_PROJECT_DIR=<output_dir_for_gwp0>
python scripts/train.py --config giga_world_policy_0_agilex_finetuneOpen-loop inference uses a server/client split. Start the server in one terminal, then run the client in another terminal.
export CHECKPOINT=<trained_transformer_checkpoint>
export NORM_STATS=<norm_stats_json>
export DATA_PATH=<lerobot_v3_dataset_root>
export BASE_MODEL=<wan_diffusers_model_or_path>Terminal 1:
cd scripts
./run_inference_openloop.sh serverTerminal 2:
cd scripts
./run_inference_openloop.sh clientTerminal 1:
cd scripts
./run_inference_openloop_gwp0.sh serverTerminal 2:
cd scripts
./run_inference_openloop_gwp0.sh client⚡ C++ Inference Runtime (wam.cpp)
For local real-time deployment, we also provide wam.cpp — a C++/CUDA inference runtime for World-Action Models. It converts GWP-0.5 checkpoints to GGUF, runs action prediction via a pinned llama.cpp/ggml backend, and keeps server-side compatibility with the open-loop client in this repository.
@article{gigaworld-policy-0.5,
title={GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch},
author={Team, GigaWorld and Ye, Angen and Ma, Angyuan and Wang, Boyuan and Ni, Chaojun and Ye, Fangzheng and Huang, Guan and Li, Guo and Zhao, Guosheng and Yan, Haodong and others},
journal={arXiv preprint arXiv:2607.13960},
year={2026}
}