CG-World: A Large-Scale World-State Dataset and Protocol for World Models
06:00 · July 30, 2026 · arXiv cs.AI RSS

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.
Summary
CG-World supplies a large-scale world-state dataset and accompanying protocol drawn from industrial computer graphics production pipelines. The release, labeled v1, comprises roughly 850,000 temporally aligned segments, each lasting between one and five seconds. Unlike conventional video or robotics corpora that record only final observations, CG-World stores explicit intermediate representations that include multimodal semantic labels, spatial scene graphs, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. These elements are organized into unified spatiotemporal samples that separate latent states, observations, relations, events, and branch metadata.
A distinctive feature is the branch-lineage mechanism designed to support intervention learning and counterfactual reasoning. Each trajectory is annotated as factual, an observational intervention, an action intervention, a mechanism intervention, or a strict counterfactual branch, with intervention targets, invariant variables, and alternative outcomes recorded explicitly. This structure supplies the structured supervision needed for models that must remain intervenable and physically grounded.
The authors evaluate the dataset on three tasks relevant to world-model research. Geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies all show measurable gains when models are trained with the additional state and branch information. The work positions CG-World as a foundation for continued community expansion toward shared infrastructure supporting general-purpose world models, Physical AI, and embodied intelligence.
Why it matters
High technical depth and novelty in structured supervision for world models; directly actionable for Dutch robotics and AI labs; supports EU synthetic-data and counterfactual research priorities.









