Self-supervised 3D world models

Neural Voxel Dynamics: Learning Volumetric Feature
Advection for 3D Physics in V-JEPA Latent Space

Zican Wang1 · Niloy Mitra1,2

1University College London   2Adobe Research

TL;DR — We learn implicit 3D physics by unprojecting V-JEPA video latents into a lifted 3D voxel grid and advecting them with an action-conditioned, flow-matching transition operator. Essentially, we learn a reusable V-JEPA dynamics module.

Flow-guided video generation on unseen real footage. Wan-Move guided by our predicted tracks; use the buttons under a clip to switch to a baseline, or see more in Experiments.

Abstract

We present Neural Voxel Dynamics, a self-supervised framework for learning 3D latent dynamics from monocular video. While generative video models now produce visually compelling motion, their predominantly 2D representations provide limited geometric structure for modeling and controlling physical interactions. Instead, we learn dynamics in a lifted volumetric latent space: we unproject semantic Video Joint-Embedding Predictive Architecture (V-JEPA) features into a voxel grid using monocular depth priors, producing a geometrically grounded representation that retains rich video features. We then introduce Volumetric Feature Advection, an action-conditioned transition model that predicts the evolution of these latent features directly in 3D.

Unlike hybrid approaches that assume access to privileged physics-engine states or explicit simulation during training and/or inference, we learn only using video-derived supervision and actions, allowing a single latent dynamics model to represent heterogeneous material-dependent phenomena, unifying rigid-body and fluid dynamics. We evaluate across multiple datasets, synthetic and real, and unseen scenarios such as unseen boundary conditions and out-of-distribution generalization, measuring both predictive fidelity and 3D geometric consistency. We further evaluate physical plausibility with Physics-IQ, using videos decoded from predicted latent states and videos generated from predicted optical flow. Our results show that geometrically lifting pretrained video latent representations provides a scalable route toward dynamic world models that learn structured 3D physical dynamics without access to privileged simulator state.

The problem

Physical dynamics unfold in 3D: objects persist through occlusion, motion occurs in a common spatial frame, and interactions depend on geometry and material properties. 2D latent spaces keep those factors entangled with the camera.

The idea

Move the predictive bottleneck out of image space: carry features in a world-frame voxel grid and let one operator advect them through it — physics as state transport rather than pixel morphing.

Why it matters

No material labels and no solver state: the same operator covers rigid bodies, fluids and smoke, and what comes out is a 4D trajectory you can look at from angles the camera never saw.

Contributions

Three claims, in the order the paper makes them:

1

A persistent, grounded latent world state

Pretrained video features are lifted into 3D, and the state is maintained through time while only partially observed — not re-derived from each frame in isolation.

2

Volumetric Feature Advection

The transition operator itself, which generatively evolves that state under an applied action — with no material-specific solvers and no privileged information.

3

One shared model across diverse dynamics

Empirical evaluation showing that a single shared 3D latent transition model captures diverse dynamics, outperforming hybrid and 2D-latent baselines on flow and video prediction.

Method

Pixel observations entangle distinct factors: camera perspective, scene geometry, and material-dependent dynamics. Against 2D generative models on one side and explicit 3D simulators on the other, Neural Voxel Dynamics is a latent middle ground. The design is materially implicit — one operator handles heterogeneous phenomena without any material label; geometrically grounded — the voxel structure makes occupancy and occlusion explicit rather than something to infer from entangled 2D latents; and perspective-independent — anchored in a world frame, so it accepts multiview input and a predicted trajectory can be rendered from new viewpoints without re-running it.

1

Geometric lifting (points → voxels)

MoGe estimates per-frame depth and camera parameters; a frozen V-JEPA 2.1 encoder produces patch-level features (d = 1408, one latent step per two source frames). Unprojecting them into a world-frame grid — the first camera frame is the world frame — gives a state Zt that carries depth and semantic features and marks every voxel empty, occupied, or occluded. Reusing a strong pretrained backbone this way avoids training a 4D video encoder from scratch.

2

Implicit feature advection (physics as transport)

Given a history Z1:t and an optional action ut, the advector predicts Ẑt+1 = fθ(Z1:t, ut). Rather than deterministically regressing the next state, fθ is trained by flow matching: it learns a time-dependent velocity field that transports a noisy source distribution toward the target latent manifold. That gives a generative prior over physically plausible transitions and lets the model represent the conditional uncertainty arising from partial observation and unresolved dynamics, without the rigid constraints of a predefined analytical solver.

3

Differentiable projection

A projection head maps any predicted state back to patch-aligned latents for a target camera. Supervision therefore stays video-derived end to end, and rolling the advector forward autoregressively yields a full 4D latent trajectory that can be read out from a viewpoint chosen after the fact.

Sparse tokenization and the advector

To avoid the cubic cost of dense grids, fθ operates on a sparse set of active voxels and their neighbours: the union of occupied and occluded voxels across the context window, dilated morphologically to leave room for predicted motion. Capacity concentrates on scene-relevant geometry, and cost scales with object complexity rather than with grid resolution.

The advector is a Diffusion Transformer variant whose layers alternate between two attention modes. Local spatial attention restricts each voxel to a k × k × k sparse neighbourhood (5 layers, k = 5), mimicking the step-wise force propagation of classical solvers such as MPM and giving cost linear rather than quadratic in the number of active tokens. Temporal attention regroups tokens by spatial location so each voxel attends causally along its own trajectory. Since one step propagates information only within a local neighbourhood, longer-range transport is composed over stacked layers and successive prediction steps. Adaptive layer normalization modulates the network by the flow-matching time τ, and the output feature velocities are scattered back into the lattice.

Force conditioning and viewpoint

The control is a time-varying action descriptor ut ∈ ℝ12: contact point, applied direction, magnitude, temporal duration, an active-frame indicator, and local / global flags, expressed and normalized in world coordinates to reduce ambiguity. It is injected on two complementary paths — local integration, where the force is projected to the token dimension and added element-wise to each voxel token as a spatially uniform contextual field, and global modulation, where it is encoded into conditioning tokens concatenated to the sequence so self-attention can modulate the scene-level transition.

Crucially, ut is not a privileged simulator state: it specifies no material parameters such as stiffness or viscosity, no engine state, and no per-particle variables. The advector never receives them, and material-dependent behaviour emerges purely from the latent state and the learned transition. Actions are drawn only for the CLEVRER/MuJoCo data; on datasets without such annotations the action reduces to the zero vector and the model predicts implicit dynamics.

Observations come from a single fixed viewpoint, but because the state is anchored in the world frame rather than the camera, the formulation accepts multiview input where it is available. Moving cameras remain out of scope: sparse tokenization scales poorly to them, as it does to unbounded environments and high-frequency deformation.

Experiments

Setup & evaluation protocol

We evaluate across three datasets spanning heterogeneous physics: a CLEVRER-style simulated set with explicit forces (rigid collisions), PhysInOne (fluids with implicit force), and PhysGaia (smoke and fluids). Ours-GT uses ground-truth depth (available for CLEVRER and PhysInOne); Ours-Estimate uses MoGe-estimated depth for all three. Kubric MOVI-C is held out entirely and used only for zero-shot testing.

All methods receive the first frames and generate the following ones. Because our model predicts V-JEPA latents directly, every baseline's generated RGB video is encoded through the same V-JEPA encoder before scoring, so all methods are measured in a common space over 16 frames. Force is provided in each baseline's native interface — 3D force for PhysCtrl/PhysGaussian, force projected to 2D camera space for PhysGen, and as text for CogVideoX.

To avoid favoring models optimized under the same latent objective, we report three non-V-JEPA views of physical correctness alongside latent L2: occupancy (3D geometry), 2D/3D latent flow (motion), and optical flow / PSNR on decoded RGB. A 2D-latent baseline (no 3D projection, same flow-matching objective) isolates the contribution of lifting to 3D.

1 · Latent prediction loss

We first measure prediction accuracy directly in V-JEPA latent space. Each entry reports 2D / 3D latent L2 loss (↓), under single/multi-camera and ground-truth/estimated-depth protocols, and broken down by dynamics category. Our method stays low and stable across material types and camera configurations.

Predicted dynamics in V-JEPA space vs. reference, across heterogeneous materials. Hover a tile and move across it to reveal either side.

The lifted 3D latent can be rendered from novel viewpoints for visualization, shown for a multiview and a single-view input protocol.

Table 1. Latent prediction across camera–depth protocols. Each entry reports mean ± std for 2D / 3D latent L2 loss (↓). N/A — dataset provides no ground-truth depth. The CLEVRER GT-depth columns use our synthetic dataset (Appendix A).
Data Method Single-camera Multi-camera
GT depthEst. depthGT depthEst. depth
CLEVRER CogVideoX2.98 ± 0.68 / 0.32 ± 0.122.98 ± 0.68 / 0.56 ± 0.133.87 ± 0.58 / 1.50 ± 0.114.14 ± 0.49 / 1.18 ± 0.07
PhysGen1.91 ± 0.08 / 0.17 ± 0.022.97 ± 0.19 / 0.56 ± 0.033.57 ± 0.17 / 1.89 ± 0.114.12 ± 0.49 / 1.24 ± 0.19
PhysGaussian4.13 ± 0.08 / 0.56 ± 0.014.39 ± 0.08 / 0.55 ± 0.034.34 ± 0.09 / 1.18 ± 0.074.21 ± 0.09 / 1.14 ± 0.08
PhysCtrl2.96 ± 0.37 / 0.35 ± 0.052.84 ± 0.23 / 0.57 ± 0.033.32 ± 0.31 / 1.01 ± 0.073.21 ± 0.20 / 1.15 ± 0.08
2D baseline1.39 ± 0.28 / 0.22 ± 0.081.53 ± 0.71 / 0.42 ± 0.123.19 ± 0.37 / 1.20 ± 0.273.69 ± 0.59 / 1.10 ± 0.11
Ours-GT0.98 ± 0.07 / 0.02 ± 0.001.01 ± 0.10 / 0.39 ± 0.011.12 ± 0.09 / 0.82 ± 0.052.51 ± 0.11 / 0.91 ± 0.05
Ours-Estimate1.26 ± 0.13 / 0.25 ± 0.031.01 ± 0.23 / 0.36 ± 0.042.50 ± 0.15 / 0.99 ± 0.092.25 ± 0.09 / 0.90 ± 0.09
PhysInOne CogVideoX3.02 ± 0.32 / 0.53 ± 0.163.08 ± 0.33 / 0.75 ± 0.204.05 ± 0.31 / 1.01 ± 0.114.00 ± 0.39 / 1.00 ± 0.19
PhysGen3.47 ± 0.59 / 0.52 ± 0.153.61 ± 0.27 / 0.71 ± 0.193.98 ± 0.37 / 0.87 ± 0.204.12 ± 0.91 / 1.02 ± 0.49
PhysGaussian3.98 ± 0.22 / 1.02 ± 0.223.95 ± 0.89 / 0.89 ± 0.204.57 ± 0.21 / 0.90 ± 0.154.62 ± 0.23 / 1.04 ± 0.21
PhysCtrl4.24 ± 0.37 / 0.66 ± 0.314.23 ± 0.51 / 0.81 ± 0.214.57 ± 0.35 / 0.84 ± 0.174.34 ± 0.38 / 0.98 ± 0.21
2D baseline2.99 ± 0.13 / 0.55 ± 0.192.25 ± 0.23 / 0.82 ± 0.193.98 ± 0.11 / 0.99 ± 0.124.14 ± 0.38 / 1.00 ± 0.50
Ours-GT0.43 ± 0.05 / 0.06 ± 0.013.08 ± 0.17 / 0.63 ± 0.142.62 ± 0.10 / 0.49 ± 0.072.62 ± 0.17 / 0.51 ± 0.14
Ours-Estimate3.02 ± 0.17 / 0.53 ± 0.141.67 ± 0.13 / 0.42 ± 0.143.51 ± 0.17 / 0.66 ± 0.142.59 ± 0.21 / 0.49 ± 0.11
PhysGaia CogVideoXN/A3.71 ± 0.72 / 0.96 ± 0.32N/A3.72 ± 0.72 / 1.57 ± 0.47
PhysGenN/A4.89 ± 0.50 / 1.15 ± 0.32N/A5.23 ± 0.52 / 1.68 ± 0.54
PhysGaussianN/A3.98 ± 1.16 / 0.96 ± 0.33N/A3.95 ± 1.21 / 1.58 ± 0.48
PhysCtrlN/A4.25 ± 1.05 / 0.97 ± 0.32N/A4.25 ± 1.09 / 1.58 ± 0.48
2D baselineN/A2.16 ± 0.92 / 0.95 ± 0.61N/A4.97 ± 0.99 / 1.47 ± 0.22
Ours-GTN/A1.84 ± 0.78 / 0.47 ± 0.21N/A2.10 ± 0.50 / 1.15 ± 0.35
Ours-EstimateN/A1.66 ± 0.50 / 0.41 ± 0.21N/A1.99 ± 0.51 / 1.00 ± 0.09

Throughout, bold marks the best result and underline the second best.

Table 2. Latent prediction per dynamics category on the synthetic dataset (2D / 3D latent L2, ↓). Little fluctuation across rigid, fluid, and smoke.
MethodRigid bodyFluidSmoke
CogVideoX3.39 ± 0.64 / 0.91 ± 0.334.21 ± 0.31 / 1.15 ± 0.133.07 ± 0.68 / 1.56 ± 0.48
PhysGen2.54 ± 0.35 / 0.59 ± 0.375.22 ± 1.10 / 1.22 ± 0.394.37 ± 0.45 / 1.80 ± 0.47
PhysGaussian3.32 ± 0.15 / 0.90 ± 0.314.80 ± 0.20 / 1.18 ± 0.192.67 ± 1.06 / 1.70 ± 0.47
PhysCtrl3.58 ± 0.37 / 0.83 ± 0.335.03 ± 0.34 / 1.19 ± 0.223.72 ± 0.96 / 1.73 ± 0.48
2D baseline2.09 ± 0.13 / 0.31 ± 0.212.99 ± 0.27 / 1.20 ± 0.332.51 ± 0.53 / 1.62 ± 0.25
Ours-GT1.94 ± 0.14 / 0.19 ± 0.112.12 ± 0.11 / 0.82 ± 0.092.22 ± 0.33 / 1.15 ± 0.20
Ours-Estimate1.97 ± 0.16 / 0.26 ± 0.122.12 ± 0.21 / 0.80 ± 0.132.03 ± 0.41 / 1.00 ± 0.29

2 · RGB reconstruction (decoder)

A small V-JEPA-to-video decoding head transfers the predicted latent dynamics back to RGB. We report frame-level reconstruction (PSNR / MSE) against reference videos and show decoded clips.

Table 3. RGB reconstruction fidelity of decoded predictions, PSNR (dB, ↑; higher ⇒ lower MSE) vs. reference.
MethodCLEVRER ↑PhysInOne ↑PhysGaia ↑Kubric ↑
CogVideoX16.40 ± 0.5320.46 ± 4.0716.96 ± 7.1117.19 ± 3.00
PhysGen24.71 ± 2.3019.17 ± 3.3916.45 ± 8.7116.36 ± 1.48
PhysGaussian25.87 ± 1.6419.92 ± 3.5215.99 ± 5.9417.87 ± 3.11
PhysCtrl18.21 ± 5.1419.63 ± 2.1916.80 ± 5.9517.12 ± 3.04
2D baseline26.42 ± 1.4817.20 ± 1.3617.02 ± 6.2717.30 ± 1.67
Ours-Estimate (decoder)29.83 ± 2.3818.59 ± 1.3418.39 ± 6.7817.89 ± 2.54
Ours-Estimate (Wan-Move)22.85 ± 2.6214.14 ± 4.3018.57 ± 3.8315.00 ± 2.00

Rendering uses a trainable convolutional mapper followed by a frozen pretrained image VAE decoder; V-JEPA and the VAE stay frozen while the mapper is trained. We report PSNR and omit FVD, as the number of decoded clips is too small for a meaningful distribution-level estimate. Physical plausibility is scored separately with Physics-IQ (Table 7).

3 · Flow estimation

We compare predicted latents against ground-truth trajectories via latent flow consistency (lower is better). Kubric MOVI-C is zero-shot — never used for training. Ours-Estimate is trained with diffusion forcing and history-guided diffusion throughout.

(a) 2D flow

Predicted vs. reference 2D flow for rigid bodies, smoke, and fluid, with arrows overlaid on V-JEPA features. The voxel model is not trained with a flow loss — flow is read out from the learned representation.

The same read-out on real footage with mixed dynamics, shown larger for the dense trajectories.

(a) 2D flow

Table 4. 2D latent-flow error (↓), mean ± std. Kubric MOVI-C is zero-shot.
MethodCLEVRERPhysGaiaPhysInOneKubric MOVI-C (zero-shot)
CogVideoX0.492 ± 0.0310.824 ± 0.1000.753 ± 0.0650.667 ± 0.094
PhysGen0.315 ± 0.0220.737 ± 0.1050.637 ± 0.1100.738 ± 0.133
PhysGaussian0.407 ± 0.0070.682 ± 0.1340.405 ± 0.0830.836 ± 0.126
PhysCtrl0.651 ± 0.0170.778 ± 0.1270.694 ± 0.0840.633 ± 0.091
2D baseline0.383 ± 0.0200.534 ± 0.1600.435 ± 0.0550.563 ± 0.109
Ours-Estimate0.273 ± 0.0090.464 ± 0.1180.373 ± 0.0310.488 ± 0.151

A small convolutional head maps two consecutive latent maps to an optical-flow field, with V-JEPA kept frozen; at inference it needs only the latent sequence, while baseline flows are computed from their generated videos.

Table 5. Optical-flow error (↓). Each entry reports displacement error normalized by ground-truth flow magnitude / flow-vector angle difference in degrees; a magnitude above one means less meaningful motion. Kubric MOVI-C is zero-shot.
MethodCLEVRERPhysGaiaPhysInOneKubric MOVI-C (zero-shot)
CogVideoX1.551 ± 0.915 / 17.7 ± 5.30.996 ± 0.037 / 21.8 ± 6.31.094 ± 0.113 / 24.0 ± 6.01.086 ± 0.228 / 23.4 ± 2.7
PhysGen0.993 ± 0.057 / 19.2 ± 3.61.001 ± 0.003 / 50.2 ± 7.51.072 ± 0.151 / 52.3 ± 12.91.000 ± 0.061 / 24.1 ± 20.2
PhysGaussian0.804 ± 0.338 / 26.4 ± 4.11.000 ± 0.002 / 37.8 ± 4.01.068 ± 0.201 / 60.9 ± 13.10.982 ± 0.076 / 29.7 ± 13.8
PhysCtrl10.93 ± 20.97 / 64.5 ± 16.91.085 ± 0.173 / 41.0 ± 8.51.295 ± 0.455 / 47.7 ± 13.71.047 ± 0.121 / 42.3 ± 14.1
2D baseline0.583 ± 0.032 / 21.24 ± 5.21.055 ± 0.124 / 30.35 ± 9.51.005 ± 0.314 / 43.33 ± 13.20.981 ± 0.030 / 29.24 ± 19.9
Ours-Estimate0.545 ± 0.186 / 13.5 ± 2.50.977 ± 0.165 / 29.0 ± 6.00.980 ± 0.035 / 25.9 ± 8.60.917 ± 0.059 / 23.1 ± 9.7

(b) 3D flow

Table 6. 3D latent-flow error (↓), mean ± std. Kubric MOVI-C is zero-shot.
MethodCLEVRERPhysGaiaPhysInOneKubric MOVI-C (zero-shot)
CogVideoX0.619 ± 0.0680.701 ± 0.0190.763 ± 0.0680.735 ± 0.025
PhysGen0.683 ± 0.0250.701 ± 0.0310.748 ± 0.0700.721 ± 0.022
PhysGaussian0.728 ± 0.0060.723 ± 0.0610.706 ± 0.1000.790 ± 0.072
PhysCtrl0.534 ± 0.0970.633 ± 0.0910.711 ± 0.1050.691 ± 0.099
Ours-Estimate0.356 ± 0.0150.564 ± 0.0450.537 ± 0.0310.519 ± 0.052

4 · Flow → real videos

The estimated 2D flow drives a large video model (i2v + 2D trajectories, Wan-Move) on real footage, and we score the physical plausibility of the generated videos with Physics-IQ.

These eight scenes, plus the two previewed at the top of the page, span liquids, smoke, rigid–fluid impacts, and deformable contact — none of them seen during training. Use the buttons under a clip to switch between our tracks + Wan-Move and the baselines, including Wan-Move run without trajectories; hover a clip to play it.

Physics-IQ is computed over our four domains and a 10% subset of the Physics-IQ test set.

Table 7. Physics-IQ benchmark on different video datasets (overall score ↑).
MethodCLEVRERKubricPhysGaiaPhysInOnePhysics-IQTotal
CogVideoX0.0080.3080.2960.1610.4050.237
PhysGen0.0730.1260.0310.1940.1910.163
PhysGaussian (all cameras)0.3680.2920.0230.2250.2530.272
PhysCtrl0.0470.2680.2700.2290.2890.281
Wan-Move (no tracks)0.1140.3350.3030.2260.3890.259
Ours-Estimate0.4780.4200.2710.2930.3910.384

5 · Occupancy loss

A non-V-JEPA, geometry-level test: predicted 3D occupancy vs. ground-truth occupied voxels. Occupancy IoU is the primary measure (higher is better); the mass ratio is a diagnostic. Ours-Estimate is reported at occupancy thresholds 0.4 / 0.5 / 0.6.

Table 8. Occupancy-based evaluation: occupancy IoU (↑) / mass ratio (∼1). Kubric MOVI-C is zero-shot.
MethodCLEVRERPhysGaiaPhysInOneKubric MOVI-C (zero-shot)
CogVideoX0.417 / 1.1010.085 / 0.7660.088 / 0.2610.084 / 0.562
PhysGen0.534 / 1.0170.120 / 1.3260.216 / 1.2150.424 / 1.023
PhysGaussian0.345 / 0.8300.108 / 0.6330.115 / 0.7850.204 / 0.535
PhysCtrl0.148 / 0.6850.113 / 0.6870.100 / 0.7010.129 / 0.698
Ours-Estimate, 0.40.912 / 1.0790.706 / 1.3830.754 / 1.3090.529 / 1.628
Ours-Estimate, 0.50.912 / 1.0790.706 / 1.3830.754 / 1.3100.529 / 1.612
Ours-Estimate, 0.60.912 / 1.0790.706 / 1.3820.754 / 1.3100.529 / 1.595

Occupancy IoU is the primary geometric metric; the mass ratio is a diagnostic, not a standalone ranking. A value near one is meaningful mainly for approximately closed, mass-conserving objects — it need not stay constant for phenomena such as smoke, and a method can hit ∼1 while placing the occupied volume incorrectly, so it should be read jointly with IoU.

6 · Force as input

The action is a 12-D per-frame descriptor (contact point, direction, magnitude, duration, and active / local / global flags). Here we compare driving it explicitly, as a supplied vector, against leaving it at zero so the dynamics must be inferred from past frames alone.

(a) Different explicit forces

Same initial scene, driven by different explicit force vectors. Each row shows the input frames on the left and a prediction vs. GT comparison on the right.

(b) Implicit forces

Force inferred from past-frame motion without an explicit force input. As above, the input frames are on the left and a prediction vs. GT comparison on the right.

Ablations

We study the key design choices: voxel resolution, voxel channels, the depth estimator, the projection loss, input views, the occupancy/occlusion losses, and the training objective.

Table 9. Voxel grid resolution vs. quality/compute. Finer grids improve all metrics at a memory cost.
GridMem. ↓IoU ↑2D ↓3D ↓
10³1.44 GB51.74%2.751.60
15³4.70 GB65.14%2.631.55
20³11.06 GB69.79%2.020.99
25³23.45 GB87.13%1.780.61
Table 10. Voxel state channels (monocular). Occupancy and observed channels are both important for geometry.
Voxel state# ch.IoU ↑2D ↓3D ↓
Features only1408N/A4.972.42
+ occupancy140967.52%2.961.02
+ occupancy + observed141087.13%1.780.61
Table 11. Depth / geometry estimator. Lifting is highly sensitive to geometric quality; MoGe is best.
EstimatorIoU ↑2D ↓3D ↓
VidDepthAnything + Mast3r38.64%3.191.76
VGGT61.03%2.190.88
MoGe87.13%1.780.61
Table 12. Projection loss. The 2D projection loss is a key component of the objective.
Proj. lossMem. ↓IoU ↑2D ↓3D ↓
Without100%69.47%3.241.75
With110%87.13%1.780.61
Table 13. Number of input views and view-angle span at inference (multi-camera / single-camera).
ViewsAngleIoU % ↑2D feat. ↓3D feat. ↓
10°71.87 / 80.002.99 / 2.231.47 / 0.61
2<45°71.07 / 87.622.82 / 2.221.48 / 0.53
max>90°71.50 / 90.252.18 / 1.431.52 / 0.32
Table 14. Occupancy and occlusion losses (5,000 steps). Removing either term degrades occupancy and motion.
ConfigurationIoU ↑2D flow ↓3D flow ↓
Full model0.7090.6850.655
w/o occupancy loss0.5720.8540.713
w/o occlusion loss0.5800.8370.804
w/o both losses0.5720.9460.829
Table 15. Deterministic regression vs. flow-matching variants (5,000 steps, 24-frame). DF: diffusion forcing; HG: history guidance.
MethodIoU ↑2D flow ↓3D flow ↓
Binary occupancy target0.3300.7240.733
Deterministic regression0.2080.8220.762
Conditional (DF)0.5720.7000.680
Vanilla guidance (DF + HG)0.7090.6850.655
Fractional guidance (DF + HG)0.6600.6870.694

Against an equal-capacity deterministic regression head, generative flow matching is decisively better (Table 15). For longer rollouts we switch the binary occupancy output to a continuous value compatible with flow matching / diffusion forcing, so binary focal / BCE losses no longer apply; vanilla history-guided diffusion is the best configuration and is used for Ours-Estimate.

The method is not tied to any single depth predictor — VGGT works too (slightly worse). Because we predict occupancy rather than hallucinate unseen regions with a generative prior, and a JEPA encoder can be jointly trained to absorb depth cues, geometry can in principle be learned more directly from video; that extension is outside this work's scope.

Failure examples

Typical failure modes. The main one is noise accumulation over long autoregressive rollouts. Scenes far from the training distribution also strain both the decoder and the predictor together, and geometric fidelity is bounded by the V-JEPA encoder resolution and monocular-depth accuracy, so very high-frequency deformation is hard to capture at the current voxel resolution.

Conclusion

We introduced Neural Voxel Dynamics to bridge the gap between 2D video generation and 3D physics simulation. By lifting monocular video into a sparse 3D latent voxel space and learning an action-conditioned implicit feature-advection model, we enable materially-agnostic and geometrically-grounded physical transitions using only passive monocular video. This disentangles camera perspective from material dynamics, avoids manual preprocessing such as segmentation or material estimation, and enables perspective-independent forecasting and the unified simulation of complex, heterogeneous phenomena — rigid collisions and fluid flow alike — without relying on explicit physics-engine states.

Limitations & future work

  • Geometric fidelity is bounded by the V-JEPA encoder resolution and monocular depth accuracy; estimation artifacts can propagate into the voxel grid.
  • Sparse tokenization scales poorly to unbounded environments, moving cameras, and high-frequency deformation; long rollouts accumulate noise, mitigable via larger training chunks or by feeding refined depth and V-JEPA latents back into the model.
  • Explicit actions are annotated only for the CLEVRER/MuJoCo data; elsewhere the action is zero and the dynamics must be inferred, so scaling the implicit-force path to diverse real-world video will likely require a larger model and training corpus.
  • The model does not generatively complete fully unseen regions; the control interface is a parameterized force vector.
  • Future work: constructing full voxel volumes from image input, a universal V-JEPA decoder for photorealistic video, continuous 3D representations beyond discrete voxels, and dense multi-point articulated (robotic) actions.

Broader impact

Grounding generative video in Euclidean space has strong positive potential for physical planning, autonomous driving, and robotics, where structural and physical accuracy matter for safety, and could accelerate safe offline reinforcement learning. As with any high-fidelity video generation, increased realism raises misuse risk (misinformation, deepfakes), motivating parallel investment in forensics and watermarking.