Abstract
We present Neural Voxel Dynamics, a self-supervised framework for learning 3D latent dynamics from monocular video. While generative video models now produce visually compelling motion, their predominantly 2D representations provide limited geometric structure for modeling and controlling physical interactions. Instead, we learn dynamics in a lifted volumetric latent space: we unproject semantic Video Joint-Embedding Predictive Architecture (V-JEPA) features into a voxel grid using monocular depth priors, producing a geometrically grounded representation that retains rich video features. We then introduce Volumetric Feature Advection, an action-conditioned transition model that predicts the evolution of these latent features directly in 3D.
Unlike hybrid approaches that assume access to privileged physics-engine states or explicit simulation during training and/or inference, we learn only using video-derived supervision and actions, allowing a single latent dynamics model to represent heterogeneous material-dependent phenomena, unifying rigid-body and fluid dynamics. We evaluate across multiple datasets, synthetic and real, and unseen scenarios such as unseen boundary conditions and out-of-distribution generalization, measuring both predictive fidelity and 3D geometric consistency. We further evaluate physical plausibility with Physics-IQ, using videos decoded from predicted latent states and videos generated from predicted optical flow. Our results show that geometrically lifting pretrained video latent representations provides a scalable route toward dynamic world models that learn structured 3D physical dynamics without access to privileged simulator state.
The problem
Physical dynamics unfold in 3D: objects persist through occlusion, motion occurs in a common spatial frame, and interactions depend on geometry and material properties. 2D latent spaces keep those factors entangled with the camera.
The idea
Move the predictive bottleneck out of image space: carry features in a world-frame voxel grid and let one operator advect them through it — physics as state transport rather than pixel morphing.
Why it matters
No material labels and no solver state: the same operator covers rigid bodies, fluids and smoke, and what comes out is a 4D trajectory you can look at from angles the camera never saw.
Contributions
Three claims, in the order the paper makes them:
A persistent, grounded latent world state
Pretrained video features are lifted into 3D, and the state is maintained through time while only partially observed — not re-derived from each frame in isolation.
Volumetric Feature Advection
The transition operator itself, which generatively evolves that state under an applied action — with no material-specific solvers and no privileged information.
One shared model across diverse dynamics
Empirical evaluation showing that a single shared 3D latent transition model captures diverse dynamics, outperforming hybrid and 2D-latent baselines on flow and video prediction.
Method
Pixel observations entangle distinct factors: camera perspective, scene geometry, and material-dependent dynamics. Against 2D generative models on one side and explicit 3D simulators on the other, Neural Voxel Dynamics is a latent middle ground. The design is materially implicit — one operator handles heterogeneous phenomena without any material label; geometrically grounded — the voxel structure makes occupancy and occlusion explicit rather than something to infer from entangled 2D latents; and perspective-independent — anchored in a world frame, so it accepts multiview input and a predicted trajectory can be rendered from new viewpoints without re-running it.
Geometric lifting (points → voxels)
MoGe estimates per-frame depth and camera parameters; a frozen V-JEPA 2.1 encoder produces patch-level features (d = 1408, one latent step per two source frames). Unprojecting them into a world-frame grid — the first camera frame is the world frame — gives a state Zt that carries depth and semantic features and marks every voxel empty, occupied, or occluded. Reusing a strong pretrained backbone this way avoids training a 4D video encoder from scratch.
Implicit feature advection (physics as transport)
Given a history Z1:t and an optional action ut, the advector predicts Ẑt+1 = fθ(Z1:t, ut). Rather than deterministically regressing the next state, fθ is trained by flow matching: it learns a time-dependent velocity field that transports a noisy source distribution toward the target latent manifold. That gives a generative prior over physically plausible transitions and lets the model represent the conditional uncertainty arising from partial observation and unresolved dynamics, without the rigid constraints of a predefined analytical solver.
Differentiable projection
A projection head maps any predicted state back to patch-aligned latents for a target camera. Supervision therefore stays video-derived end to end, and rolling the advector forward autoregressively yields a full 4D latent trajectory that can be read out from a viewpoint chosen after the fact.
Sparse tokenization and the advector
To avoid the cubic cost of dense grids, fθ operates on a sparse set of active voxels and their neighbours: the union of occupied and occluded voxels across the context window, dilated morphologically to leave room for predicted motion. Capacity concentrates on scene-relevant geometry, and cost scales with object complexity rather than with grid resolution.
The advector is a Diffusion Transformer variant whose layers alternate between two attention modes. Local spatial attention restricts each voxel to a k × k × k sparse neighbourhood (5 layers, k = 5), mimicking the step-wise force propagation of classical solvers such as MPM and giving cost linear rather than quadratic in the number of active tokens. Temporal attention regroups tokens by spatial location so each voxel attends causally along its own trajectory. Since one step propagates information only within a local neighbourhood, longer-range transport is composed over stacked layers and successive prediction steps. Adaptive layer normalization modulates the network by the flow-matching time τ, and the output feature velocities are scattered back into the lattice.
Force conditioning and viewpoint
The control is a time-varying action descriptor ut ∈ ℝ12: contact point, applied direction, magnitude, temporal duration, an active-frame indicator, and local / global flags, expressed and normalized in world coordinates to reduce ambiguity. It is injected on two complementary paths — local integration, where the force is projected to the token dimension and added element-wise to each voxel token as a spatially uniform contextual field, and global modulation, where it is encoded into conditioning tokens concatenated to the sequence so self-attention can modulate the scene-level transition.
Crucially, ut is not a privileged simulator state: it specifies no material parameters such as stiffness or viscosity, no engine state, and no per-particle variables. The advector never receives them, and material-dependent behaviour emerges purely from the latent state and the learned transition. Actions are drawn only for the CLEVRER/MuJoCo data; on datasets without such annotations the action reduces to the zero vector and the model predicts implicit dynamics.
Observations come from a single fixed viewpoint, but because the state is anchored in the world frame rather than the camera, the formulation accepts multiview input where it is available. Moving cameras remain out of scope: sparse tokenization scales poorly to them, as it does to unbounded environments and high-frequency deformation.
Experiments
Setup & evaluation protocol
We evaluate across three datasets spanning heterogeneous physics: a CLEVRER-style simulated set with explicit forces (rigid collisions), PhysInOne (fluids with implicit force), and PhysGaia (smoke and fluids). Ours-GT uses ground-truth depth (available for CLEVRER and PhysInOne); Ours-Estimate uses MoGe-estimated depth for all three. Kubric MOVI-C is held out entirely and used only for zero-shot testing.
All methods receive the first frames and generate the following ones. Because our model predicts V-JEPA latents directly, every baseline's generated RGB video is encoded through the same V-JEPA encoder before scoring, so all methods are measured in a common space over 16 frames. Force is provided in each baseline's native interface — 3D force for PhysCtrl/PhysGaussian, force projected to 2D camera space for PhysGen, and as text for CogVideoX.
To avoid favoring models optimized under the same latent objective, we report three non-V-JEPA views of physical correctness alongside latent L2: occupancy (3D geometry), 2D/3D latent flow (motion), and optical flow / PSNR on decoded RGB. A 2D-latent baseline (no 3D projection, same flow-matching objective) isolates the contribution of lifting to 3D.
1 · Latent prediction loss
We first measure prediction accuracy directly in V-JEPA latent space. Each entry reports 2D / 3D latent L2 loss (↓), under single/multi-camera and ground-truth/estimated-depth protocols, and broken down by dynamics category. Our method stays low and stable across material types and camera configurations.
Predicted dynamics in V-JEPA space vs. reference, across heterogeneous materials. Hover a tile and move across it to reveal either side.
The lifted 3D latent can be rendered from novel viewpoints for visualization, shown for a multiview and a single-view input protocol.
| Data | Method | Single-camera | Multi-camera | ||
|---|---|---|---|---|---|
| GT depth | Est. depth | GT depth | Est. depth | ||
| CLEVRER | CogVideoX | 2.98 ± 0.68 / 0.32 ± 0.12 | 2.98 ± 0.68 / 0.56 ± 0.13 | 3.87 ± 0.58 / 1.50 ± 0.11 | 4.14 ± 0.49 / 1.18 ± 0.07 |
| PhysGen | 1.91 ± 0.08 / 0.17 ± 0.02 | 2.97 ± 0.19 / 0.56 ± 0.03 | 3.57 ± 0.17 / 1.89 ± 0.11 | 4.12 ± 0.49 / 1.24 ± 0.19 | |
| PhysGaussian | 4.13 ± 0.08 / 0.56 ± 0.01 | 4.39 ± 0.08 / 0.55 ± 0.03 | 4.34 ± 0.09 / 1.18 ± 0.07 | 4.21 ± 0.09 / 1.14 ± 0.08 | |
| PhysCtrl | 2.96 ± 0.37 / 0.35 ± 0.05 | 2.84 ± 0.23 / 0.57 ± 0.03 | 3.32 ± 0.31 / 1.01 ± 0.07 | 3.21 ± 0.20 / 1.15 ± 0.08 | |
| 2D baseline | 1.39 ± 0.28 / 0.22 ± 0.08 | 1.53 ± 0.71 / 0.42 ± 0.12 | 3.19 ± 0.37 / 1.20 ± 0.27 | 3.69 ± 0.59 / 1.10 ± 0.11 | |
| Ours-GT | 0.98 ± 0.07 / 0.02 ± 0.00 | 1.01 ± 0.10 / 0.39 ± 0.01 | 1.12 ± 0.09 / 0.82 ± 0.05 | 2.51 ± 0.11 / 0.91 ± 0.05 | |
| Ours-Estimate | 1.26 ± 0.13 / 0.25 ± 0.03 | 1.01 ± 0.23 / 0.36 ± 0.04 | 2.50 ± 0.15 / 0.99 ± 0.09 | 2.25 ± 0.09 / 0.90 ± 0.09 | |
| PhysInOne | CogVideoX | 3.02 ± 0.32 / 0.53 ± 0.16 | 3.08 ± 0.33 / 0.75 ± 0.20 | 4.05 ± 0.31 / 1.01 ± 0.11 | 4.00 ± 0.39 / 1.00 ± 0.19 |
| PhysGen | 3.47 ± 0.59 / 0.52 ± 0.15 | 3.61 ± 0.27 / 0.71 ± 0.19 | 3.98 ± 0.37 / 0.87 ± 0.20 | 4.12 ± 0.91 / 1.02 ± 0.49 | |
| PhysGaussian | 3.98 ± 0.22 / 1.02 ± 0.22 | 3.95 ± 0.89 / 0.89 ± 0.20 | 4.57 ± 0.21 / 0.90 ± 0.15 | 4.62 ± 0.23 / 1.04 ± 0.21 | |
| PhysCtrl | 4.24 ± 0.37 / 0.66 ± 0.31 | 4.23 ± 0.51 / 0.81 ± 0.21 | 4.57 ± 0.35 / 0.84 ± 0.17 | 4.34 ± 0.38 / 0.98 ± 0.21 | |
| 2D baseline | 2.99 ± 0.13 / 0.55 ± 0.19 | 2.25 ± 0.23 / 0.82 ± 0.19 | 3.98 ± 0.11 / 0.99 ± 0.12 | 4.14 ± 0.38 / 1.00 ± 0.50 | |
| Ours-GT | 0.43 ± 0.05 / 0.06 ± 0.01 | 3.08 ± 0.17 / 0.63 ± 0.14 | 2.62 ± 0.10 / 0.49 ± 0.07 | 2.62 ± 0.17 / 0.51 ± 0.14 | |
| Ours-Estimate | 3.02 ± 0.17 / 0.53 ± 0.14 | 1.67 ± 0.13 / 0.42 ± 0.14 | 3.51 ± 0.17 / 0.66 ± 0.14 | 2.59 ± 0.21 / 0.49 ± 0.11 | |
| PhysGaia | CogVideoX | N/A | 3.71 ± 0.72 / 0.96 ± 0.32 | N/A | 3.72 ± 0.72 / 1.57 ± 0.47 |
| PhysGen | N/A | 4.89 ± 0.50 / 1.15 ± 0.32 | N/A | 5.23 ± 0.52 / 1.68 ± 0.54 | |
| PhysGaussian | N/A | 3.98 ± 1.16 / 0.96 ± 0.33 | N/A | 3.95 ± 1.21 / 1.58 ± 0.48 | |
| PhysCtrl | N/A | 4.25 ± 1.05 / 0.97 ± 0.32 | N/A | 4.25 ± 1.09 / 1.58 ± 0.48 | |
| 2D baseline | N/A | 2.16 ± 0.92 / 0.95 ± 0.61 | N/A | 4.97 ± 0.99 / 1.47 ± 0.22 | |
| Ours-GT | N/A | 1.84 ± 0.78 / 0.47 ± 0.21 | N/A | 2.10 ± 0.50 / 1.15 ± 0.35 | |
| Ours-Estimate | N/A | 1.66 ± 0.50 / 0.41 ± 0.21 | N/A | 1.99 ± 0.51 / 1.00 ± 0.09 | |
Throughout, bold marks the best result and underline the second best.
| Method | Rigid body | Fluid | Smoke |
|---|---|---|---|
| CogVideoX | 3.39 ± 0.64 / 0.91 ± 0.33 | 4.21 ± 0.31 / 1.15 ± 0.13 | 3.07 ± 0.68 / 1.56 ± 0.48 |
| PhysGen | 2.54 ± 0.35 / 0.59 ± 0.37 | 5.22 ± 1.10 / 1.22 ± 0.39 | 4.37 ± 0.45 / 1.80 ± 0.47 |
| PhysGaussian | 3.32 ± 0.15 / 0.90 ± 0.31 | 4.80 ± 0.20 / 1.18 ± 0.19 | 2.67 ± 1.06 / 1.70 ± 0.47 |
| PhysCtrl | 3.58 ± 0.37 / 0.83 ± 0.33 | 5.03 ± 0.34 / 1.19 ± 0.22 | 3.72 ± 0.96 / 1.73 ± 0.48 |
| 2D baseline | 2.09 ± 0.13 / 0.31 ± 0.21 | 2.99 ± 0.27 / 1.20 ± 0.33 | 2.51 ± 0.53 / 1.62 ± 0.25 |
| Ours-GT | 1.94 ± 0.14 / 0.19 ± 0.11 | 2.12 ± 0.11 / 0.82 ± 0.09 | 2.22 ± 0.33 / 1.15 ± 0.20 |
| Ours-Estimate | 1.97 ± 0.16 / 0.26 ± 0.12 | 2.12 ± 0.21 / 0.80 ± 0.13 | 2.03 ± 0.41 / 1.00 ± 0.29 |
2 · RGB reconstruction (decoder)
A small V-JEPA-to-video decoding head transfers the predicted latent dynamics back to RGB. We report frame-level reconstruction (PSNR / MSE) against reference videos and show decoded clips.
Each decoded clip rendered from a novel viewpoint.
| Method | CLEVRER ↑ | PhysInOne ↑ | PhysGaia ↑ | Kubric ↑ |
|---|---|---|---|---|
| CogVideoX | 16.40 ± 0.53 | 20.46 ± 4.07 | 16.96 ± 7.11 | 17.19 ± 3.00 |
| PhysGen | 24.71 ± 2.30 | 19.17 ± 3.39 | 16.45 ± 8.71 | 16.36 ± 1.48 |
| PhysGaussian | 25.87 ± 1.64 | 19.92 ± 3.52 | 15.99 ± 5.94 | 17.87 ± 3.11 |
| PhysCtrl | 18.21 ± 5.14 | 19.63 ± 2.19 | 16.80 ± 5.95 | 17.12 ± 3.04 |
| 2D baseline | 26.42 ± 1.48 | 17.20 ± 1.36 | 17.02 ± 6.27 | 17.30 ± 1.67 |
| Ours-Estimate (decoder) | 29.83 ± 2.38 | 18.59 ± 1.34 | 18.39 ± 6.78 | 17.89 ± 2.54 |
| Ours-Estimate (Wan-Move) | 22.85 ± 2.62 | 14.14 ± 4.30 | 18.57 ± 3.83 | 15.00 ± 2.00 |
Rendering uses a trainable convolutional mapper followed by a frozen pretrained image VAE decoder; V-JEPA and the VAE stay frozen while the mapper is trained. We report PSNR and omit FVD, as the number of decoded clips is too small for a meaningful distribution-level estimate. Physical plausibility is scored separately with Physics-IQ (Table 7).
3 · Flow estimation
We compare predicted latents against ground-truth trajectories via latent flow consistency (lower is better). Kubric MOVI-C is zero-shot — never used for training. Ours-Estimate is trained with diffusion forcing and history-guided diffusion throughout.
(a) 2D flow
Predicted vs. reference 2D flow for rigid bodies, smoke, and fluid, with arrows overlaid on V-JEPA features. The voxel model is not trained with a flow loss — flow is read out from the learned representation.
The same read-out on real footage with mixed dynamics, shown larger for the dense trajectories.
(a) 2D flow
| Method | CLEVRER | PhysGaia | PhysInOne | Kubric MOVI-C (zero-shot) |
|---|---|---|---|---|
| CogVideoX | 0.492 ± 0.031 | 0.824 ± 0.100 | 0.753 ± 0.065 | 0.667 ± 0.094 |
| PhysGen | 0.315 ± 0.022 | 0.737 ± 0.105 | 0.637 ± 0.110 | 0.738 ± 0.133 |
| PhysGaussian | 0.407 ± 0.007 | 0.682 ± 0.134 | 0.405 ± 0.083 | 0.836 ± 0.126 |
| PhysCtrl | 0.651 ± 0.017 | 0.778 ± 0.127 | 0.694 ± 0.084 | 0.633 ± 0.091 |
| 2D baseline | 0.383 ± 0.020 | 0.534 ± 0.160 | 0.435 ± 0.055 | 0.563 ± 0.109 |
| Ours-Estimate | 0.273 ± 0.009 | 0.464 ± 0.118 | 0.373 ± 0.031 | 0.488 ± 0.151 |
A small convolutional head maps two consecutive latent maps to an optical-flow field, with V-JEPA kept frozen; at inference it needs only the latent sequence, while baseline flows are computed from their generated videos.
| Method | CLEVRER | PhysGaia | PhysInOne | Kubric MOVI-C (zero-shot) |
|---|---|---|---|---|
| CogVideoX | 1.551 ± 0.915 / 17.7 ± 5.3 | 0.996 ± 0.037 / 21.8 ± 6.3 | 1.094 ± 0.113 / 24.0 ± 6.0 | 1.086 ± 0.228 / 23.4 ± 2.7 |
| PhysGen | 0.993 ± 0.057 / 19.2 ± 3.6 | 1.001 ± 0.003 / 50.2 ± 7.5 | 1.072 ± 0.151 / 52.3 ± 12.9 | 1.000 ± 0.061 / 24.1 ± 20.2 |
| PhysGaussian | 0.804 ± 0.338 / 26.4 ± 4.1 | 1.000 ± 0.002 / 37.8 ± 4.0 | 1.068 ± 0.201 / 60.9 ± 13.1 | 0.982 ± 0.076 / 29.7 ± 13.8 |
| PhysCtrl | 10.93 ± 20.97 / 64.5 ± 16.9 | 1.085 ± 0.173 / 41.0 ± 8.5 | 1.295 ± 0.455 / 47.7 ± 13.7 | 1.047 ± 0.121 / 42.3 ± 14.1 |
| 2D baseline | 0.583 ± 0.032 / 21.24 ± 5.2 | 1.055 ± 0.124 / 30.35 ± 9.5 | 1.005 ± 0.314 / 43.33 ± 13.2 | 0.981 ± 0.030 / 29.24 ± 19.9 |
| Ours-Estimate | 0.545 ± 0.186 / 13.5 ± 2.5 | 0.977 ± 0.165 / 29.0 ± 6.0 | 0.980 ± 0.035 / 25.9 ± 8.6 | 0.917 ± 0.059 / 23.1 ± 9.7 |
(b) 3D flow
| Method | CLEVRER | PhysGaia | PhysInOne | Kubric MOVI-C (zero-shot) |
|---|---|---|---|---|
| CogVideoX | 0.619 ± 0.068 | 0.701 ± 0.019 | 0.763 ± 0.068 | 0.735 ± 0.025 |
| PhysGen | 0.683 ± 0.025 | 0.701 ± 0.031 | 0.748 ± 0.070 | 0.721 ± 0.022 |
| PhysGaussian | 0.728 ± 0.006 | 0.723 ± 0.061 | 0.706 ± 0.100 | 0.790 ± 0.072 |
| PhysCtrl | 0.534 ± 0.097 | 0.633 ± 0.091 | 0.711 ± 0.105 | 0.691 ± 0.099 |
| Ours-Estimate | 0.356 ± 0.015 | 0.564 ± 0.045 | 0.537 ± 0.031 | 0.519 ± 0.052 |
4 · Flow → real videos
The estimated 2D flow drives a large video model (i2v + 2D trajectories, Wan-Move) on real footage, and we score the physical plausibility of the generated videos with Physics-IQ.
These eight scenes, plus the two previewed at the top of the page, span liquids, smoke, rigid–fluid impacts, and deformable contact — none of them seen during training. Use the buttons under a clip to switch between our tracks + Wan-Move and the baselines, including Wan-Move run without trajectories; hover a clip to play it.
Physics-IQ is computed over our four domains and a 10% subset of the Physics-IQ test set.
| Method | CLEVRER | Kubric | PhysGaia | PhysInOne | Physics-IQ | Total |
|---|---|---|---|---|---|---|
| CogVideoX | 0.008 | 0.308 | 0.296 | 0.161 | 0.405 | 0.237 |
| PhysGen | 0.073 | 0.126 | 0.031 | 0.194 | 0.191 | 0.163 |
| PhysGaussian (all cameras) | 0.368 | 0.292 | 0.023 | 0.225 | 0.253 | 0.272 |
| PhysCtrl | 0.047 | 0.268 | 0.270 | 0.229 | 0.289 | 0.281 |
| Wan-Move (no tracks) | 0.114 | 0.335 | 0.303 | 0.226 | 0.389 | 0.259 |
| Ours-Estimate | 0.478 | 0.420 | 0.271 | 0.293 | 0.391 | 0.384 |
5 · Occupancy loss
A non-V-JEPA, geometry-level test: predicted 3D occupancy vs. ground-truth occupied voxels. Occupancy IoU is the primary measure (higher is better); the mass ratio is a diagnostic. Ours-Estimate is reported at occupancy thresholds 0.4 / 0.5 / 0.6.
| Method | CLEVRER | PhysGaia | PhysInOne | Kubric MOVI-C (zero-shot) |
|---|---|---|---|---|
| CogVideoX | 0.417 / 1.101 | 0.085 / 0.766 | 0.088 / 0.261 | 0.084 / 0.562 |
| PhysGen | 0.534 / 1.017 | 0.120 / 1.326 | 0.216 / 1.215 | 0.424 / 1.023 |
| PhysGaussian | 0.345 / 0.830 | 0.108 / 0.633 | 0.115 / 0.785 | 0.204 / 0.535 |
| PhysCtrl | 0.148 / 0.685 | 0.113 / 0.687 | 0.100 / 0.701 | 0.129 / 0.698 |
| Ours-Estimate, 0.4 | 0.912 / 1.079 | 0.706 / 1.383 | 0.754 / 1.309 | 0.529 / 1.628 |
| Ours-Estimate, 0.5 | 0.912 / 1.079 | 0.706 / 1.383 | 0.754 / 1.310 | 0.529 / 1.612 |
| Ours-Estimate, 0.6 | 0.912 / 1.079 | 0.706 / 1.382 | 0.754 / 1.310 | 0.529 / 1.595 |
Occupancy IoU is the primary geometric metric; the mass ratio is a diagnostic, not a standalone ranking. A value near one is meaningful mainly for approximately closed, mass-conserving objects — it need not stay constant for phenomena such as smoke, and a method can hit ∼1 while placing the occupied volume incorrectly, so it should be read jointly with IoU.
6 · Force as input
The action is a 12-D per-frame descriptor (contact point, direction, magnitude, duration, and active / local / global flags). Here we compare driving it explicitly, as a supplied vector, against leaving it at zero so the dynamics must be inferred from past frames alone.
(a) Different explicit forces
Same initial scene, driven by different explicit force vectors. Each row shows the input frames on the left and a prediction vs. GT comparison on the right.
(b) Implicit forces
Force inferred from past-frame motion without an explicit force input. As above, the input frames are on the left and a prediction vs. GT comparison on the right.
Ablations
We study the key design choices: voxel resolution, voxel channels, the depth estimator, the projection loss, input views, the occupancy/occlusion losses, and the training objective.
| Grid | Mem. ↓ | IoU ↑ | 2D ↓ | 3D ↓ |
|---|---|---|---|---|
| 10³ | 1.44 GB | 51.74% | 2.75 | 1.60 |
| 15³ | 4.70 GB | 65.14% | 2.63 | 1.55 |
| 20³ | 11.06 GB | 69.79% | 2.02 | 0.99 |
| 25³ | 23.45 GB | 87.13% | 1.78 | 0.61 |
| Voxel state | # ch. | IoU ↑ | 2D ↓ | 3D ↓ |
|---|---|---|---|---|
| Features only | 1408 | N/A | 4.97 | 2.42 |
| + occupancy | 1409 | 67.52% | 2.96 | 1.02 |
| + occupancy + observed | 1410 | 87.13% | 1.78 | 0.61 |
| Estimator | IoU ↑ | 2D ↓ | 3D ↓ |
|---|---|---|---|
| VidDepthAnything + Mast3r | 38.64% | 3.19 | 1.76 |
| VGGT | 61.03% | 2.19 | 0.88 |
| MoGe | 87.13% | 1.78 | 0.61 |
| Proj. loss | Mem. ↓ | IoU ↑ | 2D ↓ | 3D ↓ |
|---|---|---|---|---|
| Without | 100% | 69.47% | 3.24 | 1.75 |
| With | 110% | 87.13% | 1.78 | 0.61 |
| Views | Angle | IoU % ↑ | 2D feat. ↓ | 3D feat. ↓ |
|---|---|---|---|---|
| 1 | 0° | 71.87 / 80.00 | 2.99 / 2.23 | 1.47 / 0.61 |
| 2 | <45° | 71.07 / 87.62 | 2.82 / 2.22 | 1.48 / 0.53 |
| max | >90° | 71.50 / 90.25 | 2.18 / 1.43 | 1.52 / 0.32 |
| Configuration | IoU ↑ | 2D flow ↓ | 3D flow ↓ |
|---|---|---|---|
| Full model | 0.709 | 0.685 | 0.655 |
| w/o occupancy loss | 0.572 | 0.854 | 0.713 |
| w/o occlusion loss | 0.580 | 0.837 | 0.804 |
| w/o both losses | 0.572 | 0.946 | 0.829 |
| Method | IoU ↑ | 2D flow ↓ | 3D flow ↓ |
|---|---|---|---|
| Binary occupancy target | 0.330 | 0.724 | 0.733 |
| Deterministic regression | 0.208 | 0.822 | 0.762 |
| Conditional (DF) | 0.572 | 0.700 | 0.680 |
| Vanilla guidance (DF + HG) | 0.709 | 0.685 | 0.655 |
| Fractional guidance (DF + HG) | 0.660 | 0.687 | 0.694 |
Against an equal-capacity deterministic regression head, generative flow matching is decisively better (Table 15). For longer rollouts we switch the binary occupancy output to a continuous value compatible with flow matching / diffusion forcing, so binary focal / BCE losses no longer apply; vanilla history-guided diffusion is the best configuration and is used for Ours-Estimate.
The method is not tied to any single depth predictor — VGGT works too (slightly worse). Because we predict occupancy rather than hallucinate unseen regions with a generative prior, and a JEPA encoder can be jointly trained to absorb depth cues, geometry can in principle be learned more directly from video; that extension is outside this work's scope.
Failure examples
Typical failure modes. The main one is noise accumulation over long autoregressive rollouts. Scenes far from the training distribution also strain both the decoder and the predictor together, and geometric fidelity is bounded by the V-JEPA encoder resolution and monocular-depth accuracy, so very high-frequency deformation is hard to capture at the current voxel resolution.
Conclusion
We introduced Neural Voxel Dynamics to bridge the gap between 2D video generation and 3D physics simulation. By lifting monocular video into a sparse 3D latent voxel space and learning an action-conditioned implicit feature-advection model, we enable materially-agnostic and geometrically-grounded physical transitions using only passive monocular video. This disentangles camera perspective from material dynamics, avoids manual preprocessing such as segmentation or material estimation, and enables perspective-independent forecasting and the unified simulation of complex, heterogeneous phenomena — rigid collisions and fluid flow alike — without relying on explicit physics-engine states.
Limitations & future work
- Geometric fidelity is bounded by the V-JEPA encoder resolution and monocular depth accuracy; estimation artifacts can propagate into the voxel grid.
- Sparse tokenization scales poorly to unbounded environments, moving cameras, and high-frequency deformation; long rollouts accumulate noise, mitigable via larger training chunks or by feeding refined depth and V-JEPA latents back into the model.
- Explicit actions are annotated only for the CLEVRER/MuJoCo data; elsewhere the action is zero and the dynamics must be inferred, so scaling the implicit-force path to diverse real-world video will likely require a larger model and training corpus.
- The model does not generatively complete fully unseen regions; the control interface is a parameterized force vector.
- Future work: constructing full voxel volumes from image input, a universal V-JEPA decoder for photorealistic video, continuous 3D representations beyond discrete voxels, and dense multi-point articulated (robotic) actions.
Broader impact
Grounding generative video in Euclidean space has strong positive potential for physical planning, autonomous driving, and robotics, where structural and physical accuracy matter for safety, and could accelerate safe offline reinforcement learning. As with any high-fidelity video generation, increased realism raises misuse risk (misinformation, deepfakes), motivating parallel investment in forensics and watermarking.