Appearance
The shared bet: diffusion models as geometry engines
Three papers landed on arXiv this month, from three different groups, and they're all making the same move. 4DAnyone, DreamHand, and AvatarDynamizer each take a video diffusion model and turn it into a geometry engine. Not a pixel generator, not a stylization tool. Geometry.
That's a shift from the standard pipeline, which used per-frame regressors plus temporal smoothing. It works when the body is visible, calibrated, and mostly unoccluded. It falls apart when hands leave the frame, when clothing wrinkles need to move with the skeleton, or when you need thirty consistent views of one person from a single casual video.
The generative route looks different. Let a video diffusion model imagine the views you didn't capture, then lift those views into an explicit 3D representation. The hard part is that diffusion models weren't built for reconstruction. They hallucinate. They drift. They don't care about metric accuracy. Each of these papers finds a different way to discipline the generator, and I think that's the right instinct.
Why camera-controlled video diffusion isn't enough
The motivation for 4DAnyone starts with a concrete failure. Camera-controlled video diffusion models can synthesize plausible novel views of a person from a monocular video. One view, two views, fine. But 4D Gaussian Splatting needs tens of target views to reconstruct anything usable, and a phone video gives you one viewpoint. The model has to invent the rest. When you scale past the capacity of a single DiT forward pass, the views have to be split into groups. That's where consistency dies.
The authors call this the bounded-attention-context problem, and they break it into two coupled bottlenecks. On the reference side, conditioning on every previously generated view grows as O(N). Each new view weakens the cross-view appearance guidance, because the attention context is saturated. On the target side, disjoint groups can't exchange information directly. Group one drifts structurally from group two, and the reconstructed person ends up with inconsistent proportions across viewpoints.
I found this framing the clearest explanation I've seen for why novel-view synthesis breaks at scale. Both bottlenecks come from the same root cause: the attention context is a fixed-size pipe, and reconstruction wants to pour an unbounded stream of views through it.
4DAnyone: pack the references, route the targets
4DAnyone attacks both bottlenecks with two complementary designs. Reference Context Packing (RCP) compresses the growing set of reference views into a fixed-length, mixed-resolution context. Instead of appending every new view to the conditioning stack, it packs them, bringing reference-context complexity from O(N) down to O(1). Constant memory, constant attention load, no matter how many views you've generated. That's the difference between synthesizing 10 views and 50 views at the same conditioning cost.
Target Context Routing (TCR) handles the other side. Rather than fixing the target-view groupings for the whole denoising process, TCR rotates them. At high-noise steps, the groupings shuffle so context is shared across all groups while global structure is being decided. At low-noise steps, the groupings stabilize so fine details don't get averaged away. Simple scheduling idea, and it directly targets the structural drift that kills multi-view consistency.
For training, the team built MVGameHuman, a dataset rendered from an in-house game engine, and combined it with light-stage and in-the-wild video datasets. The game-engine data matters because real multi-view captures are expensive and scarce; synthetic data gives you ground-truth geometry for free. On DNA-Rendering and DyMVHumans, 4DAnyone beats prior methods on both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.
Quick Take: video diffusion models are becoming the geometry engine for 4D reconstruction, not just the pixel generator.
DreamHand: the diffusion model as encoder
DreamHand makes a different bet, and it's the most contrarian of the three. Instead of sampling from a video diffusion model, it repurposes the VDM as a deterministic geometry encoder. One forward pass over the clean latent. No multi-step sampling, no stochasticity.
That works because the latent of a video diffusion model is trained to represent the full scene, including content that isn't visible in the current frame. A hand occluded by an object, or gone from the frame entirely, still leaves traces in the latent. A single-frame regressor can't see those traces. A windowed temporal regressor can't either, because the hand is absent for the whole window. But the diffusion latent encodes the scene beyond current observations, and DreamHand learns to read it.
The architecture is straightforward: a Deterministic Clean-Latent Encoder extracts features from the clean latent, and a Bidirectional Spatiotemporal Decoder turns them into continuous bimanual trajectories with metric placement. No external hand detector required. A Ray-Based Camera Solver configuration drops the need for test-time camera intrinsics, which matters for real egocentric footage where intrinsics are often unknown.
The results are the strongest numbers in this batch. Across five egocentric benchmarks, DreamHand sets a new state of the art. On ARCTIC, the occlusion-heavy benchmark, it cuts MPJPE-p by 30%, nearly a third less error on the hardest case. On HOT3D, 40%. Those are big jumps for benchmarks where prior methods move a point or two at a time. And once out-of-sight hands are included in the evaluation, the gains reach 46-61%, because prior methods simply collapse when the hand isn't in frame. I'd quote those out-of-sight numbers in any meeting about robot manipulation data.
The practical implication: egocentric video is the cheapest source of manipulation data for embodied AI, but occlusion has made it hard to use for training. DreamHand is a step toward turning everyday human video into robot training data at scale.
AvatarDynamizer: dynamics live in texture space
The third paper attacks a different problem. Person-agnostic avatar methods can recover a static 3D body from a monocular video, but skeleton-driven animation leaves clothing static. No wrinkles, no fabric motion, no surface dynamics. That's the uncanny valley in full effect. Person-specific methods solve it, but they need expensive multi-view capture for every individual.
AvatarDynamizer embeds surface dynamics in texture space and treats avatar animation as conditional texture generation. An encoder-decoder representation maps pose-dependent dynamics into dynamic texture maps, which are compatible with pre-trained video diffusion models. The dynamic textures are then decoded into 3D Gaussians for multi-view-consistent rendering.
The key insight is that texture space is an easier place to generate dynamics than 3D space. Video diffusion models are already good at appearance-consistent generation. You don't need to teach them 3D consistency from scratch; the Gaussian decoding handles that. You just need the dynamics to stay consistent across views.
The team also collected a large-scale multi-view dataset with long sequences covering diverse skeletal motions and surface dynamics, because existing datasets are too short and too narrow. Under limited dynamic training data, AvatarDynamizer still beats competing generalizable methods on visual fidelity.
Key numbersO(1): reference-context complexity after packing, down from O(N) growth per added view. That's the difference between 10 and 50 target views at the same conditioning cost. Tens of views: what 4DGS reconstruction needs from a single casual video, and where prior video diffusion models start to drift. Five benchmarks: the egocentric datasets where DreamHand comes out on top, including ARCTIC and HOT3D. 46-61%: how much further the error gap grows once out-of-sight hands are included, because prior methods are effectively guessing.
Three bets, one direction
| 4DAnyone | DreamHand | AvatarDynamizer | |
|---|---|---|---|
| Input | Casual monocular video | Egocentric bimanual video | Static avatar + driving video |
| Diffusion role | Multiview video generator | Deterministic geometry encoder | Conditional texture generator |
| Output | 4D Gaussian Splatting | Metric 3D hand trajectories | Animated avatar with dynamic textures |
| Key mechanism | RCP + TCR | Clean-latent encoder + bidirectional decoder | Texture-space dynamics embedding |
| Bottleneck solved | Attention context at tens of views | Occlusion and out-of-sight hands | Static avatars lack surface dynamics |
| Strongest result | Beats prior methods on DNA-Rendering and DyMVHumans | 30-40% MPJPE-p cut, 46-61% with out-of-sight | Wins on visual fidelity under limited data |
Three different bottlenecks, three different output representations, one shared substrate. That's usually how you can tell a technology is becoming infrastructure: everyone starts building on it, even when they disagree about everything else.
Common pitfalls
- Don't split target views into fixed groups and hope for the best. If the model can't attend to all target views at once, the groups will drift structurally. Rotate the groupings during denoising, like TCR does, so global structure is decided before details are locked in.
- Don't sample from a diffusion model when you need geometry. DreamHand's whole point is that a single deterministic forward pass over the clean latent gives occlusion-robust features. Multi-step stochastic sampling is for pixel quality, not metric accuracy. You want the encoder, not the renderer.
- Don't assume temporal regressors handle out-of-sight objects. A windowed model fails when the hand is absent for the entire window. The diffusion latent is the only representation here that carries scene content beyond the current observation. Skip this and you get the 46-61% collapse, not the 30-40% gains.
- Don't animate static avatars with skeleton transforms alone. Rigid skinning moves the body but leaves clothing frozen. The uncanny valley isn't about shape, it's about surface dynamics. If you can't do multi-view capture per person, texture-space dynamics is the cheapest way to get wrinkles that move.
- Don't let reference context grow without bound. Every view appended to the conditioning stack weakens the guidance from earlier views. Pack the references into fixed-length context. O(N) conditioning is a silent killer of multi-view consistency.
What this means for the field
These three papers converge on the same conclusion from different directions: video diffusion models are becoming the common substrate for dynamic 3D reconstruction. The implications are practical. 4DAnyone points toward consumer-grade avatar capture from a phone video, no multi-view rig needed. DreamHand points toward robot manipulation training from everyday egocentric footage, which is a much larger dataset than anything a lab can collect. AvatarDynamizer points toward animating existing avatar assets with believable clothing dynamics, which matters for games, film, and virtual try-on.
The interesting question is which output representation wins. 4DAnyone outputs 4D Gaussian Splatting. DreamHand outputs trajectories. AvatarDynamizer outputs dynamic textures decoded into Gaussians. The field hasn't settled, and it doesn't need to yet. What matters is that the input is now just a video.
The data play is worth watching too. All three papers had to build or collect new datasets: MVGameHuman from a game engine, a large-scale multi-view dynamic dataset for AvatarDynamizer, and the egocentric benchmarks for DreamHand. Game-engine synthetic data is the interesting one. It gives you ground-truth geometry at scale, and it's only going to get more realistic.
One thing to remember
The common thread is that diffusion models are no longer just generators. They're being repurposed as structured representations: a way to hold scene content that isn't visible, a way to enforce multi-view consistency, a way to embed surface dynamics. If you're building a reconstruction pipeline, the question isn't whether to use diffusion. It's which part of the pipeline you're going to make it.
The bottom line
If you're building 4D human avatars from monocular video, adopt the 4DAnyone pipeline, because RCP and TCR are what make tens of consistent views feasible for 4DGS reconstruction.
If you're working on egocentric hand tracking or robot manipulation data, use DreamHand's deterministic clean-latent encoder, because it's the only approach here that survives occlusion and out-of-sight hands, and it cuts MPJPE-p by 30-40% without test-time intrinsics.
If you're animating static avatars and the uncanny valley is your problem, embed surface dynamics in texture space like AvatarDynamizer, because skeleton-driven animation can't produce clothing wrinkles and per-person multi-view capture doesn't scale.
One thing to watch: game-engine synthetic data is becoming the training fuel for all three approaches, so expect in-the-wild generalization to keep improving within the next year as synthetic multi-view datasets get more realistic.