Skip to content

Spatial Reasoning Is the New Frontier in 3D Vision

#3d-vision #spatial-reasoning #scene-representation #multimodal-models #transformers

Give a vision-language model two photos of the same room from different angles, then ask where the sofa sits relative to the window. You'll often get a confident answer that's simply wrong. The model reasoned about a 2D projection because that's what it was trained on, while the question lived in 3D space.

That dimensional mismatch is the target of five recent arXiv papers. They approach it from different directions: mesh correspondence, scene completion, novel-view depth, and VLM spatial reasoning. The shared thread is a refusal to treat geometry as something you can pattern-match from flat pixels.

This cluster marks a transition in 3D vision. The field is moving from perception, naming what's there, to reasoning, knowing where it is and how it connects. Each paper here is a different answer to the same question: how do you get models to think in 3D when they were trained on 2D?

Two camps, one goal ​

The five papers split into two camps. TokenMatch and SPAR3S take the geometry-first route: build explicit 3D structure and derive everything from it. Z3D, FactoSR, and GraFT take the reasoning-first route: keep the model architecture, and give it geometry or geometric constraints through the input, the loss, or the latent space.

The two camps aren't in competition. One says learn the geometry and the rest follows. The other says your model already knows a lot about the world, it just needs geometry handed to it. Both are producing results, and the real action over the next year will be where the two routes merge.

TokenMatch: let curvature decide where tokens go ​

Mesh correspondence is one of those problems that sounds solved until you touch it. Given two 3D shapes, find the matching points: the knee on a standing person to the knee on a squatting one, the nose on a partial scan to the nose on a full model. Non-isometric deformation stretches and compresses surfaces, and partial observations remove most of the evidence. Classical descriptors and template-based methods handle the easy cases, but they degrade quickly.

TokenMatch is a transformer that treats mesh matching as a token-matching problem. Self-attention learns patch-level relations inside each shape, cross-attention learns relations between the pair, and the output is dense point-level correspondence. The core idea is how those patches are chosen. Instead of fixed-size or uniform patches, TokenMatch tokenizes the mesh adaptively, guided by curvature. High-curvature regions carry the most geometric information, so they get finer treatment, and the learned descriptors end up shape-specific rather than generic.

The training setup is the boldest part. TokenMatch trains exclusively on BeCoS, a partial-to-partial dataset built for non-isometric matching, then transfers to full-shape matching without retraining or fine-tuning. It holds up on CP2P, PSMAL, FAUST, SCAPE, and SHREC'19 on mean geodesic error and intersection-over-union, and inference runs in under a second. Sub-second is what makes this usable in a pipeline. You can run it in a loop, score matches, and feed the result to the next stage without building a batch job around it.

SPAR3S and Z3D: generating what no camera saw ​

Two papers go after the harder case: generating what no camera saw. SPAR3S completes scenes from sparse, unconstrained views. The cost problem is the usual one. Dense volumetric representations need enormous compute, and large-scale 3D supervision barely exists. SPAR3S sidesteps both by working in a sparse voxel-aligned latent space, where only occupied voxels are represented.

The supervision trick is what makes it work. The sparse latent space is learned directly from multi-view images through a photometric loss, backpropagated through differentiable 3D Gaussian splatting. No ground-truth 3D required. At inference, observed voxels are encoded from the sparse views, and a masked autoregressive transformer predicts the missing latent tokens plus their spatial support. Modeling occupancy and token values together keeps completed regions spatially consistent. SPAR3S beats prior work on synthetic indoor scenes and transfers to RealEstate10k.

Z3D takes a different route to adjacent territory. It asks what 3D foundation models such as VGGT already know internally. These feed-forward transformers predict rich scene representations, and Z3D's hypothesis is that doing reconstruction well requires encoding general knowledge about 3D scenes, including surfaces no single image reveals. So Z3D runs latent diffusion on the 3DFM representation to decode pointmaps in unseen views. The depth maps it produces for new views are realistic across multiple datasets.

Both papers are bets that explicit 3D structure beats implicit guessing. SPAR3S builds the structure, Z3D extracts it.

Quick take: every method in this cluster gains by constructing 3D geometry explicitly or handing it to a model that wasn't trained to imagine it.

FactoSR: splitting the objective into verifiable pieces ​

FactoSR starts from a sharp complaint: vision-language models are fundamentally flat. Trained to interpret 2D projections, they miss the latent 3D geometry and temporal continuity that spatial reasoning demands. The paper's argument is that you can't fix this by training bigger on the same kind of data. The objective itself has to change.

The change is a decomposition. FactoSR splits world-consistent reasoning into three orthogonal sub-objectives: XY for planar correspondence, Z for depth consistency, T for temporal reversibility. Each is a constraint you can verify, and a unified policy-learning mechanism optimizes all three. That transforms an ill-posed projection-recovery problem into a set of steps the optimizer can actually get credit for.

The results: 5.9% gain on VSI-Bench and 4.5% on All-Angles-Bench. These benchmarks are crowded at the top, so a five-point swing is the difference between a model that reasons about depth and layout and one that pattern-matches appearance.

5.9% on VSI-Bench and 4.5% on All-Angles-Bench from factorizing spatial objectives into XY, Z, and T. 27% CIDEr improvement on ScanQA from a training-free scene graph, with no fine-tuning. Up to 65% on VSI-Bench, a frozen MLLM passing fine-tuned and proprietary baselines. Under a second for mesh correspondence at inference, after training only on partial-to-partial shapes.

GraFT: scene graphs without fine-tuning ​

GraFT makes the opposite bet from FactoSR: keep the training untouched, change the input. The paper's complaint about the standard remedies for spatial reasoning is fair. Fine-tuning on curated datasets costs supervision, and attaching 3D encoders couples you to a backbone. GraFT builds a compact 3D scene graph from the input instead, and feeds three capabilities to a frozen MLLM: deterministic geometry through symbolic tools, allocentric layout through a bird's-eye-view rendering, and visual-attribute grounding through task-relevant egocentric frames.

It works. GraFT improves every metric on ScanQA over the same-backbone baseline, with CIDEr up 27%. On VSI-Bench it improves frozen MLLMs by up to 65%, surpassing every proprietary baseline and several prominent fine-tuned spatial models. No training run, no new encoder, no backbone change.

For cost-constrained deployments, this is the most direct paper in the cluster. The price you pay is maintaining the scene graph. But that's a data engineering problem, not a model training problem.

Choosing between the five ​

PaperTaskMethodSupervision neededHeadline result
TokenMatch3D shape correspondence, partial and fullTransformer with curvature-guided mesh tokenizationTrained on BeCoS onlySub-second inference, transfers to full shapes
Z3DNovel-view depth synthesisLatent diffusion on 3DFM representationsPretrained 3DFM, no task fine-tuningRealistic depth for unseen views across datasets
SPAR3SScene completion from sparse viewsSparse voxel latents via autoregressive transformerMulti-view images, no 3D ground truthHigher novel-view quality, transfers to RealEstate10k
FactoSRSpatial reasoning in VLMsFactorized RL over XY, Z, T constraintsRL with verifiable geometric rewards+5.9% VSI-Bench, +4.5% All-Angles-Bench
GraFTSpatial reasoning in MLLMs3D scene graph plus symbolic tools, frozen modelNone+27% CIDEr on ScanQA, up to +65% VSI-Bench

Sources ​

  • TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation, arXiv 2609.04202
  • Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations, arXiv 2609.04174
  • Sparse Auto-Regressive Modeling for Scene Generation from Multi-View Images, arXiv 2609.03931
  • Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning, arXiv 2609.03729
  • GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs, arXiv 2609.03892

Common pitfalls ​

A few months back I ran a functional-map baseline on a partial-shape benchmark. The inference time alone convinced me to shelve it. Watching a transformer do the same job in under a second is TokenMatch's pitch in miniature, and the broader lesson carries across this cluster: the methods that succeed are the ones that made geometry load-bearing in the input, the loss, or the latent space. But the surrounding practice is full of traps. These are the ones I keep hitting, in myself and in other people's code.

  1. Don't assume partial-to-full transfer is symmetric. TokenMatch transfers from partial training to full inference, which is the direction its design supports. Try to go the other way and descriptor statistics drift. If your failures are partial shapes but your training data is mostly full scans, retrain on partial data or accept the gap.

  2. Don't fine-tune before trying the cheap route. GraFT's 27% CIDEr gain on ScanQA comes from a frozen model plus a scene graph. If your instinct is to collect a curated spatial dataset and fine-tune a backbone, you're spending compute on something the frozen path might already deliver. Probe the cheap route first.

  3. Don't treat decoded depth as measured depth. Z3D produces realistic depth for unseen views, but realistic is not calibrated. Push a 3DFM far outside its training distribution and the decoder leans on priors, producing confident hallucinations. Validate against known geometry before you plan with it.

  4. Don't model the full volume when you only need surfaces. SPAR3S exists because dense volumetric generation is brutally expensive. If your scene generator burns GPU memory at high voxel resolution, the fix is not a bigger machine. Sparse occupancy-aware latents with splatting-based supervision get you the same coverage for a fraction of the memory.

  5. Don't squeeze spatial correctness into one reward. FactoSR separates XY, Z, and T into verifiable sub-objectives because a single scalar reward makes credit assignment nearly impossible. Decompose the objective, or the policy never learns what it did right.

One thing to remember ​

Here's the pattern to keep: geometry is becoming a first-class input and objective, not a byproduct of pixel statistics. TokenMatch puts it in the tokenization, SPAR3S in the latent structure, FactoSR in the reward, GraFT in the context. Every one of these methods beats its flat baseline by a visible margin. Models that treat the world as 3D from the start are the ones that answer spatially grounded questions correctly.

The bottom line ​

If you're shipping spatial question answering or robot planning without a training budget, adopt GraFT's scene-graph pipeline. A frozen backbone plus explicit geometry beats fine-tuned models here, and the cost is a data pipeline, not a training run.

If you're completing scenes from sparse views and your dense generator is the bottleneck, build on SPAR3S. Occupancy-aware sparse latents supervised through Gaussian splatting remove the density constraint and don't need 3D ground truth.

If you're betting on the next foundation model generation, watch the 3DFM representation line. Z3D shows those latents can decode hidden geometry, FactoSR shows factorized geometric rewards push VLMs into reliable spatial reasoning, and the two will likely merge within six months into models that output geometry natively.