Appearance
Three video problems, one shared wall
Video MLLMs have a memory problem, and it isn't the context window. Stream a 90-minute lecture into a model and the frames don't fit. Stream a live camera feed and they never stop coming. The standard answer has been to compress history into an external memory bank and retrieve query-relevant clips on demand. That works, but it has a ceiling: retrieved evidence stays outside the model, so every query re-processes the same history, and the model's reasoning never compounds.
Four recent releases attack that ceiling from different sides. LatentStream turns streaming memory into trainable latent state. KnowVis reframes lecture summarization as a cognitive-load problem. NeoMME removes the decoder and vision tower from multimodal encoding. SolarWM-Data publishes camera-conditioned world-model data at scale. Different bottlenecks, same wall: video tasks don't scale gracefully on memory, compute, or data.
LatentStream: retrieve, then internalize
The core move in LatentStream is shifting from store-and-retrieve to retrieve-and-internalize. Instead of pulling historical clips out of a bank and appending them to the prompt as extra visual context, the model retrieves evidence and folds it into a compact, fixed-length latent memory that carries forward. The distinction matters in practice: with a fixed-length latent memory, the token budget stays flat regardless of video length. A 10-minute stream and a 10-hour stream cost the same at query time.
Three components coordinate. Query-agnostic hierarchical streaming memory organizes visual history into short-, mid-, and long-term levels under a fixed budget, with Jenks-guided adaptive consolidation deciding what gets compressed and when. Once a query arrives, hierarchical latent memory evolution gives groups of latent tokens progressively expanding memory receptive fields, so each group retrieves evidence from the scope it covers and internalizes it. Progressive confidence-guided optimization then builds a reward from group-wise predictive entropy, pushing the memory toward increasingly confident reasoning as history accumulates.
The result is state-of-the-art performance on existing online and offline video benchmarks, which is the least interesting part. The interesting part is that memory becomes a learned function of history rather than a lookup table.
Quick take: memory for streaming video should be trained state, not a retrieval index. The store-and-retrieve pattern re-pays the same processing cost on every query. Internalized latent memory pays once.
KnowVis: summarization as a cognition problem
KnowVis looks at the same wall from the consumer side. Video lectures overwhelm novices because they deliver transient information linearly, while learning requires building interconnected knowledge structures. Text-heavy summaries don't fix that mismatch, they just compress the same linear structure. KnowVis instead extracts a concept map from multimodal video, identifies threshold concepts (the ideas that gate understanding of everything after them), builds structured knowledge units, and renders visual narrative summaries.
The dataset is modest: 125 educational videos across 10 disciplines, paired with 1,079 generated visual summaries. That smallness is informative; this is early-stage territory. The human study reported reduced cognitive load and improved retention compared to state-of-the-art baselines. For lecture tools, the takeaway is that summary output should be structured by conceptual relationships, not transcript order. If your summarizer emits dense paragraphs, you've reproduced the original problem, just shorter.
NeoMME: the encoding bottleneck
The third bottleneck is architectural. Most visual document retrievers are assembled from generative VLMs: a pretrained vision tower, a projector, a causal decoder, and a lot of parameters doing work the task never asked for. Retrieval, classification, and token labeling don't generate text. NeoMME's argument is that you shouldn't pay for a text generator when your output is a vector.
NeoMME is a 260M or 800M bidirectional Transformer that processes text tokens and raw 32×32 image patches through the same encoder stack. It's trained from scratch with a masked discrete-diffusion objective, uses sliding-window attention with global attention every sixth layer plus the final layer, and holds a 16,384-token context, enough for two 4K UHD images in a single pass. No separate vision tower, no causal head, no projector. Fine-tuned with a dual dense and late-interaction head for visual document retrieval, both sizes land on the ViDoRe v3 Pareto frontier. The numbers are hard to argue with.
| Model | Params | ViDoRe v3 (nDCG@10) | v2 (@5) | v1 (@5) |
|---|---|---|---|---|
| ColModernVBERT | 250M | 0.261 | 0.407 | 0.806 |
| NeoMME-260M | 260M | 0.523 | 0.522 | 0.860 |
| ColSmol-500M | 500M | 0.340 | 0.455 | 0.825 |
| NeoMME-800M | 800M | 0.556 | 0.559 | 0.874 |
| Vultron Flash | 850M | 0.565 | 0.604 | 0.882 |
| ColPali v1.3 | 2.92B | 0.430 | 0.547 | 0.848 |
| ColQwen2.5 | 3.75B | 0.524 | 0.601 | 0.895 |
Look at the 260M row. It sits within 0.002 nDCG@10 of ColQwen2.5, a model 14 times larger, on ViDoRe v3. On v1 it beats ColPali v1.3 outright with a fraction of the parameters. Throughput is just as concrete: at matched 2048×2048 input on one NVIDIA L40S, the 260M model encodes about 51 pages per second, roughly double ColModernVBERT. Indexing 100,000 pages takes about 33 minutes of single-GPU time. Re-embed the corpus whenever the data changes, the GPU bill won't punish you.
The practical pain I've hit with late-interaction retrieval is storage. A 2048×2048 page produces about 4,200 vectors, roughly 2.1 MB in float32. Across a big corpus that's terabytes of index. NeoMME's answer combines hierarchical token pooling with asymmetric quantization: cluster document vectors and replace each cluster with its mean, quantize the stored side to int8 or binary, keep query embeddings at higher precision. Storage drops from about 1.5 MB to 6 kB per page, a 255× reduction, while retaining more than 95% of baseline nDCG@10. A million-page corpus goes from 1.5 TB of index to 6 GB. That's the difference between a plausible deployment and a fantasy.
If your task doesn't generate text, a bidirectional encoder is the right tool. Encoder-only multimodal models close the gap to VLMs on retrieval at a fraction of the cost.
SolarWM-Data: the data foundation
The quietest bottleneck is data. World models need video paired with camera trajectories and intrinsics so they can learn geometry and motion, not just pixels. That data has been locked inside companies or scattered across incompatible formats. SolarWM-Data publishes a 14-source raw corpus of 1,425,694 samples: 471,708 kept at high quality, 404,545 at extra-high quality, and 549,441 explicitly rejected with reasons attached. Releasing the rejected shards matters; negative mining is a real training signal, and most datasets just throw those frames away.
The design choice that matters is separating qualification from mixture construction. Every sample keeps its camera, motion, quality, scene, and VLM measurements along with reject reasons, so you can change thresholds, tier policies, and sampling weights without re-running video decoding, camera estimation, VMAF, UniMatch, or DOVER. In most pipelines, filtering is baked into preprocessing, so every mixture change re-runs the expensive chain. This design cuts that loop. Preencoded latents are being published per backend, Wan 2.2 at 81, 153, and 957 frames, MiniMax-H3 at 158, LTX-2.5 at 153 and 953, each record carrying camera trajectories and intrinsics ready for the model reader. You can train without running an encoder locally. Local setup is a config change:
yaml
data:
index_root: /path/to/SolarWM-Data/releases-v1
transport:
kind: local
root: /path/to/SolarWM-Data/releases-v1The main repository shows 5,632 downloads in the past month. For a dataset release with this much infrastructure attached, that's a strong signal that the community was waiting for exactly this.
Key numbers 1.4M total samples, with 549k explicitly rejected but released for negative mining 51 pages per second for NeoMME-260M at 2048×2048 on an L40S 255× index compression, from 1.5 MB to 6 kB per page, at >95% nDCG@10 16,384-token context, enough for two 4K UHD images per forward pass
Common pitfalls
The biggest mistake I've seen in streaming video pipelines is treating retrieved evidence as external context instead of trainable state. When a system pulls relevant clips from a memory bank and appends them to the prompt, the model re-processes that evidence on every query, and reasoning never compounds. The fix is to internalize retrieval into compact latent memory the way LatentStream does.
The second mistake is using a causal VLM for tasks that don't generate text. Document retrieval, classification, and token labeling don't autoregress, but teams default to a 3-4B generative model because it's familiar. NeoMME-260M matches a 3.75B VLM within 0.002 nDCG@10 on ViDoRe v3 with 14× fewer parameters and double the throughput. If you're not generating, don't pay for a generator.
Third, lecture summarizers that output linear text haven't solved anything. Novice overload comes from sequential presentation colliding with networked understanding. A summary that preserves linearity just makes the same content shorter. Structure output around concept maps and threshold concepts.
Fourth, don't bake filtering decisions into preprocessing. If your training pipeline re-runs VMAF, scene detection, and camera estimation every time you tweak a quality threshold, you're burning compute on identical work. SolarWM-Data's separation of quality measurements from mixture construction is the template.
Fifth, watch your late-interaction index size. Multi-vector embeddings at 4,200 vectors per high-res page in fp32 will quietly consume terabytes. Pooling and asymmetric quantization keep most of the quality at a fraction of the footprint. Measure index size before you commit to a storage plan.
One thing to remember: the bottleneck stack for video MLLMs is memory, compute, and data. Each of these four projects targets exactly one layer. They don't compete; they slot together.
The bottom line for video reasoning stacks
- If you're building streaming video understanding, adopt retrieve-and-internalize memory over store-and-retrieve. Fixed-length latent memory keeps your token budget flat as video grows, and reasoning compounds instead of re-processing the same evidence.
- If you're building visual retrieval without generation, drop the causal VLM. NeoMME-260M lands within 0.002 nDCG@10 of ColQwen2.5 on ViDoRe v3 with 14× fewer parameters, and 51 pages per second means a 100k-page corpus indexes in about half an hour on one L40S.
- If you're training a world model, watch SolarWM-Data's latent payload releases. Camera-conditioned data with precomputed Wan, LTX, and MiniMax latents turns months of preprocessing into a download, and that changes the cost of entry for open world-model work within the next couple of quarters.