Appearance
The prompt is no longer the workflow
Type a prompt, click generate, inspect the result, adjust the prompt. That loop defined diffusion for most people since Stable Diffusion 1.4. The current wave of projects doesn't fit the loop. A CRT TV installation where you turn a physical antenna to control how much noise remains in the generated images and sound. An editing method that breaks a reference image into per-object attributes instead of one global style. Browser Spaces that stack multiple LoRAs into a single FLUX.2 generation. And an open-source 11B video model trained on a $200K budget that sits 0.69% behind OpenAI's Sora on VBench.
Four projects, one direction: the prompt box is no longer the interface. Control is. Whether that control is a physical knob, attribute-level conditioning, adapter composition, or a motion-score flag in a video pipeline, the materials differ but the goal is the same. Each project points at a different problem in the field, and together they define what the next generation of diffusion tooling looks like.
Diffusion TV: denoising you can feel
Diffusion TV is an interactive art installation built from a modified CRT television, described in arXiv 2609.05404. Audiences physically turn the TV's antenna, and the rotation controls the clarity of AI-generated images and sound. More signal, sharper image. Less signal, noisier image. The metaphor is exact: the audience is enacting the denoising loop that diffusion models run internally. The tuning knob switches between three channels of AI-generated animals, Past (extinct species), Present (endangered species), and Future (speculative creatures), which sets the interaction inside an ecological and temporal narrative.
The paper frames the work as an embodied form of explainable AI, and that's fair, but the practitioner read is sharper. Standard diffusion UIs expose the denoising trajectory as a progress bar. Diffusion TV turns the trajectory into the artifact. The curdled intermediate states, the moment a field of noise locks into a recognizable shape, that's the part of generation that makes people react. Most products hide it; this installation puts it in your hands.
Organizing the channels along extinction timelines is the clever part. Past, Present, Future makes the model's training data visible as history. What the model knows about a dodo, versus a pangolin, versus a speculative future animal, exposes the distribution of the corpus without a single line of technical explanation. Steal that pattern: let users experience the model's boundaries instead of reading about them.
RefDiT: local control over reference attributes
RefDiT (arXiv 2609.04976) attacks a failure mode that anyone who has run personalization models has hit. Give the model a reference image with a single subject, and it works: global guidance captures everything it needs. Give it a real-world scene with multiple objects, each with distinct attributes, and global guidance breaks down. The model can't localize which elements of the reference correspond to which elements of the generated image.
The cause is structural. Current methods use a single identifier token to absorb all details from the reference, which rules out attribute-level control from the start. RefDiT decomposes that token at the attribute level, constructs an attribute-aware conditioning signal from the reference image, adjusts the inference prompt context, and trains LoRA blocks on a DiT-based backbone. The correspondence between identifier tokens and local regions is learned explicitly. That's what makes it local guidance instead of global style transfer.
The framework takes three inputs: the reference image, a text prompt, and an optional user-provided guidance context. That third input matters, because it's how you express which aspect of the reference should drive the generation without relying on the model to guess.
For product builders, this is the difference between editing and pastiche. "Make the jacket in this scene look like the jacket in the reference" requires local steering. "Make the whole image feel like the reference" is what global guidance gives you, and most mainstream editors claiming reference-based editing are still implementing the second one.
Quick take: the control problem in diffusion is shifting from what the model understands to how precisely you can point at the part you care about.
The browser LoRA workbench
Between the research papers sits a tooling wave visible in Hugging Face Spaces. Two entries in this cluster show the pattern: Omni-Image-Editor at 2.54k likes, and FLUX.2-Klein-Multi-LoRA at 532.
| Space | Likes | Runtime tier | What it is |
|---|---|---|---|
| Omni-Image-Editor | 2.54k | CPU, GPU upgrade offered | browser-based image editing |
| FLUX.2-Klein-Multi-LoRA | 532 | Zero GPU | multi-LoRA composition on FLUX.2 Klein |
The telling detail is the runtime tier. Omni-Image-Editor defaults to CPU, which tells you who it's for: people who want a working editor in a browser tab without renting a GPU. The GPU upgrade exists for when the CPU ceiling becomes obvious. FLUX.2-Klein-Multi-LoRA pins to the Zero tier, because stacking several LoRAs onto a FLUX.2 model gets heavy fast.
The community pattern I keep seeing in Spaces like these: people stack a character LoRA, a style LoRA, and sometimes a concept LoRA onto FLUX-lineage models, then fight with merge weights until the output holds together. It's the manual version of what RefDiT automates, composition by hand, adapter by adapter, with mixed results. The demand for this workflow is real. 2.54k likes clears the noise floor for a space with no marketing behind it, and the workflow it represents, reference plus prompt plus composable adapters, is still easier to assemble from open parts than from any closed editor.
Open-Sora 2.0: the cost curve breaks
Open-Sora's changelog reads like a cost curve for open video generation. The project shipped its first version in March 2024 with a 3-day training run for 2-second 512x512 videos, published a 46% training cost reduction that same month, and kept pushing on both quality and price through 2024 and early 2025.
| Version | Date | Model size | What landed |
|---|---|---|---|
| 1.0 | 2024-03-18 | not listed | full pipeline; 2s 512x512 video after a 3-day training run |
| 1.1 | 2024-04-25 | not listed | 2s-15s, 144p-720p, any aspect ratio, image-to-video, video-to-video |
| 1.2 | 2024-06-17 | not listed | 3D-VAE, rectified flow, score conditioning |
| 1.3 | 2025-02-20 | 1B | shift-window attention, unified spatial-temporal VAE |
| 2.0 | 2025-03-12 | 11B | on par with HunyuanVideo 11B and Step-Video 30B on VBench and human preference |
The 2.0 release is the one to internalize. 11B parameters means the checkpoint fits on a single H100-class GPU, and the inference docs confirm it: 256px generation runs on one GPU, 768px scales across 8 with sequence parallelism. One checkpoint handles text-to-video and image-to-video. The training recipe is fully open, and the total budget is $200K. For context, Step-Video comes in at 30B parameters, nearly three times larger, and Open-Sora 2.0 matches its human preference scores. The 1.3 release's shift-window attention and unified spatial-temporal VAE are the architectural details that show up as temporal coherence, and 2.0 inherits them.
The VBench gap to Sora tells the story more precisely.
Going from a 4.52 percentage-point gap to 0.69 is the difference between "visibly behind the frontier" and "within striking distance." The human preference result against a 30B model makes the point sharper: parameter count is no longer the binding constraint for open video.
Key numbers:
- $200K total training budget for the 11B model, with checkpoints and training code fully open
- 11B parameters: 256px inference on one GPU, 768px across 8 GPUs
- One model does text-to-video and image-to-video; the text path routes through FLUX for the first frame
- Motion score is injected into the prompt at training time and exposed as --motion-score at inference, default 4
The motion-score detail is easy to miss. Open-Sora 2.0 is optimized for image-to-video, and the supported text-to-video route is two-stage: FLUX generates the first frame, Open-Sora animates it. If you run raw text-to-video and get mediocre output, that's why. The motion score gives you a direct handle on how much movement the model produces, which is the dial most video models hide from you.
What trips people up
The reference trap. Feed a personalization model a reference image with multiple subjects and you get a global average of the scene, not per-object control. Busy references produce muddled transfers. Crop the reference to the subject you care about, or use a method that does attribute-level decomposition. RefDiT exists because the single-identifier-token approach fails on natural scenes, and your product will hit the same wall.
Stacking LoRAs without a rank budget. Multi-LoRA composition works on FLUX-lineage models, but every adapter competes for the same rank budget. The browser Spaces hide this behind a friendly UI, which is convenient until you pile on three adapters and coherence collapses. Test each adapter alone first. Compose second. Don't treat merge weights as independent sliders.
Ignoring the frame-count constraint. Open-Sora's num_frames must satisfy 4k+1 and stay under 129. Requesting 128 or 130 frames fails at inference time. Use 65 for a standard short clip. The constraint comes from the temporal VAE structure, and it will burn an afternoon if you hit it cold.
Treating text-to-video as the primary path. Open-Sora 2.0 is optimized for image-to-video. The text-to-video route is text-to-image with FLUX, then image-to-video. Skip the first hop and you lose quality, then wonder why the open model underperforms the demos.
Benchmarking without the budget context. On VBench, Open-Sora 2.0 still trails Sora by 0.69%, and that's the actual headline. The point isn't that open is better. It's that $200K gets within striking distance of a commercial frontier model. Judge it as a cost curve, not as an absolute quality ranking.
One thing to remember
The prompt is no longer the only input to a diffusion model. Antenna knobs, attribute-level conditioning, LoRA stacks, and motion scores are all control surfaces now. The projects that matter are the ones giving you a precise handle on the part of generation you care about, instead of forcing another round of prompt rephrasing.
The bottom line
If you're building an image editor, adopt the RefDiT pattern of attribute-level token decomposition and local LoRA guidance, because global-style methods will fail the moment users bring natural multi-object reference photos.
If you're a small team shipping video features, adopt Open-Sora 2.0 rather than renting commercial APIs at per-minute prices. 11B parameters runs single-GPU at 256px, the checkpoint and training code are open, and a $200K training budget means this model class is about to get cheap fast.
If you're designing diffusion UX, expose intermediate states and let users control them directly, like Diffusion TV's antenna. The denoising process is the most engaging part of the stack, and hiding it behind a progress bar is a product mistake. One thing to watch: multi-LoRA composition tooling is consolidating into editor-native features, so expect attribute-level local control in mainstream editors within two quarters.