Fewer updates. 120 → 20 generator updates
Representation Distribution Matching · Few-Step Causal Video
ViRDM
Taming Representation Distribution Matching for
Few-Step Causal Video Generation
Northeastern University · Adobe Research
TL;DR ViRDM enables teacher- and critic-free post-training for few-step causal video generation by matching generated videos directly to a fixed video-text representation distribution. With only 20 generator updates, it reaches 84.87 VBench Total while reducing peak memory from 77.1 to 48.3 GB per GPU and post-training time from 22 to 2 hours on eight A100 GPUs.
Better quality. 84.51 → 84.87 Total Score
Faster post-training. 22 → 2 hours
Less memory. 77.1 → 48.3 GB per GPU
Feasible on one GPU. 68.5 GB on one A100 GPU
0. ViRDM Generation Showcase
1. Motivation
DMD-based methods keep three diffusion models active.
Existing causal post-training methods (e.g., Self-Forcing and Causal Forcing) rely on a frozen 14B score teacher, a learned 1.3B critic, and the 1.3B generator. This teacher–critic stack makes every update expensive in both memory and compute.
Replace the teacher and critic with a fixed representation distribution.
Image RDM suggests a simpler route: precompute the real representation distribution once, then train only the generator by directly matching its generated distribution to that fixed target.
2. Challenges
The transfer problem
Image RDM does not directly transfer to few-step causal video.
Video introduces a multi-step autoregressive rollout, a distinct optimization regime, and temporal behavior that the image recipe was never designed to constrain. These expose three video-specific barriers.
The end-to-end gradient path is memory-intractable
Backpropagating through a multi-step autoregressive rollout, VAE decoder, and representation encoder requires the full composite graph to remain resident and cannot fit at the target video resolution.
The image RDM optimization regime does not transfer
Image RDM relies on very large fresh populations and can begin from a bidirectional model. At video scale, those populations are prohibitively expensive and the same initialization fails to recover a causal transport.
Representation distributions underconstrain temporal dynamics
Image features largely capture appearance, while video features do not reliably distinguish moderate from strong motion.
3. ViRDM Objective
Representation extraction
How we extract joint representations
Frozen video and text encoders first extract modality-specific features; ViRDM then concatenates them into the joint feature used for distribution matching.
- FIXED
Encode the target once. The reference contains 6,505 real video–text pairs and remains frozen throughout post-training.
- JOINT
Match video and text jointly. V-JEPA 2.1 video features and normalized SigLIP2 text features make visual quality and prompt fidelity part of the same objective.
Distribution comparison
How the distributions are compared
Generated samples are attracted toward the fixed reference distribution while the within-generation term keeps them from collapsing.
- MMD
Three terms, one distribution match. Following RDM, ViRDM compares the generated and reference distributions with a three-term MMD objective.
4. Memory-Feasible ViRDM Training
ViRDM makes training memory-feasible through stochastic exits and staged vector–Jacobian products.
4.1
Stochastic Exits for Few-Step Video RDM
RDM requires only a clean endpoint, with no intermediate supervision. We therefore follow Self-Forcing’s consistency-sampling rollout, in which each denoising step directly predicts the clean endpoint. We sample one stochastic exit, run preceding steps without gradients, and backpropagate only through the clean prediction at the selected exit.
4.2
Staged VJPs Avoid Co-Resident Backward Graphs
Naive end-to-end backpropagation requires the generator, decoder, and representation-encoder graphs to coexist in memory. ViRDM factors the backward pass into staged VJPs, materializing each upstream gradient before releasing the current graph, so only one module’s backward graph is resident at a time.
Measured for 81-frame, 832×480 videos with one-video-per-GPU microbatches on 80 GB A100 GPUs.
5. Video-Specific RDM Recipe
Few-step causal video changes both the useful population scale and the initialization required for RDM.
5.1
How Many Fresh Generated Videos Are Needed?
Unlike image RDM, video RDM already works with small fresh populations, and its quality–efficiency trade-off largely saturates at B=64 videos.
Increasing batch size increases accumulation and update time, but delivers far less improvement than expected.
Total
Quality
Semantic
5.2
Causal Initialization Is Essential
RDM effectively refines a causal transport, but does not reliably create one from a bidirectional model in only 20 updates. We adopt the causal initialization families from Causal Forcing.
Total
Quality
Semantic
Dynamic Degree
Directly starting from the bidirectional model causes severe drift under chunkwise causal generation. Teacher forcing prevents this failure and preserves the scene, but produces substantially weaker dynamics; causal few-step initializations maintain coherence while supporting stronger temporal evolution.
The Recipe So Far
The recipe so far leads Total without Dynamic Degree, while motion remains the residual gap.
Total w/o Dynamic Degree
Dynamic Degree
6. Closing the Residual Dynamics Gap
Image Marginals Miss Dynamics; Video Representations Capture It Nonlinearly
We test motion sensitivity by freezing different percentages of each video.
The image representation distribution barely responds to motion changes.
The video representation distribution responds nonlinearly and cannot reliably distinguish moderate from high motion.
ViRDM Trained on Image Marginals vs. Video Marginals
Video marginals provide the missing temporal signal, raising Dynamic Degree by 30.55 points (18.06 → 48.61) while improving Total and Quality.
Total
Quality
Semantic
Dynamic Degree
Image marginals preserve frame appearance but produce little temporal evolution.
Video marginals create clearer motion while maintaining the scene and subject.
A Simple Flow-Based Dynamics Regularizer
A small flow-based correction using frozen RAFT closes the remaining gap without degrading other dimensions.
Ours w/o Dyn. Reg. vs. Ours w/ Dyn. Reg.
Total
Quality
Semantic
Dynamic Degree
7. Main Results
ViRDM achieves better video quality with 11× less post-training time and 28.8 GB less memory per GPU.
VBench Comparison on Four-Step Causal Video Generation Methods
Total
Quality
Semantic
User Study
People prefer ViRDM.
Across 25 participants and 30 prompts, ViRDM leads the matched four-way comparison and is preferred over the variant without dynamics regularization.
Matched Four-Step Methods
20 prompts · four-way preference
Text Alignment
Visual Quality
Dynamics Regularization
10 prompts · pairwise preference
Text Alignment
Visual Quality
Dynamics
Teacher- and critic-free
−28.8 GB on A100
Eight A100 GPUs
Qualitative Comparison Between Matched Four-Step Methods
8. Extension Beyond Four-Step Causal Generation
Causal
1 / 2 / 4 Steps
We follow ASD, and the first block remains four-step.Total
Quality
Semantic
Bidirectional
1 / 2 / 4 Steps
The complete video is generated as one block.Total
Quality
Semantic
9. References
- 1One-step Diffusion with Distribution Matching Distillation
T. Yin et al. · CVPR 2024
Paper ↗ - 2Improved Distribution Matching Distillation for Fast Image Synthesis
T. Yin et al. · NeurIPS 2024
Paper ↗ - 3Representation Distribution Matching for One-Step Visual Generation
L. Feng et al. · arXiv 2026
Paper ↗ - 4Consistency Models
Y. Song et al. · ICML 2023
Paper ↗ - 5Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
X. Huang et al. · NeurIPS 2025
Paper ↗ - 6Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
H. Zhu et al. · ICML 2026
Paper ↗ - 7DINOv2: Learning Robust Visual Features without Supervision
M. Oquab et al. · TMLR 2024
Paper ↗ - 8V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
L. Mur-Labadia et al. · arXiv 2026
Paper ↗ - 9RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
Z. Teed and J. Deng · ECCV 2020
Paper ↗ - 10VBench: Comprehensive Benchmark Suite for Video Generative Models
Z. Huang et al. · arXiv 2023
Paper ↗ - 11Towards One-Step Causal Video Generation via Adversarial Self-Distillation
Y. Yang et al. · ICLR 2026
Paper ↗ - 12From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
T. Yin et al. · CVPR 2025
Paper ↗ - 13Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan et al. · arXiv 2025
Paper ↗ - 14TAEHV: Tiny AutoEncoder for Hunyuan Video
O. Boer Bohan · 2025
Project ↗ - 15SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
M. Tschannen et al. · 2025
Paper ↗
10. BibTeX
@article{meng2026virdm,
title = {ViRDM: Taming Representation Distribution Matching for
Few-Step Causal Video Generation},
author = {Meng, Zichong and Ge, Chongjian and Huang, Chun-Hao P. and Zhou, Yang and Jiang, Huaizu},
journal = {arXiv preprint},
year = {2026}
}