Representation Distribution Matching · Few-Step Causal Video

ViRDM

Taming Representation Distribution Matching for
Few-Step Causal Video Generation

Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou†, Huaizu Jiang†

Northeastern University · Adobe Research

† Equal advising · Technical Report

See showcase Paper · coming soon Code ↗ Models ↗

TL;DR ViRDM enables teacher- and critic-free post-training for few-step causal video generation by matching generated videos directly to a fixed video-text representation distribution. With only 20 generator updates, it reaches 84.87 VBench Total while reducing peak memory from 77.1 to 48.3 GB per GPU and post-training time from 22 to 2 hours on eight A100 GPUs.

Fewer updates. 120 → 20 generator updates

Better quality. 84.51 → 84.87 Total Score

Faster post-training. 22 → 2 hours

Less memory. 77.1 → 48.3 GB per GPU

Feasible on one GPU. 68.5 GB on one A100 GPU

0. ViRDM Generation Showcase

1. Motivation

Why

DMD-based methods keep three diffusion models active.

Existing causal post-training methods (e.g., Self-Forcing and Causal Forcing) rely on a frozen 14B score teacher, a learned 1.3B critic, and the 1.3B generator. This teacher–critic stack makes every update expensive in both memory and compute.

Idea

Replace the teacher and critic with a fixed representation distribution.

Image RDM suggests a simpler route: precompute the real representation distribution once, then train only the generator by directly matching its generated distribution to that fixed target.

2. Challenges

The transfer problem

Image RDM does not directly transfer to few-step causal video.

Video introduces a multi-step autoregressive rollout, a distinct optimization regime, and temporal behavior that the image recipe was never designed to constrain. These expose three video-specific barriers.

Three barriers exposed by the image-to-video RDM transfer:

01

The end-to-end gradient path is memory-intractable

Backpropagating through a multi-step autoregressive rollout, VAE decoder, and representation encoder requires the full composite graph to remain resident and cannot fit at the target video resolution.

02

The image RDM optimization regime does not transfer

Image RDM relies on very large fresh populations and can begin from a bidirectional model. At video scale, those populations are prohibitively expensive and the same initialization fails to recover a causal transport.

03

Representation distributions underconstrain temporal dynamics

Image features largely capture appearance, while video features do not reliably distinguish moderate from strong motion.

3. ViRDM Objective

Representation extraction

How we extract joint representations

Frozen video and text encoders first extract modality-specific features; ViRDM then concatenates them into the joint feature used for distribution matching.

  • FIXED

    Encode the target once. The reference contains 6,505 real video–text pairs and remains frozen throughout post-training.

  • JOINT

    Match video and text jointly. V-JEPA 2.1 video features and normalized SigLIP2 text features make visual quality and prompt fidelity part of the same objective.

Distribution comparison

How the distributions are compared

Generated samples are attracted toward the fixed reference distribution while the within-generation term keeps them from collapsing.

reference distribution generated
RDM LOSS=
GEN–GEN REPULSION +GEN–REF ATTRACTION +REF–REF CONSTANT
  • MMD

    Three terms, one distribution match. Following RDM, ViRDM compares the generated and reference distributions with a three-term MMD objective.

4. Memory-Feasible ViRDM Training

ViRDM makes training memory-feasible through stochastic exits and staged vector–Jacobian products.

4.1

Stochastic Exits for Few-Step Video RDM

RDM requires only a clean endpoint, with no intermediate supervision. We therefore follow Self-Forcing’s consistency-sampling rollout, in which each denoising step directly predicts the clean endpoint. We sample one stochastic exit, run preceding steps without gradients, and backpropagate only through the clean prediction at the selected exit.

4.2

Staged VJPs Avoid Co-Resident Backward Graphs

Naive end-to-end backpropagation requires the generator, decoder, and representation-encoder graphs to coexist in memory. ViRDM factors the backward pass into staged VJPs, materializing each upstream gradient before releasing the current graph, so only one module’s backward graph is resident at a time.

Peak Memory for One Complete ViRDM Update

Measured for 81-frame, 832×480 videos with one-video-per-GPU microbatches on 80 GB A100 GPUs.

5. Video-Specific RDM Recipe

Few-step causal video changes both the useful population scale and the initialization required for RDM.

5.1

How Many Fresh Generated Videos Are Needed?

Unlike image RDM, video RDM already works with small fresh populations, and its quality–efficiency trade-off largely saturates at B=64 videos.

Increasing batch size increases accumulation and update time, but delivers far less improvement than expected.

Measured scoreExpected if the B=8→64 return continued
Total
Quality
Semantic

5.2

Causal Initialization Is Essential

RDM effectively refines a causal transport, but does not reliably create one from a bidirectional model in only 20 updates. We adopt the causal initialization families from Causal Forcing.

Total
Bidirectional65.80
Teacher Forcing83.21
Causal CD82.99
Causal ODE default83.41
Quality
Bidirectional68.82
Teacher Forcing84.05
Causal CD83.92
Causal ODE default84.11
Semantic
Bidirectional53.72
Teacher Forcing79.85
Causal CD79.25
Causal ODE default80.61
Dynamic Degree
Bidirectionalincoherent drift51.39
Teacher Forcing43.18
Causal CD47.44
Causal ODE default48.61
Qualitative Comparison of Different Initializations

Directly starting from the bidirectional model causes severe drift under chunkwise causal generation. Teacher forcing prevents this failure and preserves the scene, but produces substantially weaker dynamics; causal few-step initializations maintain coherence while supporting stronger temporal evolution.

The Recipe So Far

The recipe so far leads Total without Dynamic Degree, while motion remains the residual gap.

Total w/o Dynamic Degree
Self Forcing85.34
Causal Forcing84.64
ViRDM · no dyn.85.77
Dynamic Degree
Self Forcing67.78
Causal Forcing83.33
ViRDM · no dyn.48.61

6. Closing the Residual Dynamics Gap

6.1

Image Marginals Miss Dynamics; Video Representations Capture It Nonlinearly

Does the Representation Capture Motion?

We test motion sensitivity by freezing different percentages of each video.
The image representation distribution barely responds to motion changes.
The video representation distribution responds nonlinearly and cannot reliably distinguish moderate from high motion.

ViRDM Trained on Image Marginals vs. Video Marginals

Video marginals provide the missing temporal signal, raising Dynamic Degree by 30.55 points (18.06 → 48.61) while improving Total and Quality.

Total
Image82.70
Video83.41
Quality
Image83.16
Video84.11
Semantic
Image80.88
Video80.61
Dynamic Degree
Image18.06
Video48.61
Representation Choice Changes Temporal Behavior

Image marginals preserve frame appearance but produce little temporal evolution.
Video marginals create clearer motion while maintaining the scene and subject.

6.2

A Simple Flow-Based Dynamics Regularizer

A small flow-based correction using frozen RAFT closes the remaining gap without degrading other dimensions.

Ours w/o Dyn. Reg. vs. Ours w/ Dyn. Reg.

Total
Ours w/o Dyn. Reg.83.41
Ours w/ Dyn. Reg.84.87
Quality
Ours w/o Dyn. Reg.84.11
Ours w/ Dyn. Reg.85.82
Semantic
Ours w/o Dyn. Reg.80.61
Ours w/ Dyn. Reg.81.09
Dynamic Degree
Ours w/o Dyn. Reg.48.61
Ours w/ Dyn. Reg.72.02
Qualitative Comparison between ViRDM w/ and w/o Dynamics Regularization

7. Main Results

ViRDM achieves better video quality with 11× less post-training time and 28.8 GB less memory per GPU.

VBench Comparison on Four-Step Causal Video Generation Methods

User Study

People prefer ViRDM.

Across 25 participants and 30 prompts, ViRDM leads the matched four-way comparison and is preferred over the variant without dynamics regularization.

(a)

Matched Four-Step Methods

20 prompts · four-way preference

Text Alignment
ViRDM40.4%
Self-Forcing32.6%
Causal Forcing20.2%
CausVid6.8%
Visual Quality
ViRDM43.0%
Self-Forcing32.6%
Causal Forcing20.6%
CausVid3.8%
(b)

Dynamics Regularization

10 prompts · pairwise preference

Text Alignment
ViRDM60.8%
w/o Dyn. Reg.39.2%
Visual Quality
ViRDM76.4%
w/o Dyn. Reg.23.6%
Dynamics
ViRDM85.6%
w/o Dyn. Reg.14.4%
Training stack
3 models→1 generator

Teacher- and critic-free

Peak memory / GPU
77.1 GB→48.3 GB

−28.8 GB on A100

Post-training time
22 h→2 h

Eight A100 GPUs

8. Extension Beyond Four-Step Causal Generation

Causal

1 / 2 / 4 Steps

We follow ASD, and the first block remains four-step.

Total

1 step84.27
2 steps84.42
4 steps84.87

Quality

1 step85.04
2 steps85.44
4 steps85.82

Semantic

1 step81.17
2 steps80.35
4 steps81.09

Bidirectional

1 / 2 / 4 Steps

The complete video is generated as one block.

Total

1 step83.12
2 steps84.53
4 steps84.56

Quality

1 step84.61
2 steps85.48
4 steps85.84

Semantic

1 step77.16
2 steps80.72
4 steps79.43
Qualitative Showcase of Four-Step Bidirectional Generation

9. References

  1. 1
    One-step Diffusion with Distribution Matching Distillation

    T. Yin et al. · CVPR 2024

    Paper ↗
  2. 2
    Improved Distribution Matching Distillation for Fast Image Synthesis

    T. Yin et al. · NeurIPS 2024

    Paper ↗
  3. 3
    Representation Distribution Matching for One-Step Visual Generation

    L. Feng et al. · arXiv 2026

    Paper ↗
  4. 4
    Consistency Models

    Y. Song et al. · ICML 2023

    Paper ↗
  5. 5
    Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

    X. Huang et al. · NeurIPS 2025

    Paper ↗
  6. 6
    Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

    H. Zhu et al. · ICML 2026

    Paper ↗
  7. 7
    DINOv2: Learning Robust Visual Features without Supervision

    M. Oquab et al. · TMLR 2024

    Paper ↗
  8. 8
    V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

    L. Mur-Labadia et al. · arXiv 2026

    Paper ↗
  9. 9
    RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

    Z. Teed and J. Deng · ECCV 2020

    Paper ↗
  10. 10
    VBench: Comprehensive Benchmark Suite for Video Generative Models

    Z. Huang et al. · arXiv 2023

    Paper ↗
  11. 11
    Towards One-Step Causal Video Generation via Adversarial Self-Distillation

    Y. Yang et al. · ICLR 2026

    Paper ↗
  12. 12
    From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

    T. Yin et al. · CVPR 2025

    Paper ↗
  13. 13
    Wan: Open and Advanced Large-Scale Video Generative Models

    Team Wan et al. · arXiv 2025

    Paper ↗
  14. 14
    TAEHV: Tiny AutoEncoder for Hunyuan Video

    O. Boer Bohan · 2025

    Project ↗
  15. 15
    SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

    M. Tschannen et al. · 2025

    Paper ↗

10. BibTeX

@article{meng2026virdm,
  title   = {ViRDM: Taming Representation Distribution Matching for
             Few-Step Causal Video Generation},
  author  = {Meng, Zichong and Ge, Chongjian and Huang, Chun-Hao P. and Zhou, Yang and Jiang, Huaizu},
  journal = {arXiv preprint},
  year    = {2026}
}