Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

1MMLab, The Chinese University of Hong Kong 2Tencent AIPD

*Equal contribution   †Corresponding authors

Train on 5 seconds. Generate far beyond.

Real-score predictions for an owl rollout: video-level DMD retains historical artifacts, while RMD independently provides cleaner appearance targets.
Video-level DMD can retain artifacts from the generated history. RMD scores each chunk independently, providing a clean target despite degraded context.

TL;DR

Rollout-Marginal Distillation (RMD) keeps history for generation and evaluates appearance independently. By scoring each chunk without temporal context, RMD provides a cleaner visual-quality target, followed by video-level refinement to restore temporal coherence. Trained on 5-second rollouts, it helps autoregressive video generators preserve visual quality beyond the training horizon.

Long-Horizon Generation

Surf foam · Watch on YouTube
Flame · Watch on YouTube

Multi-Prompt Generation

Comparisons

Comparison 1 · Watch on YouTube
Comparison 2 · Watch on YouTube

Rollout-Marginal Training

A causal generator produces a rollout. Adapted score networks supervise chunks independently, and gradients flow through differentiable causal replay.
Independent appearance supervision, followed by video-level refinement. Teachers and critics are used only during training.

BibTeX

@article{gao2026rmd,
  title   = {Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation},
  author  = {Gao, Chenjian and Hu, Zhihao and Ma, Jianqi and Zhang, Jun and Zhang, Weidong and Xue, Tianfan},
  journal = {arXiv preprint arXiv:2609.37925},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.37925}
}