Replace convolutional denoisers with transformer-based diffusion architectures and connect diffusion to flow-matching objectives. Compare training stability, sampling speed, scaling behavior, and use in modern image and video systems.