Video Diffusion Inference Optimization

GB200 · 端到端加速消融 + Baseline vs Full-Opt 生成对比 · 同 seed / 同配置 / median

① 加速方法与各方法贡献 (Ablation)

Wan2.2 TI2V-5B

单卡 · 704×1280 · 121 帧 · 50 步

2.885×
70.25s → 24.35s (baseline → full-opt)
方法系数
Kernel regional-compile + qkv-fusion + cross-kv-cache + bf16-glue(无损)1.52×
Cache EasyCache 0.036(47% 步复用)1.90×
Attention (PISA)弃(0.85×)

Wan2.2-A14B (14B)

4-GPU CP4 · 720×1280 · 81 帧 · 40 步

2.19×
同拓扑纯优化:~129s (CP4) → 58.89s
方法系数
Kernel fused-qkv + compiled + async a2a(CP4 通信稀�)1.12×
Cache EasyCache 0.30(35% 复用)1.52×
Attention PISA density 0.101.28×

LingBot-Video (MoE 30B-A3B)

4-GPU CP4+FSDP · base 480p → refiner 1080p

2.6×
同拓扑纯优化:375.53s → 144.36s
方法系数
Attention backend fa2 → cuDNN(base+refiner;无独立通用 kernel 优化)1.79×
Attention sparsity refiner-only PISA 0.10(替换 refiner 的 cuDNN)1.12×
Cache EasyCache base(1.18×) + refiner(1.10×)1.30×
系数为增量口径(每个方法在前一个之上),相乘 ≈ 该模型总加速。各模型主导方法不同:5B=Cache 主导、14B=Cache+Attention、LingBot=Attention-backend(fa2→cuDNN) 主导。注意 LingBot 没有独立的通用 kernel 优化:1.79× 就是 attention backend 更换,PISA 再在 refiner 把 cuDNN 换成稀疏 attention —— 两项都是 attention 改动。PISA 只在大注意力上有效(14B / LingBot-refiner ✓;5B ✗)。

② 端到端生成对比:Baseline vs Full-Opt

每个 prompt 一个独立视频,左 = Baseline,右 = Full-Opt。

Wan2.2 TI2V-5B 2.885×

单卡 · 704×1280 · 121 帧 · 50 步 · seed 42 · 左 Baseline 70.25s / 右 Full-Opt 24.35s (kernel + EasyCache)

Prompt 0Will Smith casually eats noodles at a street food market

Prompt 1A lone hiker atop a towering cliff against the vast horizon

Prompt 2A hand tossing a bright yellow lemon from a wooden bowl

Prompt 3A curious raccoon peers through a field of yellow sunflowers

Prompt 4A superintelligent humanoid robot waking up in a lab

Wan2.2-A14B (14B) 2.19× (纯优化)

4-GPU CP4 · 720×1280 · 81 帧 · 40 步 · seed 1024 · 左 Baseline (单卡 naive) / 右 Full-Opt 58.89s

Prompt 0Will Smith casually eats noodles at a street food market

Prompt 1A lone hiker atop a towering cliff against the vast horizon

Prompt 2A hand tossing a bright yellow lemon from a wooden bowl

Prompt 3A curious raccoon peers through a field of yellow sunflowers

Prompt 4A superintelligent humanoid robot waking up in a lab

视频对比是相对单卡 naive baseline(含 4-GPU 并行化,视觉比值更大);Ablation 的 2.19× 是相对同拓扑 CP4 baseline 的纯优化贡献。

LingBot-Video (MoE 30B-A3B) 2.6×

4-GPU CP4+FSDP · refiner 1088×1920 · 121 帧 · seed 42 · 左 Baseline 375.5s / 右 Full-Opt 144.4s (cuDNN + PISA + EasyCache)

Prompt 0A young woman modeling an outfit in a bright modern apartment

Prompt 1A child blowing shimmering bubbles outdoors on a sunny day

Prompt 2First-person view of a desk workspace with a game controller