videoJul 18, 2026 · 17 min
Optimizing block-sparse INT8 attention and multi-expert step caching for Wan 2.2
Analyzing hardware execution limits and block selector heuristics in Wan 2.2 video generation. Demonstrates 1.30–1.76× speedups in block-sparse INT8 attention using a mean-pool selector and 1.78× speedups via per-expert diffusion step caching.
video generationattentionsparsityquantization
1.78×two-expert step cacheno cache ▸ per-expert