Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Home
Log in
HomeWorldPoliticsBusinessTechHealthScienceClimateEducation
Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.Kodawire — SEE EVERY SIDE. KNOW THE TRUTH.

Balanced coverage of the week's biggest story. See every side. Know the truth.

By signing up for the Story of the Week newsletter you agree to our privacy policy.

Product

  • For You
  • Trending
  • Blindspot
  • Local
  • Compare
  • Brief
  • Trust

Topics

  • World
  • Politics
  • Business
  • Tech
  • Health
  • Science

Regions

  • Africa
  • Europe
  • Asia-Pacific
  • Americas
  • Middle East

Company

  • About
  • Contact
  • All topics
  • My Bias
  • Privacy
  • Terms & DMCA
  • Partner Offers

KODAWIRE

KODAWIRE

© 2026 Kodawire

AboutContactPrivacyTermsTop
/
/
Back

Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

1 sources

Google Developers Blog (XX)37m ago

Center

Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

To address the quadratic latency bottleneck of self-attention in high-resolution video diffusion models, developers implemented Sparse VideoGen (SVG) to dynamically route attention heads to highly structured spatial or temporal sparse masks. Translating this algorithmic sparsity into physical hardware speedups on TPUs required optimizing the Splash Attention kernel by bypassing empty memory tiles, restricting exact coordinate masking strictly to boundary tiles, and permuting token memory layouts into temporal-major order for contiguous access. By aligning these sparse masks with actual hardware tile execution, the combined optimizations significantly reduced wasted matrix operations and achieved up to a 1.69x end-to-end inference speedup for 1440p video generation.

Original

Spectrum

L 0%C 100%R 0%