FlashAttention computes exact attention with an input-output-aware tiled algorithm that reduces expensive transfers between accelerator memory levels. This guide explains the mechanism, trade-offs, ...