Published results

What we have measured so far.

Every number on this site comes from a public report. Each is listed here with the conditions it was measured under.

+50.8%

FlashAttention-2 on A100

Best case from an automated run; 25 to 51% at sequence length 1,024 and up with batch size 4 and up.

Report, March 2026

+13.3%

FlashAttention-4 on B300

Over the open-source kernel, after an agent studied traces of the closed-source cuDNN kernel. 5.2 to 13.3% across shapes.

G-Watch, September 2026

+4.2%

FlashAttention-3 on H200

At best, on a kernel already above 90% of peak throughput. Typically 1 to 3%. Submitted upstream.

Report, March 2026

+1.3%

FlashAttention-4 on Blackwell

Average across attention modes, precisions and sequence lengths from 1k to 32k, with no accuracy regression.

Report, March 2026

A better trace saves an agent iterations.

We gave a coding agent one task four times: raise the throughput of FlashAttention-3 on an H100. Each session read a different profile. Everything else stayed the same.

All four sessions ended near the same plateau, about 566 TFLOPS. The chart shows the first iteration at which each came within 2.5 TFLOPS of it. The existing traces reported a wait that their own probes had created, which sent the agent in the wrong direction.

Iterations until the agent reached the target

Xtrace trace
14
Existing tracer A
54
Existing tracer B
71
Hardware counters only
85

FlashAttention-3 on NVIDIA H100. Source: G-Watch, September 2026.


Anatomy of one optimization

From a stalled warp to a faster kernel.

One real run, on the FlashAttention-4 forward kernel for NVIDIA Blackwell. Step through it.

The starting kernel runs at about 1,541 TFLOPS at batch 8 and sequence length 8,192. It is already a heavily tuned kernel.

What the trace showed, per iteration

Softmax stage
3091
Matrix multiply
1712
Matrix multiply stalled
1290

Nanoseconds, approximate. FlashAttention-4 forward on Blackwell. Source: engineering report, March 2026.


Learning from a kernel with no source code.

On B300, cuDNN's attention kernel ran faster than open-source FlashAttention-4, and cuDNN ships no source. We traced both on the same inputs and asked an agent to close the gap using only the two traces.

What the trace showed What the agent changed Gain
The kernel ran its thread blocks in four waves, with idle time between them. Made the grid persistent, so each block fetches its next tile in place. +9.4%
Persistent blocks finished up to 1.49× apart. Sorted tiles across all groups, not only within each group. +2.0%
Long tiles were paired with long tiles. Used a fixed schedule on those shapes, pairing long with short. +11.1%
After the last tile, the kernel still waited 544 ns to free a buffer nothing would use. Dropped the final wait. +1.2%

Gains as reported for each change; not every change applies to every shape. Together they lift FlashAttention-4 by 5.2 to 13.3% across four shapes. Source: G-Watch, September 2026.

All published kernel results.

Kernel Hardware Result Conditions Source
FlashAttention-2 forward NVIDIA A100 up to +50.8% 25 to 51% at sequence length 1,024 and up with batch size 4 and up. Two small configurations were 7 to 8% slower. March 2026
FlashAttention-3 forward NVIDIA H200 up to +4.19% Typically 1 to 3%, on a baseline already above 90% of peak throughput. Submitted to the FlashAttention-3 repository. March 2026
FlashAttention-4 forward NVIDIA Blackwell +1.3% average Across attention modes, two precisions and sequence lengths from 1k to 32k. No accuracy regression. March 2026
FlashAttention-4 NVIDIA B300 +5.2 to +13.3% Four shapes, from traces of cuDNN's closed-source attention kernel. September 2026
FlashAttention-3, agent run NVIDIA H100 3.9× sooner Iteration 14 with Xtrace, against 54 with the best existing tracer. September 2026

Want numbers like these on your own kernels?

Talk to the optimization team, or start with the open-source tools.