Processor and graphics card specifications are live
Standalone test · WebGPU compute

WebGPU compute. A shader with its count stated.

A WGSL compute shader does 256 dependent multiply-adds on each of 4,194,304 numbers, 2,147,483,648 floating-point operations a dispatch, submitted one at a time with a fence on each for 10 seconds; then a 64 MiB buffer is copied 20 times the same way. GFLOP/s, dispatches per second, GB/s, and the adapter's own name for itself.

SystemCheck · WebGPU compute
WebGPU computedispatch, fence, repeat; then copy
one dispatch, one fence, for 10 s64 MiB copy4,194,304 elements x 256 FMAs x 2 = 2,147,483,648 flops per dispatchx 20, GB/stimed from submit to the browser’s own fence, overhead included

An illustration of the shape, not a measurement. The test reports your own figures.

Compute

4,194,304 elements x 256 FMAs

Per dispatch

2,147,483,648 flops

Duration

10 s, one fence per dispatch

Copy

64 MiB, 20 times

Run it

The compute test. About 12 seconds, in this tab.

10 seconds of compute dispatches, then 20 buffer copies. Keep this tab in front.

What you will get

Billions of floating-point operations per second from a compute shader whose count is stated, dispatches per second, how steady they were, the copy bandwidth in GB/s, and the adapter’s own name for itself when it gives one. Nothing is uploaded.

One shader, one fence, repeated.

WebGPU gives a page a compute shader: a program the graphics chip runs over a buffer with no picture involved. It is the first browser path where a page can count the arithmetic it asked for and time the chip finishing it, so this test does exactly that and states the count.

The shader
Each invocation loads one 32-bit float, performs 256 fused multiply-adds that each depend on the last, and stores it back. The dependency is the point: the compiler cannot fold the loop, and the store means the work cannot be dropped. A workgroup is 256 invocations; one dispatch covers 4,194,304 elements.
The count
4,194,304 elements x 256 FMAs x 2 operations per FMA (a multiply and an add) = 2,147,483,648 floating-point operations per dispatch. That is the number every GFLOP/s figure on this page is worked from. It is the arithmetic asked for, not the chip's rated peak.
The fence
Each dispatch is submitted alone and timed from the submit to the browser's own promise that the work is done (queue.onSubmittedWorkDone). 3 warm-up dispatches are timed and dropped. The time includes the browser's submission and completion overhead, which on a fast chip is a real share of a two-billion-operation dispatch; the page says so rather than subtracting a guess.
The copy
A 64 MiB buffer, filled first so a driver cannot shortcut a never-written one, copied to another buffer 20 times with the same fence. Each copy's bytes are counted once, though a copy reads them and writes them, and the figure is in decimal gigabytes.

Four figures, and the adapter's name.

No score. The figures are what a web page gets through WebGPU with a fence on every dispatch, stated in units whose counts are on this page.

GFLOP/s

2,147,483,648 operations times the dispatches timed, over the sum of their times, in billions per second. On a discrete chip the fence overhead is most of a dispatch and the figure sits well under the chip's peak; that is a fact about the browser path, and it is the same for every run in the same browser.

Dispatches per second

The timed dispatches over the sum of their times, with the mean milliseconds per dispatch beside it. This is the figure the fence overhead shows in directly.

Consistency

One minus the coefficient of variation of the dispatch times, out of 100. A chip left alone gives near 100; a chip sharing its time with a browser compositor, a video, or a power limit that moves during the run gives less.

Copy GB/s and the adapter

64 MiB times the copies over the sum of their times, decimal. The adapter's name is what adapter.info reports, joined from its description, vendor, architecture and device fields; when it reports nothing, the page says so and does not guess. If it reports itself as a fallback adapter, the shader ran on the processor and the result says so.

SystemCheck · WebGPU compute
WebGPU computedispatch, fence, repeat; then copy
one dispatch, one fence, for 10 s64 MiB copy4,194,304 elements x 256 FMAs x 2 = 2,147,483,648 flops per dispatchx 20, GB/stimed from submit to the browser’s own fence, overhead included

An illustration of the shape, not a measurement. The test reports your own figures.

Hardware

What it needs. Three tiers, no fine print.

Minimum

A browser with WebGPU

Chrome or Edge 113 or later on Windows, macOS and ChromeOS; Safari 26; Firefox 141 or later on Windows. Firefox elsewhere and older Safari report not supported.

Recommended

Nothing else on the graphics chip

A video playing, a game, or another tab drawing shares the chip and lowers consistency before it lowers the rate.

Optimal

Plugged in, run twice

A laptop on battery caps its graphics clock; a second run right after the first shows whether the clock moved during the first.

Questions

WebGPU compute, answered

What does a WebGPU compute test measure?
How fast the graphics chip finishes a known amount of arithmetic when asked through the browser. The shader does 2,147,483,648 floating-point operations per dispatch, the browser reports when each dispatch is done, and the figure is operations over time. A 64 MiB buffer copy timed the same way gives the chip's copy bandwidth through the same path.
Why is my GFLOP/s so far under the chip's rated number?
Three reasons, all stated. The rated number counts a fused multiply-add as two operations at every unit running flat out; so does this page, but the shader's chain of dependent FMAs cannot keep every unit busy. The fence on every dispatch adds the browser's submission and completion time to each one. And the work is one dispatch at a time, with nothing queued behind it. The figure is what a page gets, which is the honest thing to report.
Why does it say not supported?
The browser has no navigator.gpu. Firefox stable has WebGPU only on Windows from version 141, Safari from version 26, and older versions of either have none. A browser without it has no compute path to measure, so the page reports not supported instead of a made-up figure or a slower fallback presented as the same thing.
Where does the adapter name come from?
From adapter.info, which the browser fills from the driver. Chrome often gives a vendor and an architecture and no model, and some browsers give nothing at all. The page shows what it is given and says when it is given nothing. It never looks the name up or infers it from a figure.
Is this part of the main SystemCheck run?
No. It runs on this page only, on its own, and it does not change the score or the report. The main run's graphics stages use WebGL2, which every current browser has; this test uses WebGPU, which not every browser has, which is why it stands alone.
What can it not see?
The chip's clock, its temperature, and how much of it the browser was given. A GFLOP/s figure below your expectation might be the fence overhead, a power limit, a shared chip, or a driver that ran the shader on a fallback path; the consistency figure and the fallback flag narrow it, and nothing on this page claims to name the cause.

Compute through the browser, with the count stated.