Peak benchmark scores measure a machine at its coldest and least loaded, a state it occupies for seconds and never returns to during real work.
- Boost algorithms are explicitly tuned for short bursts, so a benchmark that finishes inside the boost window measures the tuning rather than the hardware.
- The gap between peak and sustained is largest exactly where buyers are least equipped to see it: thin laptops, small-form-factor builds, and dense enclosures.
- Optimising for peak scores has real product consequences. Cooling and sustained power budgets are where cost gets cut, because nobody benchmarks them.
- The fix is not a better single number but reporting the shape of the run: sustained throughput, decay, and consistency alongside the peak.
The measurement everyone agrees to take
There is a shared convention in consumer hardware that performance is a scalar. One number, bigger is better, sorted into a leaderboard. It is enormously convenient (it makes comparison possible for people who cannot evaluate an architecture) and it is achieved by throwing away almost everything about how a machine behaves.
The specific thing thrown away is time. A benchmark that runs for fifteen seconds measures a cold chip with a full boost budget, an empty thermal mass, and no accumulated heat in the cooler. That is a genuine state the machine can enter. It is not a state it stays in, and it is not the state your ten-minute export happens in.
Boost is designed to win this test
Modern boost algorithms explicitly allow a chip to exceed its sustained power budget for a bounded window, on the reasoning that the cooler and heat spreader can absorb the excess energy before temperature catches up. Intel calls its version a turbo power limit with a time constant; AMD and every mobile vendor implement equivalents. It is a legitimate engineering technique and it makes machines feel responsive.
It also means that any benchmark completing inside that window is measuring the boost policy rather than the sustained capability of the silicon. Two laptops with identical CPUs can post nearly identical peak scores and differ by 40% in sustained throughput, because one has a 15W sustained budget and a heat pipe, and the other has 45W and a vapour chamber. The peak score cannot see the difference. The user finds it on the first long job.
Where the gap is worst
Peak and sustained diverge in proportion to how thermally constrained the design is, which means the error is largest exactly where the buyer has the least ability to detect it.
A large desktop with a serious air cooler decays a few percent. Peak is close enough to sustained that the distinction rarely changes a decision. A thin laptop can decay 30% or more within two minutes, by design, and the peak score in the review says nothing about it. Small-form-factor builds, all-in-ones, and mini PCs sit in between and vary enormously depending on how the enclosure was engineered.
The result is a systematic bias: the more a machine's real-world performance depends on its thermal design, the less the headline number tells you about it.
What optimising for peak actually costs
This is not merely a measurement inaccuracy; it feeds back into what gets built. When the number that sells a product is set in the first fifteen seconds, the rational engineering response is to spend the budget on hitting a high boost clock and to economise on everything that only matters after the first minute: heat pipe count, vapour chamber area, fan quality, VRM cooling, sustained power headroom.
None of those cuts show up in a peak score. All of them show up in a long export, a long compile, or the second hour of a gaming session. The reviewer benchmark and the user experience have been decoupled, and the decoupling is not accidental.
The same dynamic explains why thermal and acoustic behaviour so often feels like an afterthought on machines that reviewed well. It was not an afterthought. It was a budget line with no scoreboard attached.
What to report instead
The answer is not a better single number, because any single number has the same problem. It is to report the shape of the run alongside the total: sustained throughput at equilibrium, decay from peak to plateau, and consistency of delivery. Three numbers instead of one is still comparable and carries most of the information the scalar destroyed.
That is the design SystemCheck takes: a sustained multi-core load long enough to reach equilibrium, throughput retained across the whole run rather than averaged away, and decay and consistency reported as first-class results rather than footnotes. The composite grade still exists, because comparison needs one, but it is built on sustained behaviour rather than on the best interval observed.
It is worth being clear about the limits of that too. SystemCheck runs in a browser and cannot read a temperature sensor, a power rail, or a fan tachometer. No browser can. It infers thermal behaviour from throughput decay over sustained load. When you need the physical cause rather than the behavioural effect, run HWiNFO or your vendor's utility alongside it.
The question worth asking of any benchmark
Before trusting a number, ask how long the measurement ran and whether the machine reached equilibrium during it. If the answer is under about thirty seconds of sustained load, the number describes a cold machine, and you should treat it as an upper bound rather than an expectation.
Then ask what was discarded to produce it. Every score is a reduction of a time series. Knowing what was reduced away is most of the skill in reading benchmarks at all.
Terms used here
Questions
- Are peak benchmark scores useless?
- No, they accurately describe short, bursty workloads, which is a large share of interactive computing. They become misleading when used to predict long workloads, because the machine leaves the boost window and never returns to it during the job.
- How long does a benchmark need to run to be meaningful?
- Long enough for the system to reach thermal and power equilibrium, which on most machines takes between thirty seconds and two minutes of genuinely saturating load. A measurement that ends before equilibrium is measuring the boost policy, not the sustained capability.