There is no context-free best processor

A processor with more physical qubits can perform worse on a circuit that needs high-fidelity entangling gates or dense connectivity. A smaller device can excel on one structured workload and still be unsuitable for another. Benchmark choice should follow the intended task rather than the metric with the largest headline number.

Start with workload width, depth, gate family, connectivity, mid-circuit measurement, and required output accuracy. These requirements determine which hardware characteristics constrain the run. Comparing devices without the workload is closer to comparing instruments by the number of controls than by the measurement they must make.

Component metrics need operating conditions

Single- and two-qubit gate fidelities, readout error, coherence times, and leakage describe components under a calibration and test protocol. Average values can hide a weak edge in the connectivity graph, and the best reported value can describe only a selected pair. Performance can drift between calibration and user execution.

Check whether a value is mean, median, or selected best; whether uncertainty and date are reported; and whether gates were measured simultaneously. A circuit accumulates interacting errors, so multiplying isolated average fidelities is at best a rough model.

System benchmarks trade detail for comparability

Randomized benchmarking estimates aspects of gate performance under its assumptions. Quantum Volume combines width and depth through random circuits and a pass criterion. Volumetric benchmarks generalize the view across rectangular circuit shapes. Application-oriented suites execute algorithms or kernels chosen to resemble useful workloads.

Each benchmark samples a region of behavior. A single scalar is easy to compare but can conceal why a system passed, how circuits were compiled, or which workloads remain difficult. A fuller result reports the circuit family, mapping rules, number of instances and shots, confidence, and any mitigation.

  • Match benchmark circuit shape to the target workload.
  • Inspect native connectivity and compilation overhead.
  • Separate calibration data from end-to-end execution results.
  • Include queue, control, and repetition time only when the claim concerns throughput.
  • Do not combine vendor-defined scores as if they shared one protocol.

Build a comparison table with honest blanks

Use one row per device and columns for the same protocol, software version, calibration window, circuit set, and success measure. Mark data unavailable when a provider does not publish a comparable value. Converting a missing measure from an unrelated benchmark creates precision without equivalence.

The best conclusion may be conditional: device A performed better for the tested circuit family under a particular compilation and date, while device B exposes capabilities the test did not measure. That statement is more useful for selecting an experiment than a universal rank.