Pricing…Open Lab
Catalog — Metric definition

Two-qubit gate fidelity

Two-qubit gate fidelity is the chance that a gate acting on two qubits does exactly what it should. A fidelity of 99.5% means an error rate of about 0.5% per gate. It is the most important number on most spec sheets, because two-qubit gates cause most of the errors in today's quantum computers.

What is gate fidelity in quantum computing?

A gate is one step in a quantum program. It changes the state of one or more qubits. A qubit is the quantum version of a bit: the basic unit that holds information. Fidelity is a score from 0% to 100%. It says how close the real gate came to the perfect gate the math describes.

The flip side of fidelity is the error rate. You get it by taking fidelity away from 100%. So 99.5% fidelity is a 0.5% error rate. Spec sheets often write this as 5×10⁻³ or 5E-3. All three mean the same thing: about 5 mistakes per 1,000 gates, on average.

A two-qubit gate (a "2-qubit gate") acts on two qubits at once. It is the only kind of gate that can link qubits so their results depend on each other. That link is called entanglement. Almost every useful program needs many of these gates. Common names are CNOT, CZ, ECR, iSWAP and the Mølmer–Sørensen (MS) gate. Each machine has its own "native" two-qubit gate, the one its hardware does directly. Our chapter on native gate sets explains why they differ.

An everyday example. Think of a photocopier that makes a copy of a copy of a copy. Each copy is 99.5% as sharp as the one before. One copy looks perfect. But after 100 copies, the page is clearly blurry. Gate errors add up the same way. That is why a tiny change in fidelity matters so much.

Where the example breaks. A photocopy only fades. A quantum error can do other things. It can flip a qubit, shift its phase (the hidden timing that makes interference work), or push it out of the 0/1 states entirely. That last one is called leakage. Some errors can also add up faster than a simple fade, or partly cancel. And fidelity is an average over many possible inputs. A gate can do worse than its average on the input your program happens to use.

Why does two-qubit fidelity matter more than single-qubit fidelity?

On current hardware, two-qubit gates are much worse than one-qubit gates. They take longer, and they need two qubits to interact without disturbing their neighbours. Our record for ibm_brisbane shows this. It lists a median two-qubit (ECR) error of 8.335E-3 and a single-qubit (SX) error of 2.669E-4. Both come from IBM's own calibration data, as reported in third-party papers. (The first is a median and the second is labelled a mean, so treat the gap as a rough size, not an exact ratio.)

So when you look at a compiled program, count its two-qubit gates first. That count, with the error per gate, tells you most of whether the result will be signal or noise.

What methods are used to measure quantum gate fidelity?

You cannot run a gate once and check it. Measuring gives only a 0 or a 1, and measuring makes its own mistakes. So scientists use repeated tests. These are the main ones.

  • Randomized benchmarking (RB). Run a long random chain of gates, then one last gate that should undo the whole chain. With perfect gates, you always get back to the start. With real gates, you get back less often as the chain gets longer. The speed of that drop gives the average error per gate. A key strength: mistakes in setting up and reading the qubits change the starting level of the curve, not its slope. So RB separates gate errors from readout errors. The gates used are a special set called Clifford gates. A two-qubit Clifford is built from several native gates, so the result is converted to a per-gate number. See Knill et al., 2007 and Magesan, Gambetta and Emerson, 2011.
  • Interleaved randomized benchmarking (IRB). Run normal RB. Then run it again with your one chosen gate slipped in after every random gate. Compare the two drop rates. The difference is the error of that one gate. This is how you test a specific gate, not an average. See Magesan et al., 2012.
  • Direct randomized benchmarking. A variant that builds its random chains more directly from the machine's native gates. IonQ uses this name for its two-qubit figures.
  • Cross-entropy benchmarking (XEB). Run random circuits. A normal computer works out the ideal odds of each answer. Then you check how well the real machine's answers match those odds. A closer match means higher fidelity. Google used this on Sycamore. See Arute et al., 2019.
  • Layer fidelity and error per layered gate (EPLG). Instead of testing one pair of qubits alone, run gates on many pairs at the same time, in a long chain. This catches crosstalk: one gate disturbing its neighbours. It usually gives worse numbers than testing one pair alone, and those numbers are closer to what a real program sees. See IBM's layer fidelity paper (2023).

Each method answers a slightly different question. RB gives an average over many gates. IRB targets one gate. XEB uses random circuits. Layered tests include crosstalk. None of them is "the" true fidelity.

Why can't you compare fidelity numbers from different vendors directly?

Two fidelity figures are not directly comparable unless four things match. Spec sheets rarely make all four match.

  1. Aggregation: best pair or typical pair? A chip has many qubit pairs, and each has its own fidelity. A median is the middle pair. A mean is the average. A best figure is the single luckiest pair. A best-pair number can look many times better than the median on the very same chip.
  2. Method. RB, interleaved RB, XEB and layered benchmarks measure different things. A layered number that includes crosstalk is expected to be worse than an isolated-pair RB number.
  3. Alone or together. A pair tested while the rest of the chip is quiet ("isolated") usually scores better than the same pair tested while everything runs ("simultaneous").
  4. Date. Machines are re-tuned (calibrated) often, and their numbers drift. A figure from two years ago describes a different machine state.

The gate type matters too. A CZ gate and an MS gate have very different durations and failure modes.

Our rule: before comparing two numbers, check the method, aggregation and date on each record. If they differ, the comparison needs a warning. Our comparison tool generates that warning for you.

What does 99.5% fidelity mean over 100 gates?

Here is a simple rule of thumb. If each gate works with a chance of 0.995, the chance that every gate works is 0.995 multiplied by itself once per gate.

  • Two gates: 0.995 × 0.995 = 0.990025. About 99.0%.
  • Ten gates: 0.995 multiplied ten times ≈ 0.951. About 95%.
  • One hundred gates: 0.995¹⁰⁰ ≈ 0.606. About 61%.

Now change the fidelity and keep 100 gates:

  • 99.0% per gate: 0.99¹⁰⁰ ≈ 0.366. About 37%.
  • 99.9% per gate: 0.999¹⁰⁰ ≈ 0.905. About 90%.

Look at the jump. Going from 99.0% to 99.9% sounds like less than one point. But over 100 gates it takes you from about 37% to about 90%. That is why vendors fight over each extra "9".

What this estimate leaves out. It assumes errors are independent and that nothing else goes wrong. Real runs also have single-qubit errors, readout errors, idle time and crosstalk. So treat the result as a rough guide, not a prediction. But the direction is right: more two-qubit gates means less signal.

How do you read two-qubit fidelity on a QPU137 hardware page?

Each figure on our QPU pages is a sourced record. It carries the value, the method, the aggregation, the source and the date. Some vendors publish fidelity (like 99.5%). Others publish error (like 5E-3). We keep the vendor's own form, so check the unit. These three records show why the labels matter:

  • IBM Boston lists two vendor-reported figures for one chip. "2Q error (best)" is 6.43E-4: the lowest error on any single edge. "2Q error (layered)" is 2.54E-3: the average error per layered gate across a chain of 100 qubits. They measure different things. Neither is wrong.
  • IonQ Forte has a vendor-reported, peer-reviewed spread of errors over all 435 pairs of a 30-qubit chain: a median of 46.4e-4 (0.464%), a best pair of 27.8e-4 and a worst pair of 885e-4. The worst pair is more than 30 times worse than the best.
  • Google Sycamore reports a vendor-reported, XEB-based two-qubit error of 0.36% isolated and 0.62% simultaneous. Same chip, same method, but running everything at once makes the error larger.

To go deeper, read the chapter on fidelity, error and calibration. Or see how fidelity combines with readout fidelity and coherence time to limit what a machine can run.

See it in the data: the sourced catalog · compare two processors · lesson: reading hardware specs