The Army put its convoys on a quantum computer. We ran ours on IBM's, and checked the answers.
CISO AI · 8 Oct 2026 · Analysis
In February 2024 the Sydney company Q-CTRL published a case study about the Australian Army. Its quantum software, it said, had scheduled the convoys for Exercise Talisman Sabre, the largest combined exercise between the Australian Defence Force and the US military, and beaten a classical solver by two hours. In May this year the company went further. A white paper titled Quantum Computing for Battlefield Information Dominance projects that quantum computers will give "operational advantages" in defence logistics "as soon as 2027", and cites a new version of the convoy work: 5,000 vehicles, 50 convoys, 85 qubits.
Claims like these end up in board papers and capability business cases, so we wanted to see what one looks like up close. We built a convoy routing problem of our own, ran it on a real IBM quantum computer on 8 October, and checked every answer against the true optimum. The comments are open at the bottom, and we want to hear from anyone who has had to evaluate one of these proposals.
What the Army work was
The problem the Army brought to Q-CTRL was real: work out departure times and routes for 5,000 vehicles, grouped into convoys of ten, so that the whole force deploys in the shortest time while arriving in a set order, with convoys of different sizes and speeds sharing a handful of roads. Q-CTRL's solution split the job between a classical computer and a quantum one, ran the quantum part on IBM Guadalupe, a 16-qubit machine, and used its Fire Opal software to suppress the hardware's errors.
The case study reports three numbers. Fire Opal made finding the optimal answer 12 times more likely than running on the raw hardware. It doubled the number of convoys that fitted on the machine. And the final schedule was about two hours, roughly 10%, shorter than "a benchmark classical heuristic solver", so the last convoy arrived before midnight rather than after one in the morning. The Army's SO1 Quantum Technologies, Lieutenant Colonel Marcus Doherty, is quoted alongside:
"Optimally routing 120 convoys can take more than a month of classical computation. The Australian Army is evaluating the potential of quantum computing to provide improvements; however, it's been difficult to validate the feasibility of a quantum solution due to hardware noise."
Read further down and the case study is candid about the limits. The problem "was run at a scale where 'brute force' calculation could still be practically executed (quantum advantage was not achieved)". The 12 times improvement is quantum against quantum: the same hardware with and without Q-CTRL's software, not a quantum computer against a classical one. And the classical solver the schedule beat is never named. The 2026 update adds the 50 convoys and 85 qubits but, in the coverage we could find, no comparison with an optimum or a named classical method.
None of that makes the work dishonest. It makes it a research result being read as a capability.
What we built
Our problem is a simplified cousin of the Army's, chosen because its true answer can be checked. A group of convoys has to reach the same destination by one of two roads. Road A takes 60 minutes to drive and road B takes 90, and each convoy occupies its road for its own length of time, from 11 to 47 minutes depending on its size, so convoys on the same road queue behind each other. The task is to choose a road for every convoy so that the last one arrives as early as possible. It is a scheduling problem of the kind mathematicians call NP-hard, and it maps directly onto a quantum computer: one qubit per convoy, and an interaction between every pair.
We solved it four times, for 6, 8, 10 and 12 convoys, with QAOA, the quantum optimisation algorithm the Army work also used. Two layers, with the angles tuned on a classical computer first, as practitioners do. Each problem was sampled 4,096 times on ibm_fez, IBM's 156-qubit Heron processor, and the whole run used 7 seconds of quantum processor time. Then we compared it with three things a fair test needs:
- The true optimum, found by checking every possible assignment on a laptop.
- A simple classical rule: take the convoys biggest first and send each down whichever road currently finishes earlier.
- Random guessing, with the same 4,096 attempts the quantum computer had.
What the figures show
| Convoys | True optimum (laptop time) | Simple rule | Quantum, typical | Random, typical |
|---|---|---|---|---|
| 6 | 183 min (0.2 ms) | 190 min | 222 min | 221 min |
| 8 | 205 min (0.8 ms) | 211 min | 240 min | 245 min |
| 10 | 222 min (5 ms) | 227 min | 256 min | 262 min |
| 12 | 234 min (14 ms) | 238 min | 267 min | 275 min |

Four things stand out.
The quantum computer found the optimal schedule every time, and so did random guessing. At every size, the best of the quantum computer's 4,096 samples was the true optimum. But the best of 4,096 random guesses was the optimum too. When a problem is small enough to brute force, "found the optimal solution" is not evidence of anything, and that is exactly the scale at which the Army work was run.
Its typical answers were slightly better than chance. From 8 convoys up, the quantum computer's average schedule was 5 to 8 minutes shorter than random guessing, so the algorithm was doing real work. At 6 convoys it was no better at all.
The classical answers were far better. A rule a logistics officer could apply with a pencil came within 2 to 4% of the optimum at every size. The laptop found the optimum itself in 14 milliseconds at 12 convoys. The quantum computer's typical answer was 14 to 21% worse than optimal.
Noise ate most of the gain. The same circuits on a perfect simulator did noticeably better, and the gap is the hardware. The 12-convoy circuit compiled to 431 two-qubit operations on the chip, each with its own error rate. In our formulation every pair of convoys interacts, so 12 convoys means 66 pairs and 50 convoys would mean 1,225.

Two caveats, to be fair to Q-CTRL. We ran IBM's standard toolchain, not Fire Opal, and the company's own figures say its software improves hardware results considerably. And our problem is simpler than the Army's, which had ordering constraints and changing congestion. What our run shows is not that the Army's result is wrong. It is what the baseline looks like, and how much distance any software layer has to cover.
Precise is not the same as accurate
We ran one more test, on a different kind of problem, because it is the other half of how quantum results get oversold. Chemistry is the use case most experts expect to pay off first, so we asked ibm_fez for the energy of a hydrogen molecule, the textbook example. The first answer was 6.8 thousandths of a hartree off (the unit chemists use; useful answers need to be within 1.6), with an uncertainty almost as large, so we ran it again with far more measurements. That took nearly nine minutes of quantum processor time. The uncertainty shrank to 1.7 thousandths, but the answer settled 6.7 thousandths away from the true value, about four times its own error bar.
More samples made the answer precise without making it right. The bias comes from noise in the hardware, which basic error mitigation does not remove. When a vendor says a result needs "just more shots", ask which of the two problems more shots would fix.
How to read a quantum case study
The Army is doing what a sensible buyer of an emerging technology should: paying a little now for evidence and a roadmap, not buying a capability. The risk is in how the results travel. By the time a case study reaches a steering committee, "promising on a 16-qubit machine at brute-force scale" can easily become "quantum beat our solver by 10%". Four questions cut through it:
- Could the problem have been brute forced? If yes, the only meaningful comparison is with the true optimum, and the case study should report it.
- Better than what? A named, state-of-the-art classical solver, or an unnamed heuristic? Against the raw quantum hardware, or against a classical computer?
- What does random guessing score? With thousands of samples on a small problem, guessing finds good answers too.
- What happens at twice the size? Count the two-qubit operations. Noise grows with them, and they grow faster than the problem.
Where do you stand?
This piece is meant to start an argument, and we would rather hear from you than from ourselves. Sign in below with an email address and tell us:
- Should Defence fund quantum computing pilots now, while the results are at brute-force scale, or wait for error-corrected machines and spend the money on things that work today?
- Would "10% better than a heuristic" survive your business case? If a vendor brought you that number, what would you ask for before believing it?
- Is 2027 credible? Q-CTRL projects quantum advantage for defence logistics between 2027 and 2029. Is that a forecast, a roadmap or a sales line?
- Computing, sensing or navigation? Q-CTRL's quantum navigation work is much closer to fielded capability than its computing work. Which quantum technology reaches the ADF first?
We will read every comment and reply to the ones that argue.
Sources: Q-CTRL, Improving Army logistics with quantum computing, first published February 2024, including the quotation from LTCOL Marcus Doherty. Q-CTRL, Q-CTRL Defines the Path to Quantum Battlefield Information Dominance, 28 May 2026, and Quantum Computing Report's coverage of the white paper. Our results are from runs on IBM's ibm_fez processor on 8 October 2026 with Qiskit, using two-layer QAOA with classically tuned angles and 4,096 samples per problem; the hydrogen results use IBM's Estimator with readout error mitigation. The simulator comparisons use IBM's published noise model for the same chip.
Comments
No comments yet.