An approver has 35 minutes. We asked a quantum computer which AI agent actions deserve them.
Choosing which held agent actions a human reviews in 35 minutes is a knapsack problem, and the quantum sampler learnt the budget but never the best queue.
CISO AI · 9 Oct 2026 · Analysis

Every serious design for governing AI agents has a step where a human decides. The agent wants to approve a $48,000 purchase order, release a payment, delete a folder of ledger exports, send a message to all staff, or spawn twenty copies of itself, and policy says a person must say yes first. The gateway parks the call, posts a card to a console, and waits. If nobody opens the card before it times out, the call is refused.
That last sentence is where the security problem hides. An approver on a shift has an hour, or 35 minutes between meetings, and a queue that is longer than that. Whatever they do not get to is refused by default, which is safe, and also means the business grinds while the riskiest actions sit unread behind trivial ones. So which cards should the approver open? Each has a risk weight and takes a different time to review properly. Cover the most risk within the minutes available.
That is the knapsack problem, the one every quantum optimisation deck uses as its worked example. In our two previous pieces we ran a tool-splitting problem and a convoy problem on real IBM hardware. This time we ran the approver's queue on IBM's ibm_fez, on IBM's and Rigetti's simulators, and on Quantinuum's noisy emulator of its H2-1 trapped-ion machine, and checked every answer against the true optimum. The comments are open at the bottom.
The queue
Nine cards, each with a risk value from 4 to 9 and a review time from 5 to 20 minutes:
| Card | What the agent wants to do | Risk value | Minutes |
|---|---|---|---|
| 1 | Approve a $48,000 purchase order | 8 | 15 |
| 2 | Release a $12,000 payment to a vendor created this week | 7 | 10 |
| 3 | Delete 212 files from the ledger exports folder | 9 | 5 |
| 4 | Spawn 20 worker agents under a planner | 6 | 10 |
| 5 | Send a message to all staff from the HR agent | 5 | 5 |
| 6 | Post a draft press release to the public site | 4 | 10 |
| 7 | Create a new supplier with bank details supplied by email | 6 | 15 |
| 8 | Run a shell command on the build runner | 7 | 20 |
| 9 | Export 5,000 contact records to a file store | 8 | 15 |
The approver has 35 minutes. The best possible queue covers a risk value of 29: the ledger deletion, the purchase order, the payment and the all-staff mail, which fills the 35 minutes exactly. There is one other queue that also scores 29. A sensible rule of thumb, take the cards with the most risk per minute first, scores 27. A laptop finds the optimum by checking all 512 possibilities in a millisecond, and the standard dynamic-programming method would do it for thousands of cards in the same time. Nobody needs a quantum computer for this, and that is the point of running it: when the answer is known, you can see exactly what the quantum machine does and does not do.
What the quantum computer sees
The problem goes onto qubits one card per qubit, plus a wrinkle that the vendor decks skip. "Within 35 minutes" is an inequality, and the quantum formulation can only express equalities, so the budget is padded with three extra qubits that absorb any unused minutes. Five cards need eight qubits; nine cards need twelve. Those slack qubits do no useful work for the approver. They are bookkeeping, and they are the first cost of a constrained problem.
The second cost is that the budget penalty connects every qubit to every other. The circuit below is the five-card problem with one layer of the algorithm. Each vertical bar is a two-qubit gate between two cards, or between a card and a slack qubit, and there are 28 of them before anything else happens. Compare the tool-splitting circuit in our earlier piece, where a tool only touched the few tools it conflicted with.

We ran four versions: 5 cards and 7 cards with two layers of the algorithm, and 9 cards with one layer and with two. Angles were tuned on a classical computer first, as practitioners do. Each was sampled 4,096 times on the simulators and 600 times on the Quantinuum emulator. The 5-card and the one-layer 9-card versions also ran on ibm_fez, IBM's 156-qubit Heron processor, 4,096 samples each, 4 seconds of quantum processor time. Every sample was decoded into a queue, checked against the budget, and scored. A sample that overran the 35 minutes scored zero, because it is not a queue the approver can use.
What the figures show
| Cards, layers | Qubits | Two-qubit gates, ions / IBM Heron | Risk covered, typical: perfect simulator / real ibm_fez / Quantinuum emulator / random | Fits the budget: simulator / ibm_fez / random | Exactly the best queue: simulator / ibm_fez / random |
|---|---|---|---|---|---|
| 5, two | 8 | 54 / 199 | 15.8 / 15.9 / not run / 14.8 | 96% / 92% / 91% | 2.6% / 3.3% / 2.9% |
| 7, two | 10 | 90 / 368 | 14.0 / not run / 13.9 / 9.6 | 92% / not run / 55% | 0.7% / not run / 1.2% |
| 9, one | 12 | 66 / 242 | 11.0 / 9.3 / 10.5 / 3.9 | 72% / 51% / 22% | 0.7% / 1.0% / 0.4% |
| 9, two | 12 | 132 / 464 | 8.3 / not run / not run / 3.8 | 83% / not run / 22% | 0.1% / not run / 0.3% |

Three things stand out.
The quantum sampler learnt the budget. At nine cards, 72% of the perfect simulator's answers fitted the 35 minutes against 22% for random guessing, and its typical answer covered nearly three times the risk. On the real chip it was 51% and two and a half times. That is real work. The penalty for overrunning the budget is the biggest term in the problem, and the algorithm found it.
It never learnt the best answer. The share of samples that were exactly the optimal queue was no better than random guessing at any size, on any machine: 0.7% against 0.4% at nine cards on the simulator, 1.0% on the chip, 0.7% against 1.2% at seven, within noise either way. The best of thousands of samples was the optimum every time, as it was for random guessing. A typical quantum answer covered 11 of a possible 29. The simple rule covers 27 every time.

More depth made it worse. Two layers at nine cards covered less risk than one, 8.3 against 11.0, on a perfect simulator. The classical tuning minimises the penalised cost, which rewards fitting the budget more than it rewards covering risk, and the deeper circuit found a better way to fit and a worse way to cover. The thing a vendor's tuning loop optimises is not always the thing you are paying for.
Two chips, and a result we did not expect
On IBM's Heron processor the nine-card circuit compiles to 242 two-qubit operations for one layer and 464 for two, because the chip's qubits only talk to their neighbours and every long-range connection costs extra swap gates. On Quantinuum's trapped-ion architecture every qubit can talk to every other, so the same circuit needs 66, and the emulator of that machine stayed within half a point of the perfect simulator. Our convoy study, at 358 operations on the IBM chip, had been indistinguishable from random guessing, so we expected the nine-card run to drown.
It did not. At 242 operations the real chip still fitted the budget half the time and covered two and a half times the risk of guessing. The thing the algorithm learnt, stay under 35 minutes, is a coarse feature of the problem that noise blurs but does not erase, where the exact best queue is a fine feature that even the noiseless simulator never found. The lesson from our earlier pieces, that gate count decides what survives, still holds. What survives is not one number. Coarse structure survives far more gates than precise answers do, and a vendor can truthfully report either.
The honest version of the architecture comparison is this. Dense problems, the kind with a budget or a capacity, suit ion machines better than superconducting ones, and the emulator result is evidence for it. But it is an emulator with a published noise model, not the machine, and even the perfect simulator did not find the optimum. Better hardware would make a noisy answer cleaner. It would not make the algorithm learn something it did not learn with no noise at all.
What this means for a CISO
Approval queues need a priority order, and the inputs for one already exist. Risk tier, amount, blast radius and the time a card takes to review are all known at the moment the card is posted. Showing cards in posting order, which is what most consoles do, is the worst of the options above. Even the rule of thumb gets 27 of 29.
Timeouts are a security control with a cost. Refusing what nobody reviewed is the right default. It also means an unprioritised queue silently refuses the wrong things. Measure what times out, by risk.
When a vendor demonstrates knapsack on a quantum computer, ask three things. How many of the qubits are slack? How many two-qubit operations did the circuit compile to on the chip they used? And what did the sampler's typical answer score against the simple rule, not just whether its best sample ever hit the optimum. Thousands of random guesses hit it too.
Where do you stand?
Sign in below with an email address and tell us:
- How is your approval queue ordered today? By arrival, by risk, or by whoever shouts loudest?
- What should time out? If an approver cannot get to everything, should the unreviewed calls be refused, deferred, or approved under a limit?
- Would you take a solver's queue? If a system proposed the four cards above and skipped the other five, would an approver follow it?
- Ions or superconductors? The all-to-all machines ran this problem with a quarter of the gates. Does that change which vendor's roadmap you read first?
We will read every comment and reply to the ones that argue.
Sources: Our results are from runs on 8 and 9 October 2026 with Qiskit on IBM's ibm_fez processor (two circuits, 4,096 samples each), on IBM's ideal simulator and its published noise model for ibm_fez, and on Rigetti's QVM and Quantinuum's H2-1 emulator through Azure Quantum, using one- and two-layer QAOA with classically tuned angles, compared with brute force, a greedy rule and random guessing. Gate counts are from compiling for each machine. Our hardware runs of the tool-splitting problem are in Least privilege for AI agents is an optimisation problem, and of the convoy problem in The Army put its convoys on a quantum computer.
Comments
No comments yet.