June 16, 20267 min read

Which Agent Should Go First? Reversibility, Evidence, and Exception Cost

QD

By Equipo Quantum Developers

Physical four-quadrant matrix with blue, green, and yellow tokens as a person moves one token in front of an analytics dashboard.
Share

Operating thesis

The highest-volume process is not always the best first agent. Its data may be ambiguous, its actions irreversible, or its exceptions dependent on tacit expertise. The thesis is more selective: a candidate deserves priority when its action can be reversed, its decision can be tested against available evidence, and the current cost of resolving exceptions is material.

Together, those conditions produce learning without wagering the entire operation. Reversibility bounds harm; evidence distinguishes a correct decision from a plausible one; exception cost creates economic value even when the agent does not automate everything. NIST calls for understanding context, impacts, human roles, and deployment conditions before managing risk in the AI RMF Core. Prioritization is not an ordered list of attractive ideas. It is a decision about where the enterprise can learn responsibly.

Keep the three axes distinct

Reversibility asks whether an action can be undone within bounded time and cost. Coding an invoice is more reversible than releasing a payment. Proposing a route is more reversible than terminating a contract. The window matters too: a quote may be edited before sending but not after customer acceptance.

Evidence asks whether inputs, policy, decision, and outcome can be reconstructed. A documented process does not guarantee evidence; rules may live in messages or the outcome may arrive months later. Strong evidence includes stable identifiers, current sources, policy version, required approval, and a later signal confirming or contradicting the decision.

Exception cost captures all work generated by cases outside the standard path: analysis, waiting, coordination, rework, and risk. A process with few exceptions can still be poor if those exceptions are cheap. A process already highly automated may be excellent when the remainder consumes expert hours and causes delay.

Apply vetoes before scoring

Do not average every concern away. Apply vetoes first:

  • an irreversible action with material harm is ineligible for initial autonomy;
  • an unobservable outcome cannot support quality validation;
  • a mandatory approval cannot be replaced by agent confidence;
  • a process without an operating owner has nobody to accept exceptions or narrow scope;
  • data with unproven origin or permissions invalidates the pilot.

The Green Book 2026 incorporates uncertainty, options, and flexibility into appraisal. In this setting, preserving an exit is as important as estimating a benefit. A two-way-door choice can be tested and reversed; a one-way-door choice requires a different gate.

The reversibility, evidence, and exception matrix

After vetoes, use an internal one-to-five scale. It is neither a benchmark nor a probability. Its purpose is to expose disagreement:

Axis 1 3 5
Reversibility harm cannot be repaired reversal needs coordination rollback is simple, fast, and tested
Evidence outcome is unobservable sample or proxy is verifiable complete chain and frequent outcome
Exception cost residual and cheap recurring moderate effort expensive, slow, or risky queue
Control readiness no owner or boundary review partly defined owner, thresholds, and degraded mode
Learning value little transferable learning tests a common capability unlocks several future decisions

Do not sum the columns mechanically. Plot candidates with reversibility on one axis, evidence on another, and exception cost as size or color. Control readiness can be an eligibility label. Learning value breaks ties.

AWS recommends recording priorities, benefits, risks, and accountable owners when evaluating tradeoffs in its Well-Architected operational excellence guidance. That is the matrix’s real value: preserving reasoning, not manufacturing an objective number.

Three illustrative candidates

Invoice coding. The agent proposes a ledger account and cost center, while a person confirms cases not covered by policy. The proposal is reversible before posting, evidence exists in the purchase order, invoice, and policy, and exceptions consume analysis. It can be a strong learning candidate when mandatory approval remains intact.

Payment release. Volume may be attractive, but the action moves money and recovery may be difficult. Even with good evidence, the first mode should be recommendation, never initial autonomous decision. A high volume score does not erase a consequence veto.

Shipment incident classification. Routing an incident is reversible, and evidence may include the order, carrier event, and service level. Yet value is low if the current team corrects the queue in seconds. The candidate rises when exception cost includes waiting, penalties, or repeated coordination, not simply because events are numerous.

These examples do not assign universal values. The committee must test every claim using its own process data.

The decision artifact

Each candidate card should contain:

  • business object and exact action;
  • the point through which the action remains reversible;
  • required evidence and time until outcome is known;
  • exception types, frequency, and handling cost;
  • maximum consequences and applicable veto;
  • person who approves, handles, and can stop;
  • degraded mode and rollback procedure;
  • learning metric and next gate.

In Quantum Automation Center, the card can link the workflow, runs, artifacts, approvals, and timeline. A pilot should not start until an observer can associate every decision with its candidate and version. Results can then recalibrate the matrix rather than merely decorate a status report.

Portfolio decisions

A portfolio needs controlled diversity. An initial group might include one direct-value candidate, one that tests a reusable integration, and one that improves evidence for later workflows. Avoid shared hidden dependence: three pilots relying on the same unstable source do not diversify risk.

Revisit priority at each gate. If exceptions fall because policy changed, the economics changed. If outcomes become observable more frequently, an uncertain candidate may rise. If rollback has never been tested, claimed reversibility should fall. The matrix is a living hypothesis, not an annual ranking.

The strongest counterargument

The matrix can reward safe but marginal tasks and postpone the difficult processes where the largest transformation sits. A portfolio full of reversible recommendations may produce polished demonstrations without changing a material outcome.

The criticism is valid. Reversibility does not mean lack of ambition; it means sequencing. A high-impact process can begin in recommendation mode, advance to approval-bound decisions, and only then reach bounded autonomy. Learning needs an explicit path to value. A safe candidate that neither tests a reusable capability nor moves an outcome should not be prioritized either.

When not to use this approach

Do not use an aggregate score to authorize irreversible actions, high-impact human decisions, or a process with no owner able to accept residual risk. Those cases require specific assessment, formal controls, and often mandatory human judgment.

Do not use the matrix when an obligation already fixes priority or an outage requires immediate recovery. Use it to compare genuine opportunities in a portfolio. Its best output is not a declared winner. It is a visible explanation of why an idea can learn quickly, how it can be stopped, and which evidence would justify broader autonomy.

Sources