June 15, 20267 min read

The AI-Agent Business Case as an Evidence Packet: Ranges and Kill Criteria

QD

By Equipo Quantum Developers

Tablet showing a portfolio dashboard beside a printed report with charts and a notebook.
Share

Operating thesis

A business case is not a slide carrying one large savings number. It is an evidence packet that lets a later decision-maker determine whether the hypothesis survived contact with operations. The practical thesis is that a defensible case fits on one page because it links every benefit to a baseline, source, and test, and states in advance which result requires stopping.

The UK Treasury Green Book 2026 calls for comparing options with business as usual, accounting for uncertainty, and planning evaluation. The GAO Cost Estimating and Assessment Guide emphasizes documented source data, assumptions, sensitivity, and risk. Neither treats an agent as a magical category. Both force the author to show where a number came from and how it might fail.

The reusable one-page template

The template can travel across finance, sales, logistics, or service, provided that industry-specific evidence is not erased:

Block Verifiable content
Decision requested amount, owner, date, and next gate
Problem population, frequency, baseline, and cost of inaction
Scope permitted actions, exclusions, and affected systems
Evidence source, date, quality, and owner for each assumption
Options business as usual, non-AI alternative, and agent
Benefits low, central, and high range; realization mechanism
Costs build, run, people, control, and retirement
Risks possible harm, control, early signal, and owner
Test cohort, outcome, horizon, and evidence threshold
Exit criteria to stop, narrow, or retire

The options block matters. If a deterministic rule, form redesign, or policy change solves the problem with less risk, the agent should compete against it. “Do nothing” is not zero. It represents continuing operations with their expected costs and risks.

Replace a point estimate with a defensible range

A point benefit is brittle because it combines volume, adoption, accuracy, time released, and economic realization. Write the relationship explicitly:

Realized benefit = eligible volume × incremental improvement × effective adoption × realization rate × unit value.

Every factor needs a source and range. Volume may come from the system of record; improvement from a comparative pilot; adoption from observation; realization from a capacity plan; unit value from finance. The low case should not be an arbitrary haircut to the central case. It should represent a coherent set of conditions that could occur together.

Costs also need ranges. Uncertain integrations, human review, exceptions, observability, and maintenance tend to expand after a pilot. GAO recommends sensitivity analysis to identify which assumption drives the estimate. That variable should receive an early test. If the case depends almost entirely on reducing review that has not yet been proven safe, review reduction is the dominant economic risk.

Illustrative example: a reconciliation agent

Consider a monthly reconciliation process. The baseline records eligible volume, analysis minutes per case, rework, and supervision cost. The agent proposes matches and assembles evidence, but does not post. The low case assumes partial adoption and full review; the central case assumes stable adoption and sample review for low-risk cases; the high case expands eligibility only after quality is demonstrated.

The figures themselves do not matter here and are not benchmarks. The important feature is structure: every scenario keeps the same formula, changes named assumptions, and separates cash savings, redeployed capacity, and avoided risk. The committee can fund a small pilot without committing to the high case.

The next gate might require four pieces of evidence: sufficient outcome coverage, material error below the internal-control limit, exception time no worse than the current process, and total cost per case inside the approved range. A missing condition should not be “offset” with a favorable secondary metric.

Define kill criteria before launch

Exit criteria protect capital and reputation. They must be observable and owned. Examples include:

  • stop autonomous execution when a material action lacks evidence or approval;
  • narrow eligibility when harm clusters in one segment;
  • return to recommendation mode when outcomes cannot be verified;
  • close the pilot when a critical data source fails the minimum quality bar;
  • retire the solution when recurring cost exceeds its boundary without attributable improvement;
  • pause when a changed obligation invalidates the control design.

NIST calls for mechanisms to prioritize risk, respond, recover, and deactivate systems when needed in the AI RMF Core. A kill criterion is not pessimism. It is the operating expression of that management capability.

Evidence changes by industry

The template is common; the proof is not. Accounts payable depends on duplicates, approvals, and segregation of duties. Sales depends on margin, price validity, and commercial exceptions. Logistics depends on timeliness, expediting cost, and external events. Service depends on resolution, reopen rates, and customer harm. Every case should use an outcome the function already recognizes, not a metric invented because the automation emits it conveniently.

In Quantum Automation Center, the decision card can link the workflow, versions, runs, approvals, artifacts, and exceptions. The committee can then revisit the same decision object at each gate. The platform does not replace financial validation or accountable ownership. It preserves the chain between hypothesis, operation, and outcome.

Review through gates

A sensible sequence separates discovery, assisted test, bounded autonomy, and scale. Each gate re-estimates cost and benefit using new evidence. Ranges should narrow; if they widen, the document should explain why. The choice is not only continue or cancel. Scope can change, a control can be strengthened, or the non-AI option can win.

The sponsor owns the outcome; finance owns realization and cost; operations owns baseline and capacity; risk owns limits; technology owns feasibility and operational evidence. Approval never transfers accountability to the agent.

The strongest counterargument

This discipline can slow a cheap experiment and create a false appearance of precision around technology that changes rapidly. A team may spend weeks modeling uncertainty before learning whether the data is usable.

That objection is valid when spend and impact are reversible. The template should be proportional: wide ranges, a short test, and a small decision. Exploration does not require an exhaustive financial model. It still requires a statement of what the team intends to learn, how much can be lost, and when it will walk away. Speed improves when the committee does not renegotiate success conditions after every result.

When not to use this approach

Do not use this template as a financial contract for open-ended research with no stable workflow, observable baseline, or defined investment decision. Do not use it to justify a regulatory obligation that must be met even with negative ROI; compare compliance routes on cost, risk, and time instead.

Use it when capital, operating change, or autonomy is at stake. Its purpose is not to predict the future precisely. Its purpose is to leave a falsifiable claim, an honest range, and a safe exit so enthusiasm does not become a permanent obligation.

Sources