June 14, 20267 min read

Automation ROI Needs a Counterfactual, Not a Before-and-After Anecdote

QD

By Equipo Quantum Developers

Executive dashboard with a trend chart, donut chart, bars, and a table between two stacks of documents.
Share

Operating thesis

“We were faster after the agent” is a useful observation, but it is not proof of ROI. Volume, case mix, staffing, policy, or data quality may have changed at the same time. The testable thesis is stricter: a committee can attribute value to an agent only if it defines the unit of analysis, outcome, baseline, and a comparison approximating what would have happened without the intervention before the pilot starts.

The UK Treasury Magenta Book explains that impact evaluation seeks to identify change caused by an intervention and therefore requires an estimate of the counterfactual. The World Bank’s Impact Evaluation in Practice develops the same principle through comparison groups and causal inference methods. An enterprise automation pilot need not become an academic trial. It does need to distinguish an observed outcome from an attributable outcome.

Approve the measurement card first

Before enabling the agent, the sponsor, operations lead, and finance partner should approve a one-page measurement card:

Field Question that must be settled
Unit Is the unit an invoice, quote, case, order, or customer?
Intervention What exact action changes, and on what date?
Eligible population Which cases may the agent handle, and which are excluded?
Primary outcome Time to close, cost per case, recovery, error, or margin?
Baseline Which prior window and cleaning rules will be used?
Comparison Which cohort credibly represents operation without the agent?
Horizon When should effects appear, and how long is observation?
Costs Are build, run, review, integration, and control costs included?
Decision rule What evidence permits scale, adjustment, or stop?

This card prevents choosing, after the fact, whichever metric moved most favorably. It also stops incompatible benefits from being added together. If released time does not reduce paid hours and is not reassigned to measurable work, it is not cash savings. It may be recovered capacity, which is valuable, but should be named accurately.

A counterfactual proportional to risk

The cleanest comparison randomly assigns eligible cases between the current process and the agent-assisted process. That is not always practical. A staged rollout across teams, regions, or weeks can preserve comparable groups. Cases can also be matched on complexity, value, channel, and age. An interrupted time series may be appropriate when history is long enough and no major concurrent change contaminates the break.

No method removes all uncertainty. The design should record baseline imbalances, spillover between groups, exclusions, and concurrent changes. NIST calls for metrics, evaluation methods, limitations, and uncertainty to be documented and revisited throughout the lifecycle in the AI RMF Core. In executive terms, the number must travel with its conditions of validity.

The minimum calculation for one outcome is:

Incremental effect = change in the intervention cohort − change in the comparison cohort.

Financial ROI may then be expressed as monetized incremental benefit minus total cost, divided by total cost. That ratio is useful only if the numerator does not mix causal impact, theoretical capacity, and realized savings.

An illustrative example, not a benchmark

Consider two comparable cohorts of one hundred cases. At baseline, both take about ten hours per case. During the pilot, the agent cohort drops to seven hours and the comparison cohort drops to nine because the business also simplified a policy. A simple before-and-after story credits the agent with three hours. The difference in changes attributes about two: three hours of intervention improvement minus one hour that occurred without the agent.

Assume each recovered hour has an internal unit value and only half can be reassigned to verified productive work. The business case should not monetize the full two hours as savings. It should separate the operating effect, realization rate, and financial benefit. Every number here is fictional and demonstrates the calculation only; it is neither a promise nor an industry benchmark.

The cohort analysis can also reveal heterogeneity. The agent may improve standard cases substantially while making exceptions worse. A favorable average does not authorize expansion to populations the pilot did not represent. The defensible decision may be to scale only the evidenced segment and keep human review elsewhere.

Costs that disappear from the story

The denominator is not just technology consumption. It includes discovery, integration, data remediation, testing, observability, approvals, human review, exception handling, support, security, training, maintenance, and retirement. It also includes the opportunity cost of staff validating the pilot.

Report three layers separately:

  1. Operating outcome: the incremental change in time, error, recovery, or margin.
  2. Economic realization: the portion converted into cash, used capacity, or avoided risk.
  3. Total cost: initial investment, recurring cost, and control cost.

Avoided risk needs an explicit hypothesis: expected frequency, loss magnitude, and attributable reduction. It should not be presented as realized savings when the event was never observable or the control’s coverage is unknown.

The executive decision dashboard

In Quantum Automation Center, each run can retain a case identifier, workflow version, intervention, human review, outcome, and supporting artifacts. Those events make cohorts reproducible instead of forcing analysts to reconstruct stories from screenshots. The dashboard should show eligible and excluded populations, baseline balance, trends by cohort, incremental effect, uncertainty range, accumulated costs, and material failures.

Set the rule before results arrive: scale when the primary outcome improves without breaching an error limit and coverage is sufficient; adjust when direction is favorable but uncertainty remains high; stop when harm rises, traceability breaks, or cost per outcome crosses the agreed boundary. Those thresholds belong to the enterprise. They are not universal benchmarks.

The strongest counterargument

In a small operation, creating comparable groups can cost more than the pilot. Deliberately withholding an apparent improvement from one group may also feel unfair. If an effect is large, observable, and reversible, a staged launch with a stable baseline may be sufficient for a decision.

That objection is valid. The answer is proportionality, not methodological ceremony. Use the simplest design capable of falsifying the sponsor’s preferred story. A stepped introduction, queue-level comparison, or independent sample review can produce enough evidence without freezing operations. What is not defensible is claiming causality because two screenshots carry different dates.

When not to use this approach

Do not impose a heavy causal design when automation is a legal requirement, responds to an urgent incident, or processes too few cases for a credible comparison. In those situations, measure compliance, safety, full cost, and deployment quality without pretending to have causal precision.

Do not force a monetary ROI when the actual objective is exploratory learning. Define instead which uncertainty the pilot will reduce and what decision that learning will enable. For repeated workflows with sufficient volume and a consequential investment decision, the counterfactual is essential: it turns an apparent improvement into a claim that finance, operations, and risk can challenge using the same evidence.

Sources