Ninety Days to Test an Agent, Not to Promise Autonomy
By Equipo Quantum Developers

Summarize:
Operating thesis
Ninety days are not a deployment promise. They are a window for reducing uncertainty in three stages: define and observe, test without autonomy, and close real cases. The calendar only organizes learning. An agent does not advance to production action until the team sees closures, resolved exceptions, and verifiable outcomes in its eligible population.
NIST calls for AI systems to be tested before deployment and monitored in operation, with documented metrics, limitations, and conditions of use in the AI RMF Core. A demonstration measures behavior on selected examples. The program measures whether the organization can govern decisions when real data, people, and dependencies intervene.
Before day one: choose a question, not a department
The sponsor should state a decision that can change at the end: expand an assisted mode, authorize a reversible action, redesign the process, or stop. “Explore AI in logistics” is not a question. “Can an agent classify shipment exceptions with enough evidence to reduce waiting without increasing incorrect diversions?” is one.
Compare candidates through five filters:
- outcome observable inside the window;
- identifiable population and source;
- reversible initial action;
- owned, costly exceptions;
- current alternative that serves as reference.
Choose one, not a list. If several candidates share a critical dependency, testing that dependency can be a deliverable, but it does not make five pilots.
Days 1–30: contract, baseline, and historical closures
The first stage does not build a polished interface. It produces the evidence contract:
| Deliverable | Question it must answer |
|---|---|
| decision card | which object, action, and population are in scope? |
| baseline | how does work perform today and with which denominators? |
| closure outcome | which event confirms success, error, or harm? |
| exception catalog | what can the agent not decide, and who receives it? |
| data contract | which source, version, freshness, and permission apply? |
| action boundary | what remains recommendation or approval? |
| evaluation plan | which comparison and coverage support a decision? |
The UK Treasury Magenta Book distinguishes process and impact evaluation and emphasizes the counterfactual for attribution. During this stage, the team chooses a parallel cohort, staged introduction, or another proportional comparison. It does not wait until the end to select the metric that moved favorably.
Use historical cases with known outcomes to discover classes, but do not confuse that exercise with production. By day thirty, there should be an exclusion reason, data owner, and closure definition. If the outcome cannot be observed, change the case or stop the program.
Days 31–60: shadow, assistance, and operating test
The agent works on current inputs without acting first. Every proposal links to the object, source, version, rule, and later outcome. A reviewer records more than correct or incorrect: reason, severity, and missing evidence.
After enough shadow coverage, the agent may assist a person on a narrow population. Assistance means the human keeps decision authority and access to context; it is not automatic approval in disguise. Measure:
- coverage of cases with observable outcomes;
- quality by type and consequence;
- exception time and cause;
- agent–reviewer disagreement;
- complete evidence per closure;
- technical and human cost per case;
- review queue and age.
Test failures too: stale source, unavailable tool, insufficient approval capacity, and uncertain outcome. AWS describes Operational Readiness Reviews as a question-driven mechanism that uses incident lessons to remove known causes of impact. Here, the review uses findings from the pilot itself rather than a generic checklist.
The day-sixty gate does not ask, “Did the model work?” It asks whether operations can detect, contain, and reconstruct a decision.
Days 61–90: close the loop and choose the next level
The final stage keeps population stable so cases can mature to outcome. Expanding volume to show momentum is tempting. Resist it. Without closure, expansion only creates more unverified decisions.
A bounded, reversible action may be enabled only when earlier closure evidence exists, the owner approves, and rollback has been tested. The goal is not full autonomy. It is to demonstrate the chain:
input → proposal → review or action → operating outcome → closed exception → evidence.
Compare with baseline, identify segments where outcomes differ, and calculate full cost. Unobserved benefits remain hypotheses. At day ninety, four decisions are legitimate:
- stop because the hypothesis or data failed;
- redesign and repeat one stage;
- retain assistance because it adds value without justifying autonomy;
- authorize bounded autonomy for a proven population.
“Scale” without a population, boundary, and next gate is not a decision.
Artifact: the 30/30/30 board
The executive board uses one row per evidence class:
| Evidence | 1–30 | 31–60 | 61–90 | Owner | Decision |
|---|---|---|---|---|---|
| population | defined and profiled | observed live | segment proven | operations | retain or restrict |
| outcome | closure event | initial coverage | mature comparison | business | value or no value |
| risk | consequences and limits | failures rehearsed | residual risk | control | accept or stop |
| operation | owner and runbook | alert and escalation | recovery proven | technology | ready or not |
| economics | baseline and cost | cost per case | realized value | finance | fund or close |
Do not use universal percentages. Every threshold reflects consequence, volume, and review capacity.
In Quantum Automation Center, catalog, states, timelines, artifacts, logs, analytics, permissions, and approvals can assemble evidence by run. The platform does not decide whether an outcome is material; the process owner validates it.
Closure metrics that block autonomy
A closure metric occurs after action: an invoice posted without reversal, a logistics exception resolved inside its window, an accepted quote with validated margin, or a service case not reopened during the defined period. Exception closure also matters: assigned, resolved, cause recorded, and object returned to a consistent state.
If the program knows only “response generated” or “user accepted suggestion,” it still measures output or a proxy. Assistance may continue, but irreversible autonomous action is not justified.
The strongest counterargument
Waiting for closure can slow use cases whose outcomes take months. The program may finish without an executive success story even after learning that data is poor or review is too costly.
That criticism is valid. Use intermediate outcomes only when a theory of change explains their relationship to closure, and keep autonomy bounded. Negative learning is also a decision: it prevents funding a larger bet. Moving the goalpost to claim success destroys the program.
When not to use this approach
Do not use the format when outcomes arrive well after ninety days, volume cannot support an observable cohort, or an obligation requires action without an experiment. Adapt the window, use retrospective evaluation, or implement the required control without pretending to establish causality.
Use it when work is repeated, closure is observable, and the decision is reversible. The 30/30/30 discipline does not guarantee production deployment. It guarantees that deployment, if approved, did not happen because the calendar expired.
Sources
Article topics


