July 9, 20266 min read

No Catalog Agent Is “Ready” Without Data, Control, and Ownership Maturity

QD

By Equipo Quantum Developers

Four ascending levels show, on every card, icons for data access, protected control, and a person.
Share

Operating thesis

“Ready to deploy” compresses at least five questions: are data authorized, can the decision be validated, do controls work, can operations respond, and does someone own the outcome? A responsible catalog does not average those answers. Readiness equals the weakest critical dimension: a strong demo cannot compensate for unauthorized data, untested controls, or ownerless operations.

GAO organizes its AI Accountability Framework around governance, data, performance, and monitoring. NIST calls for inventory, clear roles, evaluation in deployment conditions, and continuous management in the AI RMF Core. Both contradict the idea that “agent” is a finished product selected from a shelf without context.

Treat the catalog as capability inventory, not a list of names

A useful card begins with object and decision:

  • object: invoice, shipment, quote, order, or case;
  • population: eligibility and exclusions;
  • current mode: exploration, shadow, assistance, or action;
  • outcome: event confirming success or harm;
  • data: sources, rights, freshness, and provenance;
  • control: boundary, approval, permission, and rollback;
  • ownership: owner, on-call, budget, and review date;
  • evidence: coverage, exceptions, changes, and attribution level;
  • next entry: test required to advance;
  • expiry: point when the assessment stops being current.

Do not use “finance agent” as the unit. It may contain reversible classification and irreversible payment release. Those need separate cards.

Five verifiable levels

Level 0: bounded concept

A problem, object, proposed action, and exploration owner exist. No readiness claim exists. Entry requires purpose, population, non-AI alternative, data boundary, and expiry.

Exit: continue discovery or archive.
Permitted action: none on real work.

Level 1: observable evidence

Authorized, versioned sources, historical cases, closure outcome, and baseline exist. The team knows exclusions and can join input to object. Entry requires data rights, provenance, minimum quality, and a verifiable outcome.

Exit: shadow test.
Permitted action: reading from a copy or isolated flow.

Level 2: validated assistance

The agent processes current inputs without autonomous decision. Coverage, consequence-weighted quality, disagreement, exceptions, and review cost are observed. The human receives context and keeps authority.

Entry: shadow evidence, decision contract, narrow population, and exception route.
Permitted action: recommendation; the person executes.

Level 3: bounded action

A reversible action executes for a proven population. Least privilege, approval where required, idempotency, rollback, and reconciliation exist. SLOs include quality, escalation, and evidence.

Entry: closed outcomes, tested control, accepted residual risk, and degraded mode.
Permitted action: only the verbs and boundaries in the contract.

Level 4: operating capability

The agent has an outcome owner, on-call path, runbook, budget, change management, monitoring, and retirement rule. Evidence follows the object to outcome; periodic review confirms that the level remains valid.

Entry: demonstrated operational readiness, sustainable cost, and rehearsed recovery.
Permitted action: approved autonomy inside scope, never beyond it.

There is no “autonomous at everything” level. Maturity belongs to a specific decision and population.

The dimension matrix

Assess each dimension separately:

Dimension Entry question Evidence
data do we have rights, provenance, freshness, and quality? contract, profile, and failures
outcome can we observe closure and compare? definition, coverage, and baseline
control does action have a boundary, approval, and reversal? permission and rollback test
operations can we detect, escalate, degrade, and recover? exercise and runbook
ownership who accepts outcome, risk, and cost? charter and budget
economics does observed value exceed full cost within range? measurement and finance validation

The published level is the minimum across critical dimensions. If data is level one while operations is level three, the card remains level one. Show the full profile to direct investment, but do not hide the bottleneck behind an average.

Artifact: the entry card

A catalog card should answer:

  1. What is the current level, and since when?
  2. Which evidence supports it, and with what coverage?
  3. Which dimension is limiting?
  4. Which action is allowed now?
  5. Which action remains prohibited?
  6. Which test enables the next level?
  7. Who signs that entry?
  8. Which event reduces or expires the level?

A material change to policy, source, population, permission, or outcome reduces the level automatically. Revalidation is not punishment. It recognizes that readiness is contextual.

Illustrative example: accounts payable

A card named “accounts-payable agent” is divided. Field extraction may be validated assistance; coding proposal may be bounded action for purchase-order invoices; bank-account change and payment release remain outside autonomy. They share documents but not consequences.

The catalog shows the profile instead of one badge. If exception ownership is missing, every action producing exceptions remains assistance. This example is illustrative and does not certify a product or prescribe a universal level.

Readiness review using local evidence

AWS describes Operational Readiness Reviews as tailored checklists informed by incident learning to inspect risk across the lifecycle. The catalog can use the same mechanism: every incident adds a question or test rather than an abstract rule copied from another sector.

In Quantum Automation Center, catalog, states, timelines, artifacts, logs, analytics, agents, permissions, and approvals can connect a level to its executions. A label should not be edited manually without references to supporting evidence.

Avoid maturity theater

Do not reward document count. A tested rollback is worth more than a runbook nobody executed. Do not allow unreviewed self-assessment for material actions. Do not freeze levels for a year; use dates and expiry events. Do not make levels a team ranking; they measure case readiness.

Maturity is not value either. An agent can be highly prepared and solve a small problem. Portfolio decisions still compare value and risk; the catalog says what can be operated responsibly.

The strongest counterargument

Levels can become static bureaucracy. Teams fill templates to obtain a badge, and context changes the next day. The barrier may also discourage low-risk experiments that need speed.

That criticism is valid. Keep criteria few and observable, automate evidence, and use level zero for experiments with expiry. Rigor rises before action, not before thinking. Behavior demonstrates a level; document volume does not.

When not to use this approach

Do not apply production levels to isolated prototypes with no real action. Label them exploration with a data boundary and expiry. Do not use the scale as a regulatory guarantee or external certification.

Use it when the catalog influences purchasing, deployment, or permissions. Replacing “ready” with a level and pending entry test turns a commercial promise into an operating decision that can be verified, reduced, and revoked.

Sources