Agent Reliability Engineering for reliable and governable autonomous AI agents
Agent Reliability Engineering (ARE) is the software engineering discipline for making autonomous AI agents reliable, governable, observable, recoverable, and accountable in production.
Intelligence is not enough.
AI agents are becoming capable of making decisions, using tools, writing code, interacting with systems, and taking consequential actions. Capability is increasing rapidly. Reliability must increase with it. Without reliability, you only have liability.
What is Agent Reliability Engineering?
Agent Reliability Engineering (ARE) is a software engineering discipline for making autonomous AI agents reliable, governable, observable, recoverable, and accountable in production.
ARE treats reliability as an architectural property—designed into an agent's identity, authority, contracts, controls, and recovery paths before it receives production authority.
Why do AI agents need reliability engineering?
Traditional software generally executes deterministic or bounded instructions. Autonomous AI agents introduce:
- Probabilistic behavior
- Dynamic planning
- Tool selection
- Unscoped data access
- Delegated execution
- Emergent workflows
- Model drift
- Prompt injection
- Authority expansion
- Multi-agent interactions
The more capable an agent becomes, the more important it becomes to engineer the boundaries within which that capability operates.
The Reliability Surface
The Reliability Surface encompasses the behavioral dimensions and engineering controls that determine whether an autonomous AI agent can be trusted to operate in production.
Each row is a dimension a team can audit independently — a gap in any one of them can undermine the rest, regardless of how well the others are engineered.
Reliability Debt
Reliability Debt is the accumulated risk created when an autonomous system's capabilities, complexity, or authority grow faster than its reliability engineering.
Progressive Autonomy
Agents should not receive unlimited authority simply because they are capable of performing an action. They should earn greater authority through demonstrated reliability.
Observe
Proposed actions are logged, not executed.
Draft
The agent prepares actions for a human to execute.
Act with approval
Execution requires explicit sign-off.
Act within scope
Autonomous execution inside a narrow, verified boundary.
Expanded authority
Scope widens only as evidence accumulates.
Authority should increase when reliability improves and contract when reliability deteriorates.
How ARE fits into the engineering landscape
| Discipline | Primary focus |
|---|---|
| SRE | Reliability of software services and infrastructure |
| AI Safety | Preventing harmful AI outcomes |
| AI Security | Protecting AI systems, data, models, and infrastructure |
| AI Governance | Organizational accountability, policy, risk, and compliance |
| AI Evaluation | Measuring model and agent behavior |
| ARE | Engineering autonomous systems so these concerns operate together reliably in production |
ARE does not replace these disciplines. It connects their relevant engineering practices around the operation of autonomous systems.
Agent Reliability Engineering Concepts
ARE defines a common vocabulary for engineering reliable autonomous AI systems.
AI Agent Reliability
The measurable degree to which an agent behaves as intended, within authority, and recovers safely.
Reliability Surface
The full set of behavioral dimensions and controls that determine agent trustworthiness.
Reliability Debt
Accumulated risk from capability growing faster than reliability engineering.
Progressive Autonomy
Authority that expands with demonstrated reliability and contracts when it deteriorates.
Agent Authority
The explicit, scoped set of actions an agent may take and the data those actions may touch.
Agent Governance
The organizational structures and accountability mechanisms that govern agent behavior.
Agent Reliability Controls
The concrete engineering mechanisms that implement the Reliability Surface.
Reliability Maturity
A five-level model for how systematically reliability engineering has been implemented.
Missing a term?
New concepts are added through community review, not by editorial fiat.
Propose a conceptThe Agent Reliability Engineering Manifesto
Agents are intelligent. They are not yet reliable.
The manifesto lays out the founding principles of the discipline — developed in the open, subject to community review before each version is finalized.
Principles of Agent Reliability Engineering
- 01
Governance is architectural, not operational.
WhyGovernance built as a meeting process, rather than as part of the system itself, will be bypassed under pressure the first time it's inconvenient.
ImplicationAuthority checks and audit logging are implemented in code and infrastructure, not left to a review committee's memory.
- 02
Authority is earned, not assumed.
WhyAn agent's ability to call a tool is not the same as its right to use it unsupervised.
ImplicationNew capabilities launch at the most restrictive stage of Progressive Autonomy and expand only against evidence.
- 03
Reliability must be observable, or it does not exist.
WhyA team cannot manage what it cannot see; unobserved agents accumulate Reliability Debt invisibly.
ImplicationEvery consequential action is logged with enough context to reconstruct why the agent took it. Prefer OpenTelemetry spans over proprietary logs. Enrich each consequential span with identity, authority scope, data scope, and the policy verdict so any OTEL backend can audit the run without a second control plane.
AI Agent Reliability Maturity Model
Five levels for assessing how systematically an organization has implemented the Reliability Surface across its autonomous agents.
Unmanaged
No defined identity or authority boundaries; agent actions are effectively unbounded and untracked.
Observed
Actions are logged, but authority is still broad and largely ungoverned.
Governed
Explicit authority scopes and policy enforcement exist; approval gates are in place.
Accountable
Every action is attributable to an identity and policy; incidents produce traceable root cause.
Adaptive
Authority expands and contracts automatically based on measured reliability.
Agent Reliability Engineering Glossary
Core terms in the emerging Agent Reliability Engineering vocabulary. Canonical concept pages are being published as the working draft develops.
- Agent Authority
- The explicit, scoped set of actions an agent is permitted to take.
- Reliability Surface
- The full set of behavioral dimensions and controls that determine agent trustworthiness.
- Progressive Autonomy
- Authority that expands with demonstrated reliability and contracts when reliability deteriorates.
- Reliability Debt
- Accumulated risk from capability growing faster than reliability engineering.
About Agent Reliability Engineering
ARE is intended to be an open discipline rather than a proprietary methodology.
ARE was proposed by Mike Hogan and is currently a working draft, published openly for review and revision. It has not yet been adopted as a formal standard by any industry body.
Mike Hogan's professional work includes Trustabl. ARE is published as an independent, openly licensed discipline rather than as Trustabl product documentation.
Research and further reading
This manifesto was not written in a vacuum. It builds on academic work that measured the capability/reliability gap, enterprise studies that documented agent failures after pilots, identity and runtime work from vendors who now treat agents as first-class principals, and the frustration of developers who can ship intelligent agents faster than they can govern them.
The list starts with the paper that treats reliability as a measurable engineering problem. We encourage these authors, and anyone else measuring failures, designing controls, or operating agents in production, to add their insights to the manifesto.
Academic research
Towards a Science of AI Agent Reliability
Rabanser, Kapoor, Kirgis, Liu, Utpala, and Narayanan (Princeton). Twelve metrics across consistency, robustness, predictability, and safety; capability gains have not produced matching reliability gains.
arXiv:2602.16666ReliabilityBench
Evaluates consistency, robustness to task perturbations, and fault tolerance under injected tool/API failures. Single-run success rates hide how agents behave under production-like stress.
arXiv:2601.06112τ-bench
Yao, Shinn, Razavi, and Narasimhan. Tool-using agents must follow domain policy across multi-turn user interaction; pass^k shows even strong models are inconsistent across retries.
arXiv:2406.12045On the Reliability of Computer Use Agents
Succeeding once is not the same as succeeding again. Unreliability comes from execution stochasticity, task ambiguity, and behavioral drift across repeated OSWorld runs.
arXiv:2604.17849MCP Tool Descriptions Are Smelly!
Most MCP tool descriptions contain defects that mislead tool choice and arguments. Cleaning them can raise success — and also add steps or regress some tasks — so contracts, not just prompts, matter.
arXiv:2602.14878Engineering Trustworthy Agentic AI for Critical Systems
Treats trustworthiness as an engineering property — safety, robustness, transparency, accountability, security — mapped onto an assurance workflow rather than a benchmark score.
arXiv:2607.18548Evaluation Scores Are Perishable Knowledge Claims
Averaging eval signals inflates trust. Scores have formality, scope, and a validity window; weakest-link ranking of HELM models does not match mean ranking.
arXiv:2607.26191Press and analysts
Boomi / Forrester: Agentic AI Readiness Gap
86% of surveyed enterprises have moved agents beyond pilots; only 34% trust the actions those agents take. "Agentic chaos" correlates with about $2.1M in extra failure cost.
Boomi studyGartner: 40% demote or decommission by 2027
Governance applied as binary — locked down or fully trusted — is predicted to drive rollback after production incidents, not before them.
CoverageVentureBeat: the agent evaluation gap
About half of surveyed enterprises shipped an agent that passed internal evaluations and then failed a customer. Few fully trust automated evaluation as a production gate.
VentureBeatForrester AEGIS and Why AI Agents Fail
AEGIS frames enterprise guardrails for agentic systems. Companion analysis covers compounding errors, goal misalignment, orchestration risk, and why genAI review habits do not transfer.
AEGIS overviewCommercial insights
Microsoft Entra Agent ID
Agents get distinct identities, blueprints, sponsorship, Conditional Access, and lifecycle — not shared service principals or human accounts.
Microsoft LearnZero Trust for AI agents
Microsoft's Zero Trust for AI guidance applies verify-explicitly and least-privilege to agents, memory, tools, and runtime — not only to users and devices.
Security blogCopilot Studio and Foundry governance
Zoned environments, data policies, Agent 365 inventory, and Foundry task-adherence controls treat production agents as governed workloads.
Copilot StudioNVIDIA OpenShell
Out-of-process sandbox, filesystem, network, and process policy so an agent cannot lift its own limits after prompt injection or drift.
NVIDIA Technical BlogOkta for AI Agents
Treats agents as first-class non-human identities: discovery, ownership, scoped tokens, and a kill switch when behavior goes off-policy.
OktaDatabricks Agent Bricks
Identity, tool and data access, traces, and continuous evaluation in one governed execution path instead of after-the-fact review.
DatabricksAWS Forward Deployed Engineering
$1B to embed engineers with customers so agentic systems ship against real data, governance, and operating constraints.
AWSGoogle Cloud $750M partner fund
Partner and FDE investment aimed at production agentic deployments, not model access alone.
Google CloudSmarter models make your agents less safe
Trustabl on why more capable models expand the action surface faster than guardrails, contracts, and least-privilege policy keep up.
TrustablAdd a source
Research, incident data, or a control pattern that belongs here should be proposed in the open. The list should grow with the field.
Open an issueHelp define the discipline
Agent Reliability Engineering is being developed in the open. Contributions from engineers, researchers, security professionals, operators, architects, and AI practitioners are encouraged.
Review the Manifesto
Read the current draft and leave comments directly against specific passages.
Comment on the Google DocContribute code & artifacts
Documentation, diagrams, tooling, schemas, examples, and reference implementations.
GitHubSubmit a proposal
Propose terminology, controls, metrics, and case studies with the working group.
Community