Automation is often evaluated by what it can do when everything goes normally: process a transaction, route a request, reconcile data, approve a low-risk decision, or move a customer through a workflow.
But what happens when the abnormal occurs?
What happens when an input is incomplete, a dependency fails, a downstream system becomes unavailable, or an automated decision encounters a situation it was not designed to interpret? Who is alerted? Who has authority to intervene? What information do they need to act quickly and safely?
Those questions sit at the center of a common operational problem: automated processes lack monitoring, exception handling, or human oversight, increasing operational risk.
Recent developments in autonomous vehicles provide a visible example of the issue. The governance lessons learned with robotaxis apply across industries: automation cannot be considered mature simply because it performs its normal task reliably. It must also be observable, interruptible, accountable, and designed to manage exceptions.
The exception is part of the system
On July 8, 2026, the National Highway Traffic Safety Administration called on automated-vehicle developers to address a pattern of interference with first responders. The agency described instances in which driverless vehicles entered emergency scenes, obstructed ambulances or firefighters, or did not respond appropriately to signals such as flashing lights, flares, smoke, fire, and traffic cones. NHTSA’s point was straightforward: emergency-scene interaction is an expected operational condition, not an isolated edge case. (nhtsa.gov)
A few days later, Zoox issued a software recall following a June incident involving heavy smoke at an active fire scene. According to reporting on the recall, the robotaxi braked hard and attempted to steer away before stopping; a teleoperator ultimately reversed it from the area so that first responders could place traffic cones. The company deployed a software update intended to improve recognition of heavy smoke in certain scenarios. (techcrunch.com)
The immediate issue is clearly important for road safety. However, the broader operational takeaway is equally relevant to organizations deploying automation in finance, healthcare, supply chain, customer service, cybersecurity, HR, and internal operations.
A process that works well under standard conditions may still create material risk if it cannot recognize abnormal conditions, escalate them appropriately, and hand control to a qualified person or fallback process.
In other words, an automation program needs more than functionality – it needs an operating model.
Scale increases the need for controls
When an organization scales an automated workflow, it also scales the consequences of design gaps. A failed manual process may affect one case at a time. A failed automated process can replicate an incorrect action across thousands of transactions, customers, records, or decisions before anyone notices.
The risks are familiar:
- Errors can propagate at speed and volume.
- Service-level commitments can be missed without a visible incident signal.
- Compliance violations can occur without an immediate human review point.
- Customers may receive inconsistent, incorrect, or unexplainable outcomes.
- Teams may discover a problem only after financial, operational, or reputational damage has accumulated.
- Individuals tasked with resolving incidents may lack the information, authority, or runbooks needed to act.
The goal is not to make automation slower by inserting manual review everywhere. It is to be deliberate about where human judgment is essential, where automated recovery is safe, and where escalation must occur.
Observability is a control, not a technical afterthought
Many teams treat monitoring as a post-deployment activity. The workflow is built, released, and then someone later asks for a dashboard.
That sequence creates a gap. If a team cannot observe the health and behavior of an automated process, it cannot responsibly operate it.
Observability should be designed into the process from the beginning. At a minimum, teams should establish standards for:
-
Meaningful metrics
Track more than uptime. Measure throughput, completion rates, error rates, queue age, retry volume, latency, data-quality failures, override rates, and outcome variance. These measures should reveal whether the process is performing as intended, not merely whether it is running. -
Structured logs and traceability
Each automated action should leave an understandable and auditable record: what triggered it, what data it used, what decision or action it took, what downstream systems were involved, and whether an exception occurred. Traceability supports incident response, auditing, and continuous improvement. -
Health checks and synthetic testing
Automated processes should be tested continuously, including checks that simulate critical inputs or dependency failures. A service that appears available may still be unable to complete the business function it was built to support. -
Actionable alerts
Alerts should be tied to clear thresholds, owners, and response expectations. A notification that goes to an unmonitored mailbox is not an operational control. Nor is an alert that provides no context about the impact, scope, or next action.
Monitoring is not only about detecting technical failure. It is about recognizing that business conditions have shifted beyond the automation’s approved operating boundaries.
Exception handling must be designed, not improvised
Often, the system retries indefinitely, writes an ambiguous error, skips a record, or stops quietly. These patterns turn manageable exceptions into hidden operational risk.
A stronger approach defines exception-handling patterns before deployment:
- Use bounded retry and backoff policies rather than endless retries.
- Separate transient failures from business-rule failures and security-relevant anomalies.
- Create structured error categories that support routing and trend analysis.
- Open tickets automatically when ownership or intervention is required.
- Preserve the relevant context so that a responder does not have to reconstruct the event from multiple systems.
- Establish safe fallback states, including pausing, queuing, or requiring approval before a process continues.
- Define decision rights for overrides, restarts, and changes to rules or thresholds.
The Zoox incident is a useful illustration of why this matters. The vehicle did not simply continue through a situation it could not interpret; it stopped, and a teleoperator was able to intervene. That does not eliminate the underlying issue, but it demonstrates the importance of having a defined path from automated uncertainty to human action.
For enterprise processes, the equivalent may be a payment placed on hold, an access request routed to a security reviewer, a suspicious claim sent to a specialist, or an automated customer communication paused before delivery.
Human oversight needs clear triggers and ownership
“Human in the loop” is often used loosely. Effective oversight requires specificity.
Which decisions require human approval? Which conditions trigger a handoff? Who is accountable for responding? How quickly must they respond? What happens if they do not?
Organizations should define human-in-the-loop controls around critical decision points, especially where automation affects safety, regulatory obligations, financial exposure, customer rights, or irreversible outcomes. They should also maintain on-call or roster responsibilities for systems that operate outside normal business hours.
A human escalation path without named ownership is only a diagram. A named owner without authority, access, or a runbook is not prepared to resolve an incident.
Governance turns good intentions into repeatable practice
The underlying challenge is rarely that teams do not understand the value of monitoring or escalation. More often, these controls are inconsistently implemented because they are not embedded in the automation lifecycle.
A practical governance approach includes:
- minimum observability and alerting requirements;
- reusable exception-handling and escalation patterns;
- automation gates before production release;
- documented runbooks and ownership models;
- periodic health checks and control audits;
- evidence that critical workflows have been tested under abnormal conditions;
- clear mapping from policy requirements to implementation standards.
Automation can create speed, consistency, and capacity. But value depends on whether the organization can see what it is doing, recognize when it is outside its limits, and intervene before a small exception becomes a larger incident.






