Failure is an operating condition.
Autonomous systems should assume that tools time out, APIs change, models misclassify, dependencies disappear, and returned evidence can contradict expected state. Treating those events as rare exceptions creates brittle automation. Treating them as ordinary operating conditions forces the architecture to define what happens next before the system is trusted with more responsibility.
Detect before correcting.
Recovery starts with recognizing that the route is unhealthy. A system needs explicit signals for failed tests, contradictory state, missing evidence, stale dependencies, permission errors, and unexpected side effects. Without detection, a failing route can keep producing confident output while moving farther away from a known-good condition.
Contain before retrying.
Blind retries can multiply damage. Once failure is detected, the safest next action is often containment: stop additional side effects, freeze the affected route, preserve evidence, and prevent silent fallback chains from spawning alternate implementations. Recovery should fix the root condition rather than bury it under another layer of compatibility behavior.
Restore, then verify.
A rollback is not complete because a command succeeded. The system should return to a known-safe state and then re-run the checks that define healthy operation. Tests, state inspection, diffs, and other evidence should confirm that the route is stable before normal autonomy resumes. Recovery is therefore part of verification—not merely a way to make an error disappear.