The exception path is the real workflow
An invoice needs approval. A reviewer is away. A payment times out. Each needs a different next step. Here is how to design for those cases.
An invoice matches its purchase order, the goods have arrived and the approver signs it off. Payment can proceed. That is the straightforward case.
Now change a few details. The total is £40 over tolerance. The supplier’s bank details have changed. The approver is away. Or the payment request times out, leaving you unsure whether the bank accepted it.
Each needs a different next step: approval, verification, reassignment or reconciliation. Calling them all “failed” leaves someone else to work out what happened and how to continue.
The happy path is only the summary
The happy path is a useful starting point. It shows what happens when the expected information, decisions and services are available. The workflow also needs to explain what happens when one is missing.
Take a box labelled “Validate invoice”. It may check required fields, compare the amount with a purchase order and apply a tolerance. An unreadable invoice and an amount over tolerance need different responses, even though both stop at that box.
A validation step is a family of paths
The ordinary route is only one member of the set. Each failed assumption changes what the operation needs.
- Complete and consistent Continue to the ordinary next step
- Required evidence missing Request, wait and resume
- Sources disagree Contain side effects and reconcile
- Outside this team’s scope Transfer with context and ownership
- Unsafe or irreversible Stop and escalate to authority
Process notations already recognise this. BPMN has explicit events for time, messages, errors, escalation, cancellation and compensation. CMMN exists for work whose plan evolves as facts and judgement emerge. The symbols matter, but they are not the operating contract. A timer on a diagram does not name the person who receives the case, show them the evidence or give them authority to act.
Not every exception is an error
An exception may be valid work that needs another route. A high-value request needs approval. Incomplete evidence needs a follow-up. Conflicting information may need someone to investigate.
The distinction matters because the response to one category can make another worse. Retrying missing evidence does not create the evidence. Sending a network timeout directly to an approver asks a person to diagnose software behaviour. Automatically accepting conflicting bank details turns uncertainty into an irreversible action.
Different conditions require different first moves
Start with recurring variations and failures that would have serious consequences. Give new or unclear cases a place to stop safely, with an owner who can investigate.
Find the work the diagram left out
Look at the last ten cases that did not finish normally. Follow each through the messages, corrections and decisions that finally resolved it, or left it waiting.
The evidence is usually distributed:
- queue items that have been waiting far longer than their neighbours;
- fields people repeatedly correct after import;
- support tickets raised to unblock an ordinary transaction;
- spreadsheets that classify cases more precisely than the main system;
- email and chat threads that carry approvals, context or warnings;
- reversed transactions, reopened cases and manual database changes;
- the colleague everybody calls when the official route stops making sense.
The person missing from the diagram may be doing this repair. The spreadsheet beside the system may record what the main application leaves out. Find out what both contribute before changing the workflow.
For every difficult case, identify the last trustworthy state, the assumption that stopped holding, the actions taken, the information used, the people involved and the point at which ordinary work could continue. That produces an exception path grounded in evidence rather than imagination.
Escalation is only the middle of the path
A dependable exception path reaches a safe outcome and leaves enough evidence to improve the system.
- 01DetectWhich assumption failed, and how do we know?
- 02ContainWhat must not happen while the state is uncertain?
- 03UnderstandWhat is known, missing or contradictory?
- 04AssignWho owns the next action and has authority?
- 05DecideWhat was chosen, for what reason?
- 06Recover and learnWhere does work resume, and should the design change?
Give the exception an owner and a clock
Once a case stops, show who owns the next action, what they need and when to follow up.
A queue entry should explain the problem. “Amount over tolerance: approval needed” tells a reviewer more than “Failed”. Include the evidence, any previous attempts and a route to reassign the case if the usual owner is away.
An exception should be managed state, not inbox folklore
- Detected
- Owned
- Investigating or waiting
- Decided
- Recovered or stopped
- Trigger
- Which assumption failed, and what classified the condition?
- Scope
- Which case, customer, transaction or output is affected?
- Owner
- Who is responsible for moving the case forward now?
- Authority
- Who is permitted to choose or approve each available action?
- Evidence
- What is known, missing, conflicting or explicitly uncertain?
- Clock
- When is the next review, deadline, reassignment or escalation?
- Next action
- What response, dependency or decision is the case awaiting?
- Outcome
- What was decided, why, and where does ordinary work resume?
Choose the deadline from the consequence of waiting. A quote may expire. Evidence may become stale. A suspected duplicate payment may need immediate attention even if the normal service target has hours left.
Ownership also needs an absence path. If a case can be assigned to one named person but cannot be reassigned, pooled or escalated, the workflow contains a single point of operational failure.
Choose what to automate, assist or escalate
Decide what to automate by checking the rule, evidence, authority and consequence. Frequent work may still require a person’s judgement. A rare case may be safe to handle automatically if the rule is clear and the action can be reversed.
Different exceptions need different responses
Swipe horizontally to compare every column.
| Response | Use when | The system must | Example |
|---|---|---|---|
| Automate | The distinction is stable, evidence is available, and the action is low-consequence or reversible. | Apply the rule, log the result and expose failure. | Pause a submission that is missing a required document. |
| Assist | Options are bounded, but interpretation still depends on context. | Assemble evidence, expose conflicts and preserve uncertainty. | Suggest possible duplicate customer records for comparison. |
| Human judgement | The case is novel, ambiguous, consequential or requires authority. | Route complete context to an authorised owner and record the reason. | Approve an exception to policy or verify changed bank details. |
| Redesign | The same exception recurs or consumes growing repair work. | Measure the pattern and change the upstream model or workflow. | Stop repeatedly requesting evidence that could be collected at intake. |
AI can help in the assist column: extracting evidence, proposing classifications, summarising a case or drafting a response. It does not remove the need to define the boundary. NIST’s AI Risk Management Framework makes the same distinction: human and AI roles should be differentiated, with clear responsibility for oversight and decisions. The practical version is described in Giving AI a job description.
If an automated capability cannot explain what it saw, preserve uncertainty or hand the case to an authorised person, it has not eliminated the exception. It has made the exception harder to see.
Separate business exceptions from technical failures
Technical resilience and operational exception handling meet at the same case, but they are not interchangeable. A business rule exception is often valid work that needs a different route. A transient infrastructure fault may disappear if the same intent is attempted later. An unknown outcome is more dangerous: the action may have happened even though the caller never received confirmation.
The same red status can conceal three different realities
The invoice exceeds tolerance or two verified sources disagree.
Route, request evidence, obtain authority or stop. Repeating the same calculation changes nothing.A dependency returns an explicit busy or temporarily unavailable response.
Use a bounded retry with delay, backoff and jitter, provided the operation is safe to repeat.The payment request times out after submission and no result is received.
Reconcile the outcome or repeat with an idempotency key. Never infer “not done” from silence.Retry is not a universal recovery policy
A retry is appropriate when the failure is plausibly temporary, the intent is still valid and repeating it cannot create an extra business effect. It also needs a budget. Repeated attempts at several layers can amplify one failing operation into a surge of work that prevents the dependency from recovering.
Idempotency protects the effect, not the appearance of the response. A caller supplies a stable request identifier for one business intent. If the same request arrives again, the receiving system can return the recorded outcome or continue the existing work instead of charging, booking or creating a second time. Guessing duplicates from similar payloads is weaker because two identical-looking requests may represent two genuine intentions.
Retries should therefore have a clear owner, an attempt limit, increasing delay, randomness to avoid synchronized demand and a deadline aligned with the operation. Validation errors and explicit permanent failures should not be retried. When the result is unknown, reconcile first or use a protocol designed to make repetition safe.
Compensation is another business action
Long-running workflows often cross systems that cannot share one database transaction. Some steps commit before a later step fails. Recovery may mean resuming from recorded progress, completing the remaining work, substituting an alternative or performing a compensating action.
A refund is a new transaction. It does not erase the original payment, the notification already sent or the time the customer spent waiting. Give compensating actions their own permissions, records and recovery paths.
Place irreversible actions after the strongest available validation. Record checkpoints around committed effects. If compensation itself fails, preserve its progress so the recovery can resume instead of starting the entire workflow again.
Design the human step
Define the review task as carefully as the automated steps.
The reviewer needs to understand why the case stopped, what the system knows, what remains uncertain and which actions are permitted. They need enough authority to decide, or a clear route to somebody who has it. They need a deadline, a handoff path and a safe way to disagree with any automated recommendation.
Lisanne Bainbridge described a central irony of automation in 1983: remove the routine work and people can be left handling only rare, difficult conditions, with less practice and less context. A review step that almost always asks someone to approve the system’s suggestion does not necessarily preserve judgement. It can train a habit of acceptance.
The system can still do substantial work around human judgement. It can gather evidence, contain unsafe side effects, check permissions, track time, preserve a history, prompt for a reason and carry the chosen action back into the workflow. That is useful automation because it improves the decision without pretending the judgement has disappeared.
Capacity matters too. A path that works when two cases need review can collapse when two hundred arrive after an upstream change. Test the information, authority and workload of the exception route under abnormal conditions, not only on a quiet day.
Test the paths that break the assumptions
Test incomplete, duplicated, delayed and conflicting cases as well as the straightforward one.
Start with invariants expressed in business language: never pay the same invoice twice; never release an order before the required check; never discard evidence already accepted; never let a consequential case become unowned. Then design tests that attack the assumptions behind those promises.
Test the cases that invalidate the assumptions
Known handlers need example tests. Cross-system effects need invariant and idempotency tests. Delays, dropped responses and unavailable dependencies need fault injection. Historical incidents and awkward real cases should become a regression pack. Tabletop exercises reveal whether people can understand, own and hand off the work when the usual colleague is unavailable.
Test the capacity of the exception path as well as the correctness of one case. A surge should preserve priority, oldest age and ownership. Ordinary work must not starve difficult work indefinitely, and difficult work must not absorb the entire operation without a visible decision.
Measure exception work without measuring the person
Exception records are useful operational evidence. They show where assumptions fail, where cases wait and which repairs recur. They can also become a surveillance system if every measure is turned into an individual performance score.
Measure the design before measuring the person. Ask whether work becomes unowned, whether the right evidence arrives, whether authority sits in the right place, whether cases are repeatedly rerouted and whether recovery succeeds.
Measures that reveal the shape of the exception path
Do not make a lower exception count the only target. Better detection may make previously hidden work visible. A team may classify cases more honestly. The more useful question is whether avoidable causes are being removed while necessary judgement becomes better supported.
Try it: make an uncertain payment recoverable
A timeout leaves the payment outcome unknown. This small state machine makes the next action explicit and keeps a repeated event from changing the case twice.
How to try it. Download the example and run node exception-transitions.mjs with Node.js. It uses invented records and calls no payment service.
Payment state and event handling
Included file: exception-transitions.mjs
Full code
// TreeNodes · fictional payment-state exercise. No payment API is called.
const routes = {
awaiting_payment: {
timeout: ['awaiting_reconciliation', 'payments_operations'],
provider_paid: ['paid', null],
},
awaiting_reconciliation: {
provider_paid: ['paid', null],
provider_not_found: ['needs_decision', 'payments_operations'],
},
};
export function applyPaymentEvent(state, event) {
if (!event.id || event.paymentId !== state.paymentId) {
throw new Error('The event must identify this payment intent.');
}
if (state.processedEvents.includes(event.id)) return state;
const route = routes[state.status]?.[event.type];
if (!route) throw new Error('This transition is not permitted.');
return {
...state,
status: route[0],
owner: route[1],
processedEvents: [...state.processedEvents, event.id],
};
}
// Run with Node.js: node exception-transitions.mjs
const initial = {
paymentId: 'PAY-1042', status: 'awaiting_payment',
owner: 'payments_operations', processedEvents: [],
};
const timeout = { id: 'EV-1', paymentId: 'PAY-1042', type: 'timeout' };
const uncertain = applyPaymentEvent(initial, timeout);
const replay = applyPaymentEvent(uncertain, timeout);
const resolved = applyPaymentEvent(replay, {
id: 'EV-2', paymentId: 'PAY-1042', type: 'provider_paid',
});
console.log(uncertain.status, replay.processedEvents.length, resolved.status);
What your result should show
- The output is: awaiting_reconciliation 1 paid. The timeout assigns the case to payments operations; replaying that event leaves one processed event.
- A provider-paid event resolves the same payment intent. A provider-not-found event requires a decision; it does not silently initiate another payment.
- An event for a different payment intent or an unsupported transition is rejected.
Change one condition
Replace provider_paid with provider_not_found and inspect the resulting state and owner. In a deployed system, authenticate provider events and persist the state change and event identity in one transaction. This example isolates transition logic; its in-memory event list is not a durable idempotency store.
Let the exception path change the main path
A recurring exception is a design signal. It may reveal that intake asks for the wrong evidence, two policies conflict, a tolerance is unrealistic, ownership is misplaced or the data model has compressed away a distinction the operation still needs.
Review exception patterns regularly and choose deliberately:
- remove an upstream cause rather than improving downstream repair;
- collect required evidence earlier and once;
- formalise a stable variation as an ordinary supported path;
- automate containment or evidence gathering without automating the judgement;
- repair a policy, system boundary or source of conflicting truth;
- retain a deliberate human capability for rare, high-impact cases;
- retire a rule or handler whose reason no longer exists.
Begin with the ordinary route and the important exceptions you already know. Build a small complete path, watch it in use and add the missing cases as you learn.
A useful exception path lets someone understand the case, take responsibility and recover without losing progress or repeating an action by accident. Review the cases that keep returning; they may point to a problem earlier in the process.
Takeaways
- Follow cases that did not finish normally and record what helped them move again.
- Give each recurring exception a clear next action and enough context for someone to take over.
- Use repeated exceptions to find problems earlier in the process.
REFERENCES AND FURTHER READING 12 sources
- Workflow Exception Patterns Nick Russell, Wil M. P. van der Aalst and Arthur H. M. ter Hofstede, CAiSE 2006 A technology-independent classification of workflow exceptions and the handling actions that can apply to a work item, case or process.
- Business Process Model and Notation (BPMN), Version 2.0.2 Object Management Group, 2014 The formal notation for modelling normal flow alongside timers, errors, escalation, cancellation, compensation and other events.
- Case Management Model and Notation (CMMN), Version 1.1 Object Management Group, 2016 A companion model for work whose plan evolves with evidence, events and professional judgement rather than following one fully prescribed sequence.
- Ironies of Automation Lisanne Bainbridge, Automatica 19(6), 1983 The classic warning that automation can leave people responsible for rare abnormal conditions while giving them less practice and context with which to respond.
- HTTP Semantics: Idempotent Methods IETF, RFC 9110, 2022 Defines idempotency for HTTP methods and explains why a client needs stronger knowledge before automatically retrying a non-idempotent request.
- Making retries safe with idempotent APIs Malcolm Featonby, Amazon Builders’ Library A practical account of ambiguous outcomes, caller-provided request identifiers and the design needed to repeat an intent without repeating its effect.
- Addressing Cascading Failures Mike Ulrich, Site Reliability Engineering, Google Guidance on bounded retries, exponential backoff, jitter, retry budgets and avoiding repeated retry policies across several layers.
- Sagas Hector Garcia-Molina and Kenneth Salem, ACM SIGMOD, 1987 The foundational paper on long-lived transactions made from smaller steps with compensating actions when part of the sequence cannot continue.
- Compensating Transaction pattern Microsoft Azure Architecture Center Explains why compensation is domain-specific, may not restore the original state and must itself be restartable and observable.
- Testing for Reliability Alex Perry and Max Luebbe, Site Reliability Engineering, Google Covers the system, stress and failure testing needed to establish confidence beyond the ordinary execution path.
- Workload UK Health and Safety Executive Human-factors guidance for assessing workload, staffing and performance under abnormal as well as steady operating conditions.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0) National Institute of Standards and Technology, NIST AI 100-1, 2023 Calls for documented, differentiated roles and responsibilities in human–AI configurations and oversight.