Engineering4 min read

Reliable Business Software: Recovery and Acceptance

Oleksandr Melnychenko··
On this page

When software controls orders, stock or payments, reliability includes correct decisions and recoverable records. Define what must keep working during a failure, how much interruption the business can accept and how the team will restore operations.

Start with the consequence of failure

A dispatch board, stock reservation and billing platform fail in different ways. Identify the work that stops, the decisions affected by stale or incorrect data and the person responsible for recovery.

In discovery, ask which decisions depend on the software, what happens if data is stale or wrong, who can authorise a fallback, and how long the organisation can operate without the system. These answers drive the architecture and the acceptance tests.

Separate availability from correctness. A service may answer every request while assigning an order to the wrong customer. For each important workflow, define both what counts as a successful operation and what must remain true during a failure. For a stock reservation, one such rule might be that confirmed allocations never exceed the quantity available for allocation. The rule needs a test for concurrent requests, not only a successful single-user demonstration.

Make unsafe uncertainty visible

If a dependency fails, do not present missing information as a confirmed answer. Show the last known update time, source and status. Route decisions that cannot be validated to a human reviewer or a documented fallback. The right behaviour is a domain decision agreed with the operator, not a universal software setting.

For integrations, define retries, duplicate handling, reconciliation and what happens when the external system returns conflicting records. For migration, reconcile identifiers, required relationships and business totals before cutover. Equal row counts can conceal a missing record and an unrelated duplicate; they are one check, not proof that the transfer is correct.

Record decisions you may need to explain

Decide which events must be traceable: access, changes, approvals, automated rules and exports. Capture the actor, time, relevant inputs and outcome while avoiding unnecessary sensitive data in logs. Protect the records and set retention with the client's legal and operational teams.

The audit trail should let an authorised reviewer reconstruct a business decision: which record changed, under whose authority and using which information.

Design recovery around the business

Define two recovery objectives for each workflow: RTO, the maximum acceptable time from interruption to restoration, and RPO, the maximum acceptable data-loss window measured back from the interruption. These are targets to design and test against, not results established by owning a backup. AWS's guidance on recovery objectives relates them to business impact.

Test backup restoration, dependency outages and manual fallback. Measure recovery through to usable, reconciled business operations, not just a running server. Redundancy, regional failover and 24-hour support may be appropriate, but each needs an agreed operating model, responsible team and cost.

A useful release test asks: if a key dependency is unavailable for an hour, what can users still do, what must stop, and how will records be reconciled afterwards?

For example, a warehouse outage might leave new orders visible as pending while preventing an unverified stock reservation. The recovery test should identify the affected orders, replay eligible operations and reconcile the final allocations. Record the tested outage, data volume, configuration and observed recovery time. Passing this scenario supports a claim about those conditions; it does not establish that every possible outage is covered.

Measure the system users experience

Infrastructure metrics matter, but the operator needs domain signals too: unprocessed orders, stale shipments, unmatched payments, failed data syncs, or reports awaiting validation. Alert on a condition that needs a response, link it to a runbook, and assign an owner.

What to agree before development

Turn these requirements into acceptance criteria: the failures to simulate, what users should see, the recovery targets and the person responsible for each check. Include the authoritative data sources and support responsibilities in the scope.

Describe the workflow with the greatest impact if it fails and the systems holding its records. We can define recovery and support requirements and estimate the work needed to meet them.