Pick your recovery pattern from two numbers and the decision takes minutes. How much data can you afford to lose, measured in time, and how long can you be down. AWS calls those the recovery point objective and the recovery time objective, and its Well-Architected guidance maps four patterns onto them: backup and restore, pilot light, warm standby, and multi-site active-active. Everything else, including the cost, follows from where your business sits on those two axes.

Start with two numbers, not four options
The recovery point objective is how much work you are willing to redo. If your RPO is one hour, you accept losing up to an hour of orders, bookings or edits. The recovery time objective is how long the business can run without the system. Those two numbers are commercial decisions, not technical ones, and they should come from whoever owns the revenue rather than whoever owns the servers.
Ask the question in money. What does an hour of downtime cost in lost sales, idle staff and support calls? IBM’s 2026 Cost of a Data Breach Report frames its global average of USD 4.99 million as roughly USD 1,100 per hour across the whole lifecycle of an incident, which gives you a reference point even though your own figure will differ. Once you have an hourly cost, the argument about whether to spend an extra few hundred a month on warm standby answers itself.
- RPO: the maximum acceptable data loss, in minutes or hours.
- RTO: the maximum acceptable outage, in minutes or hours.
- Both per workload. Your invoicing system and your marketing site rarely deserve the same answer.
The four patterns, as AWS defines them
AWS documents the four patterns in its Well-Architected Framework, with explicit recovery ranges. The table below follows that guidance. Azure and Google Cloud use different names for the same shapes, so the pattern language travels even if the service names do not.
| Pattern | RPO | RTO | Running in the recovery region | Relative cost |
|---|---|---|---|---|
| Backup and restore | Hours | 24 hours or less | Nothing but backups; continuous backup can cut RPO to about 5 minutes | Lowest |
| Pilot light | Minutes | Tens of minutes | Core infrastructure and replicated data, always on; application tier switched off | Low |
| Warm standby | Seconds | Minutes | A scaled-down but fully working copy, always serving | Moderate |
| Multi-site active-active | Near zero | Potentially zero | Full capacity in two or more regions, all taking traffic | Highest |
AWS adds a note that catches people out, and it is worth repeating. Pilot light and warm standby look similar on a diagram because both keep a copy of your assets in the second region. The difference is that a pilot light cannot process requests until someone takes action. A warm standby is already serving, just smaller. On a bad morning that distinction is the gap between scaling up and starting up.
A pilot light is a plan you execute. A warm standby is a plan that is already running.
Scale a warm standby all the way to production capacity and the industry calls it hot standby. At that point you are paying close to double for infrastructure, and the only reason to stop short of full active-active is that your application cannot handle writes in two places at once. That limitation is usually in the database, and it is where most multi-region projects spend their effort.
What each pattern costs to keep warm
Replication tooling is the part people assume will be expensive, and it is not. These are published list prices in October 2026. AWS Elastic Disaster Recovery charges a flat USD 0.028 per source server per hour with no minimum, and AWS’s own worked example multiplies that by 730 hours a month. Azure Site Recovery is priced per protected instance per month, and Microsoft’s Retail Prices API returns USD 25.00 for a virtual machine replicated to Azure and USD 16.00 for one replicated to System Center.
Read the small print on both. The replication licence is not the bill. AWS notes that replication consumes storage for staging disks and compute during drills and recovery, and Azure’s per-instance price excludes the storage and transactions your replicated data uses. In practice, for a handful of servers, the storage line usually exceeds the replication line.
| Availability target | Allowed downtime per month | Allowed downtime per year |
|---|---|---|
| 99.0% | About 7 hours 18 minutes | About 3 days 15 hours |
| 99.9% | About 43 minutes | About 8 hours 46 minutes |
| 99.95% | About 22 minutes | About 4 hours 23 minutes |
| 99.99% | About 4 minutes 23 seconds | About 52 minutes |
That table is the one to take into a conversation with a supplier. A 99.9% promise permits more than eight hours of outage a year, which may be entirely fine for an internal reporting tool and quite unacceptable for a payment flow. Matching your recovery pattern to the availability you have actually been sold prevents a lot of disappointment.

Three worked examples
Patterns make sense once you attach them to real businesses. These three are typical of the clients we work with, with the numbers rounded and the details changed.
- A 12-page brochure site with a contact form. Losing a day of form submissions would be irritating, not fatal. Nightly backups plus a documented rebuild, tested once, is the right answer. Backup and restore, a few dollars a month, and no second environment to maintain.
- A 40-person professional services firm running a line-of-business application. Staff cannot bill without it, so a day of downtime costs real money, but a thirty-minute gap is survivable. Pilot light: data replicated continuously, infrastructure defined as code, application tier started on demand. Replication tooling at roughly USD 20 to 25 per server per month plus storage.
- An online store taking orders around the clock. Every hour down is lost revenue and every lost order is a support conversation. Warm standby, with a scaled-down copy already serving in a second region and health checks ready to shift traffic. Expect to pay for real infrastructure in two places, and expect the database design to be the hard part.
Notice that the first example spends almost nothing deliberately. The AWS whitepaper makes the same point in its own words: for less critical workloads, a valid strategy may be no disaster recovery in place at all. Spending three hundred a month protecting a site you could rebuild in an afternoon is not prudence, it is waste that crowds out the protection your important systems actually need.
The test that decides whether your plan is real
Here is the uncomfortable part. In our experience, most small businesses that believe they have disaster recovery have backups running and have never restored one. The backup job’s green tick tells you a file was written. It tells you nothing about whether the file is complete, whether you still have the credentials to read it, or whether anyone remembers the sequence.
- Pick one real thing. A database table, a document library, a virtual machine. Not a test file you created for the drill.
- Restore it somewhere harmless. A separate environment or a different name, so a mistake cannot overwrite production.
- Time it. Wall-clock, from decision to verified data. Compare that to your stated RTO and adjust whichever number is wrong.
- Write down what surprised you. There is always something: an expired key, a missing DNS record, a licence tied to a hostname.
- Repeat quarterly. Drills on AWS Elastic Disaster Recovery do not change the hourly replication fee, so there is no billing excuse for skipping them.
Ninety minutes a quarter covers this for a small estate. The first drill will be messy and that is the point: you would rather discover the missing DNS record on a Tuesday afternoon than during an outage.
Where these patterns quietly fail
Three failure modes show up again and again, and none of them is about the pattern you chose.
- The dependency nobody replicated. Your application fails over cleanly and then cannot reach the payment gateway allowlist, the licence server or the SMTP relay that only trusts your old IP address.
- The configuration that lives in someone’s head. If the recovery region was built by hand, it has drifted. Infrastructure as code is what makes pilot light credible, because the application tier you switch on is defined rather than remembered.
- The backup that was encrypted by the same incident. Ransomware that reaches your backup target turns four patterns into none. Immutable or logically air-gapped copies exist on all major providers for exactly this reason.
The third one deserves emphasis because it has changed how we advise clients. A recovery plan built only for hardware failure is a recovery plan for the wrong decade. Design it so that a compromised administrator account cannot destroy the copies, and read our guide to choosing a cloud provider for a global business if data residency rules also constrain where those copies can sit.

Getting a plan written without a site visit
Recovery design is one of the easiest things to do remotely, because it is mostly configuration, documentation and testing. Eudora Technology’s cloud solutions service sets the two numbers with you, picks the pattern that matches them, implements the replication and then runs the first drill with your team watching so the runbook is theirs rather than ours. We work across time zones and never need to be in the building.
If your architecture is still in flux, two neighbouring pieces help. Containers and Kubernetes: do you actually need them? covers how much of this gets easier when your workloads are portable, and our serverless piece explains why managed runtimes shrink the surface area you have to recover in the first place.
Frequently asked questions
What is the difference between a backup and disaster recovery?
A backup is a copy of data. Disaster recovery is the ability to run your business again, which needs the data plus the infrastructure, the configuration, the access and a sequence somebody has rehearsed. Backup and restore is the cheapest DR pattern, but only once you have proved you can actually rebuild from it inside your recovery time objective.
How much should a small business budget for disaster recovery?
For a handful of servers, replication tooling at list price is roughly USD 20 to 25 per server per month on AWS or Azure, plus the storage and any compute used during drills. Pilot light typically lands in the low hundreds a month all in. Warm standby costs considerably more because real capacity runs continuously in a second location.
How often should we test the recovery plan?
Quarterly for most small estates, and after any significant change to the application, the data model or the team. A drill does not have to be a full failover. Restoring one real dataset and timing it catches most of the problems, and AWS does not charge extra for test launches on Elastic Disaster Recovery.
Do we need a second cloud provider for disaster recovery?
Usually not. A second region with the same provider removes the realistic failure modes for far less effort, because the tooling, identity model and skills carry across. Running a second provider makes sense when a contract or a regulator requires it, and you should price the duplicated engineering time before you commit.
Does a provider’s uptime SLA mean we can skip DR?
No. An SLA is a commercial remedy, not a guarantee of continuity, and a 99.9% target still permits more than eight hours of outage a year. Service credits do not replace lost orders. Treat the SLA as the floor you design above, not the plan itself.
Want a recovery plan sized to your actual downtime tolerance, implemented and then drilled with your team rather than handed over as a document? Get in touch with Eudora Technology to talk about your project.



