Name the three business processes that cannot stop for a day, write down how long each can actually be down and how much data it can afford to lose, then test a restore before the month is out. That is a continuity plan. The binder, the policy and the certification can come later, and for most companies under a hundred people they never need to come at all.

Start with three processes, not a binder
Sit down with the people who run operations and list what the business does for money. Taking orders. Invoicing. Delivering the service. Paying people. Then ask one question per item: if this stopped at 09:00 on a Tuesday, when would it start hurting? Some answers are in minutes, most are in days, and a few are in weeks. The ones measured in minutes and hours are your continuity scope. Everything else is a nice-to-have with a lower priority and a cheaper plan.
This is the business impact analysis that Ready.gov and ISO 22301 both put at the front of the process, stripped of the paperwork. Doing it properly takes an afternoon and the output fits on one page. Doing it badly takes six weeks and produces a document in which 14 systems are all “tier one”, which means the plan has made no decisions at all.
- What does this process need to run: which system, which supplier, which person?
- How long until the business feels it: minutes, hours, days?
- What is the manual workaround, and how long can we hold it?
- Who decides to invoke the workaround, and who tells customers?
What an outage actually costs
Downtime costs have kept climbing even as infrastructure has become more reliable. Uptime Institute’s Annual Outage Analysis 2026 reports that 57% of respondents to its 2025 survey said their most recent major outage cost more than USD 100,000, and for the second year running one in five put their most recent impactful outage above USD 1 million. Roughly one in ten said their last outage had serious or severe consequences, and this is against a backdrop of outage frequency per site declining for the fifth consecutive year.
The lesson is not that outages are becoming more common, because per site they are not. It is that each one lands on a business that depends on more systems than it did three years ago. Uptime also notes that power remains the leading cause of impactful outages, with failures involving UPS systems, transfer switches and generators prominent, while fibre and connectivity problems are rising and tend to produce longer disruptions. If your continuity plan assumes the internet connection is the reliable part, revisit it.
Security incidents belong in the same conversation. IBM’s 2026 Cost of a Data Breach Report puts the record global average at USD 4.99 million, and the difference between a bad week and a bad quarter is usually whether recovery was rehearsed. Our note on building a simple incident response plan covers the response half of that pairing.
RTO and RPO: the two numbers everything hangs off
Recovery time objective is how long a process may be unavailable. Recovery point objective is how much work you can afford to lose, measured in time. A nightly backup gives you an RPO of up to 24 hours, which is perfectly reasonable for a document store and completely unreasonable for an order queue. Writing both numbers next to each process is what turns “we should improve our backups” into a decision you can price.
| Tier | Example processes | RTO target | RPO target | What that usually needs |
|---|---|---|---|---|
| Tier 1 | Order intake, payment processing, customer-facing application | Under 4 hours | Under 15 minutes | Replicated database, warm standby, documented failover, tested quarterly |
| Tier 2 | Email, CRM, support desk, internal line-of-business apps | Same working day | Up to 1 hour | Managed SaaS with provider SLA, plus your own export of the data |
| Tier 3 | File shares, reporting, intranet, marketing tools | 1 to 3 days | Up to 24 hours | Nightly backup, restore tested twice a year |
| Tier 4 | Archives, historical reporting, dormant projects | A week or more | Up to a week | Cold storage, documented retrieval steps |
Be honest at this stage. A four-hour RTO on a system whose backup takes six hours to restore is not a target, it is a wish. Time the restore, then either accept the real number or spend money to change it. Plans lose credibility the first time a stated objective turns out to be fiction, and after that nobody trusts the rest of the document.

The standards, and the parts worth your afternoon
You do not need to buy a standard to plan well, but knowing what each one contains stops you reinventing the structure. ISO 22301:2019 is the certifiable management-system standard for business continuity, now in its second edition. NIST SP 800-34 Revision 1 is older, from May 2010, and is still the clearest free walkthrough of contingency planning for information systems. Ready.gov publishes free continuity templates and test-exercise handbooks. The multi-agency #StopRansomware Guide from CISA, the FBI, the NSA and MS-ISAC is the one to read for the ransomware-specific recovery sequence.
| Document | Publisher | Reference and date | Best used for |
|---|---|---|---|
| Security and resilience — business continuity management systems | ISO | ISO 22301:2019, edition 2, October 2019 | The audited management-system structure, if a customer or insurer asks for certification |
| Contingency Planning Guide for Information Systems | NIST | SP 800-34 Rev. 1, May 2010 | A free, concrete process for contingency planning, including plan structure and testing |
| Business Continuity Planning | Ready.gov (US DHS) | retrieved October 2026 | Templates, situation manuals and a test facilitator handbook you can run as-is |
| #StopRansomware Guide | CISA, FBI, NSA and MS-ISAC | retrieved October 2026 | Ransomware response and recovery checklists, including offline backup requirements |
If nobody is asking you for ISO 22301 certification, do not start there. Take the structure from NIST and the exercise material from Ready.gov, write six pages, and keep the money.
Backups only count once you have restored one
The 3-2-1 habit still holds: three copies of the data, on two kinds of media, with one copy off-site. The ransomware era adds a fourth condition, which is that at least one copy must be offline or otherwise out of reach of an attacker who holds your admin credentials. The #StopRansomware Guide is blunt about this, and it is the single most common gap we find: backups that run beautifully, into a storage account the same compromised administrator can delete.
- Immutable or offline: object-lock retention, a separate credential boundary, or physical media that leaves the building.
- Separate identity: the backup system should not authenticate through the same single sign-on as everything else.
- Restore test with a stopwatch: pick one system a quarter, restore to a scratch environment, record how long it took and what broke.
- Check what is not covered: SaaS mailboxes, code repositories, CI secrets, DNS zone files and the laptop where somebody keeps the only copy of the pricing model.
- Document the order: which system comes back first, and what depends on it. Identity and DNS almost always lead.
A backup you have never restored is a receipt for a service, not a recovery capability.
The single points of failure that are people
Ask who is the only person who can do each critical task. In most growing companies the answer is uncomfortable: one engineer holds the production credentials, one person knows the payroll process, one director has the bank authorisation. That is a continuity risk with a name and a holiday calendar. Fixing it is cheap: a documented runbook, a second authorised person, and break-glass credentials in a sealed envelope or an emergency vault entry.
Suppliers deserve the same treatment. For each one that your tier-one processes depend on, write down the support channel you would actually use at 03:00, your account identifier, the contractual response time, and what you would do for 48 hours without them. Most plans have a vendor list; very few have the answer to “and if they are the outage?” Remote and hybrid working adds another layer to this, which we cover in cybersecurity essentials for distributed and remote teams.
The two-hour drill that finds the real gaps
Block two hours. Pick one scenario and run it for real rather than discussing it: restore the file server to a scratch environment, fail the website over to the standby, or work through a day of order taking with the main system marked unavailable. Keep a clock running and write down every moment someone had to ask a question nobody could answer.
- Choose the scenario a week ahead and tell people the date, not the detail.
- Appoint one observer whose only job is to take notes with timestamps.
- Do not help from the sidelines. The plan is being tested, not the people.
- Stop at two hours whether or not you finished. Where you got to is the finding.
- Turn the notes into no more than five fixes with owners and dates, then put the next drill in the calendar.
The first drill is always sobering. Credentials have expired, the documented runbook refers to a console that was redesigned, the off-site copy turns out to be 11 days old. That is the point: finding it on a Tuesday afternoon with coffee costs nothing, and finding it during a real outage costs the figures in the Uptime data above.

Keeping the plan alive
Continuity plans rot quietly. Systems get replaced, people leave, a new supplier becomes critical without anybody noticing. Put a 45-minute review in the calendar twice a year with three agenda items: has the tier list changed, did the drills pass, and what new dependency did we take on? Then update the one-page summary and reissue it. If the review keeps getting cancelled, the plan is already out of date and you should assume it is wrong.
Our cybersecurity and resilience service does this work remotely with clients worldwide: the impact analysis over a call, the tier and objective setting with your operations lead, the backup and failover configuration in your own cloud accounts, and the first drill facilitated end to end. The deliverable is six pages and a timed restore, not a binder. Regulatory obligations that run alongside continuity are covered in data protection compliance basics. Clients who want someone physically on site in Sri Lanka should look at eudora.lk.
Frequently asked questions
How is business continuity different from disaster recovery?
Disaster recovery is the technical restoration of systems and data. Business continuity is how the business keeps serving customers while that restoration happens, including manual workarounds, communication and decision rights. You need both, and the second one is the part small companies skip. A perfect technical recovery that takes 36 hours still needs somebody to decide what to tell customers on hour three.
What should our recovery time objective be?
Whatever the business impact analysis says, adjusted for what you can afford. For most growing companies a four-hour RTO on the revenue-critical system and a same-day target on everything else is both achievable and sufficient. Pushing below an hour means warm standby infrastructure and regular failover testing, which roughly doubles the hosting bill for that system. That is worth it for a payments platform and wasteful for an intranet.
Do we need ISO 22301 certification?
Only if a customer, regulator or insurer asks for it. Certification proves you run a management system; it does not make you more recoverable than a six-page plan that has been drilled. If a tender requires it, budget for an audit cycle and external support. Otherwise borrow the structure, which is free to read about, and spend the money on tested backups.
How often should we test?
One timed restore per quarter, rotating through systems, and one two-hour scenario drill twice a year. That cadence is light enough to survive a busy quarter and frequent enough to catch drift. Record the results somewhere durable, because the trend over four quarters tells you more than any single test.
Does cloud hosting mean we do not need a continuity plan?
No. It changes which failures you own. The provider handles hardware and much of the availability engineering, but you still own your configuration, your identity provider, your data exports, your dependency on a single region and your own mistakes. Uptime Institute’s data on connectivity-related outages is a reminder that the link between you and the cloud is itself a single point of failure.
Want a continuity plan that fits on six pages and a restore you have actually timed? Get in touch with Eudora Technology to talk about your project.



