CIOPages
All RFP question modules

Security & Compliance

Business continuity & disaster recovery questions to ask a software vendor

Questions on what happens when the vendor's region, service or company is in trouble: recovery point and time objectives, disaster recovery architecture, continuity testing and results, and how the vendor communicates during an incident.

101
questions
27
RFI
38
RFP
36
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Provide your contractual Recovery Point Objective (RPO) and Recovery Time Objective (RTO) commitments for the production service. The response should be a table with rows for each service tier (e.g., Standard, Enterprise) and columns for: RPO, RTO, contractual backing (e.g., SLA Section 3.1), and the service credit remedy for a breach.

Why it matters. RPO and RTO are foundational resilience commitments customers depend on for their own BCP planning. Commitments must be contractually binding and have clear remedies to be meaningful. This question separates firm, tiered commitments from marketing aspirations.

Good answer
  • Response provided in the requested table format
  • Specific RPO/RTO values are stated in hours or minutes for each tier
  • Explicit references to contract or SLA sections are provided
Red flags
  • Refuses to use the table format
  • Provides targets or 'goals' instead of contractual commitments
  • RPO/RTO values are absent from the contract or SLA

2. Describe the disaster-recovery architecture of the production service, including whether it is active-active, active-passive (hot standby), warm standby, or cold standby across regions.

Why it matters. The DR posture directly determines the achievable RTO. Active-active architectures can support near-zero recovery times, while cold standby may take hours or days. The buyer needs to understand the architectural reality behind the vendor's recovery commitments.

Good answer
  • Names the specific DR pattern in use (e.g., active-active)
  • Describes data replication mode (synchronous, asynchronous, lag bounds)
  • Maps the chosen architecture to the stated RTO
Red flags
  • Cannot describe the pattern in specific, industry-standard terms
  • Claims active-active but describes a manual failover process
  • The described architecture is insufficient to meet the claimed RTO

3. State the cadence of disaster-recovery exercises (e.g., tabletop, partial failover, full failover) and confirm the date of the most recent exercise of each type.

Why it matters. A documented testing cadence with recent dates demonstrates that the vendor's resilience is operational rather than theoretical.

Good answer
  • Cadence stated for each exercise type (e.g., quarterly tabletops, annual full failover)
  • Dates of most recent tests for each type are provided and fall within the vendor's stated cadence
  • Exercise types include realistic failure scenarios
Red flags
  • No documented testing cadence
  • Last full failover exercise is older than the vendor's own stated cadence
  • Only tabletop exercises are performed; no actual failover testing

4. Provide the URL of your public status page and confirm whether it reports component-level status (e.g., per-region, per-API).

Why it matters. A public status page with component granularity is the baseline for incident communication. Status pages limited to a single rolled-up indicator can hide a partial or regional outage.

Good answer
  • A public status page URL is provided
  • The status page reports status per-region and per-major-component (e.g., API, UI)
  • A historical incident archive is publicly accessible
Red flags
  • No public status page exists
  • The status page only shows a single aggregate 'up/down' status
  • The status page is hosted on the same infrastructure as the product itself

5. Confirm whether the published RPO applies to all customer-generated content and configuration (including fine-tuned models, vector embeddings, prompt history, and agent state), and identify any data classes that are excluded from this recovery commitment.

Why it matters. AI services often create high-value derived artifacts (fine-tuned weights, vector indexes) whose loss can be catastrophic. Vendors sometimes scope RPO commitments to a narrow subset of data, leaving these critical artifacts unprotected. This question forces disclosure of any such gaps.

Good answer
  • Explicit enumeration of AI-specific data classes covered by the RPO
  • Clear disclosure of any artifacts that are best-effort only or have a weaker RPO
  • Confirms that fine-tuned model weights and vector indexes are within scope
Red flags
  • Refers only to 'customer data' without defining the scope
  • Fine-tunes, vector indexes, or agent state are explicitly excluded
  • Answer is evasive or non-committal about derived artifacts

6. Identify the geographic regions and availability zones across which the production service is deployed for redundancy.

Why it matters. Geographic redundancy bounds the blast radius of a regional outage, fiber cut, or localized disaster. Multiple availability zones in one region do not protect against a region-wide outage.

Good answer
  • Names specific cloud provider regions (e.g., 'us-east-1' and 'us-west-2')
  • States the geographic separation between primary and DR sites
  • Discloses cross-region replication topology
Red flags
  • Single-region deployment only
  • Vague reference to 'multiple data centers' without naming regions
  • DR region is in the same metropolitan area as the primary region

7. Confirm whether the most recent full DR failover exercise met the contractually stated RTO and RPO, and disclose any gaps observed between the tested result and the commitment.

Why it matters. An exercise shows whether the vendor meets its own RTO and RPO under test conditions, and the remediation plan shows what it does about any gap.

Good answer
  • Direct answer confirming that RPO/RTO were met
  • Provides the measured RPO/RTO values from the exercise
  • Transparently discloses any gaps and describes the remediation plan and status
Red flags
  • Refuses to state whether the exercise met the stated RPO/RTO
  • Cannot provide the measured values from the test
  • Admits to gaps but has no documented remediation plan

8. State your Service Level Objective (SLO) for the maximum time between the detection of a major incident and the first public status-page update.

Why it matters. Slow status updates force customers to assume the worst and burn their own engineering time on diagnosis. A documented SLO for time-to-communicate holds the vendor accountable for timely disclosure.

Good answer
  • A specific time target is provided (e.g., within 15 minutes of declaration for a Sev1)
  • The SLO is stated in a public policy or the customer SLA
  • Historical performance against the SLO is tracked and can be shared
Red flags
  • No SLO exists for the first update
  • The target is 'as soon as possible' or similarly non-specific
  • Recent incidents have missed the stated SLO without explanation

The full set: 101 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 165 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • RPO / RTO commitments & SLA backing (14)
  • DR architecture & geographic redundancy (27)
  • BCP testing cadence & results (19)
  • Incident communication & status-page discipline (24)
  • Incident response process (1)
  • Corporate resilience (10)
  • Vendor viability (6)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 165 changes. Three examples:

Wrong or outdated citation

Draft: Does your SOC 2 Type II report describe the DR testing control and include the auditor's opinion on its operating effectiveness?

Now: Does your SOC 2 Type II report describe the DR testing control, and what tests did the auditor perform on it and with what results?

A SOC 2 Type II opinion covers the system description and the controls as a whole; per-control content is the auditor's description of tests and results (AICPA SOC 2 guide, section 4 of the report). There is no per-control opinion.

Wrong or outdated citation

Draft: Service credits are the buyer's only direct contractual lever when an outage occurs.

Now: Service credits are a direct contractual remedy when an outage occurs.

Not the only lever: contracts can also carry termination rights and damages claims, which deep-dive.005's follow-up asks about.

Wrong or outdated citation

Draft: as part of annual SOC 2 or ISO 27001 attestations

Now: as part of SOC 2 reports or ISO 27001 certification audits

ISO/IEC 27001 is a certification by an accredited certification body, not an attestation; SOC 2 is the attestation report (AICPA AT-C 205).

Questions about this module

How many business continuity & disaster recovery questions are there?

101: 27 for the RFI stage, 38 for the RFP and 36 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 165 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related