8 questions from the RFI stage, free
These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.
1. Provide your contractual Recovery Point Objective (RPO) and Recovery Time Objective (RTO) commitments for the production service. The response should be a table with rows for each service tier (e.g., Standard, Enterprise) and columns for: RPO, RTO, contractual backing (e.g., SLA Section 3.1), and the service credit remedy for a breach.
Why it matters. RPO and RTO are foundational resilience commitments customers depend on for their own BCP planning. Commitments must be contractually binding and have clear remedies to be meaningful. This question separates firm, tiered commitments from marketing aspirations.
- Response provided in the requested table format
- Specific RPO/RTO values are stated in hours or minutes for each tier
- Explicit references to contract or SLA sections are provided
- Refuses to use the table format
- Provides targets or 'goals' instead of contractual commitments
- RPO/RTO values are absent from the contract or SLA
2. Describe the disaster-recovery architecture of the production service, including whether it is active-active, active-passive (hot standby), warm standby, or cold standby across regions.
Why it matters. The DR posture directly determines the achievable RTO. Active-active architectures can support near-zero recovery times, while cold standby may take hours or days. The buyer needs to understand the architectural reality behind the vendor's recovery commitments.
- Names the specific DR pattern in use (e.g., active-active)
- Describes data replication mode (synchronous, asynchronous, lag bounds)
- Maps the chosen architecture to the stated RTO
- Cannot describe the pattern in specific, industry-standard terms
- Claims active-active but describes a manual failover process
- The described architecture is insufficient to meet the claimed RTO
3. State the cadence of disaster-recovery exercises (e.g., tabletop, partial failover, full failover) and confirm the date of the most recent exercise of each type.
Why it matters. A documented testing cadence with recent dates demonstrates that the vendor's resilience is operational rather than theoretical.
- Cadence stated for each exercise type (e.g., quarterly tabletops, annual full failover)
- Dates of most recent tests for each type are provided and fall within the vendor's stated cadence
- Exercise types include realistic failure scenarios
- No documented testing cadence
- Last full failover exercise is older than the vendor's own stated cadence
- Only tabletop exercises are performed; no actual failover testing
4. Provide the URL of your public status page and confirm whether it reports component-level status (e.g., per-region, per-API).
Why it matters. A public status page with component granularity is the baseline for incident communication. Status pages limited to a single rolled-up indicator can hide a partial or regional outage.
- A public status page URL is provided
- The status page reports status per-region and per-major-component (e.g., API, UI)
- A historical incident archive is publicly accessible
- No public status page exists
- The status page only shows a single aggregate 'up/down' status
- The status page is hosted on the same infrastructure as the product itself
5. Confirm whether the published RPO applies to all customer-generated content and configuration (including fine-tuned models, vector embeddings, prompt history, and agent state), and identify any data classes that are excluded from this recovery commitment.
Why it matters. AI services often create high-value derived artifacts (fine-tuned weights, vector indexes) whose loss can be catastrophic. Vendors sometimes scope RPO commitments to a narrow subset of data, leaving these critical artifacts unprotected. This question forces disclosure of any such gaps.
- Explicit enumeration of AI-specific data classes covered by the RPO
- Clear disclosure of any artifacts that are best-effort only or have a weaker RPO
- Confirms that fine-tuned model weights and vector indexes are within scope
- Refers only to 'customer data' without defining the scope
- Fine-tunes, vector indexes, or agent state are explicitly excluded
- Answer is evasive or non-committal about derived artifacts
6. Identify the geographic regions and availability zones across which the production service is deployed for redundancy.
Why it matters. Geographic redundancy bounds the blast radius of a regional outage, fiber cut, or localized disaster. Multiple availability zones in one region do not protect against a region-wide outage.
- Names specific cloud provider regions (e.g., 'us-east-1' and 'us-west-2')
- States the geographic separation between primary and DR sites
- Discloses cross-region replication topology
- Single-region deployment only
- Vague reference to 'multiple data centers' without naming regions
- DR region is in the same metropolitan area as the primary region
7. Confirm whether the most recent full DR failover exercise met the contractually stated RTO and RPO, and disclose any gaps observed between the tested result and the commitment.
Why it matters. An exercise shows whether the vendor meets its own RTO and RPO under test conditions, and the remediation plan shows what it does about any gap.
- Direct answer confirming that RPO/RTO were met
- Provides the measured RPO/RTO values from the exercise
- Transparently discloses any gaps and describes the remediation plan and status
- Refuses to state whether the exercise met the stated RPO/RTO
- Cannot provide the measured values from the test
- Admits to gaps but has no documented remediation plan
8. State your Service Level Objective (SLO) for the maximum time between the detection of a major incident and the first public status-page update.
Why it matters. Slow status updates force customers to assume the worst and burn their own engineering time on diagnosis. A documented SLO for time-to-communicate holds the vendor accountable for timely disclosure.
- A specific time target is provided (e.g., within 15 minutes of declaration for a Sev1)
- The SLO is stated in a public policy or the customer SLA
- Historical performance against the SLO is tracked and can be shared
- No SLO exists for the first update
- The target is 'as soon as possible' or similarly non-specific
- Recent incidents have missed the stated SLO without explanation
What the audit changed
A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 165 changes. Three examples:
Wrong or outdated citation
Draft: Does your SOC 2 Type II report describe the DR testing control and include the auditor's opinion on its operating effectiveness?
Now: Does your SOC 2 Type II report describe the DR testing control, and what tests did the auditor perform on it and with what results?
A SOC 2 Type II opinion covers the system description and the controls as a whole; per-control content is the auditor's description of tests and results (AICPA SOC 2 guide, section 4 of the report). There is no per-control opinion.
Wrong or outdated citation
Draft: Service credits are the buyer's only direct contractual lever when an outage occurs.
Now: Service credits are a direct contractual remedy when an outage occurs.
Not the only lever: contracts can also carry termination rights and damages claims, which deep-dive.005's follow-up asks about.
Wrong or outdated citation
Draft: as part of annual SOC 2 or ISO 27001 attestations
Now: as part of SOC 2 reports or ISO 27001 certification audits
ISO/IEC 27001 is a certification by an accredited certification body, not an attestation; SOC 2 is the attestation report (AICPA AT-C 205).