CIOPages
All RFP question modules

Delivery & Operations

Support & operations questions to ask a software vendor

Questions on production support: severity definitions, response and restoration targets, service credits, coverage hours, support channels and escalation. Initial deployment is in Implementation & onboarding; major outages are in Business continuity & DR.

92
questions
30
RFI
31
RFP
31
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Provide your published severity level definitions (e.g., P1/P2/P3/P4 or equivalent), including the customer-impact criteria that determine each level.

Why it matters. Severity definitions drive every downstream SLA commitment. Vendors who define severity unilaterally and narrowly can deflect legitimate production incidents into lower tiers, eroding the practical value of the SLA. Buyers need to see the criteria in writing to compare across vendors.

Good answer
  • Provides written, versioned severity definitions tied to customer-impact criteria.
  • Allows customer input on initial severity classification.
  • Includes examples of incidents at each severity level.
Red flags
  • Severity assigned solely at vendor discretion with no appeal path.
  • Definitions exist only in sales decks, not in the support policy or MSA.
  • Only P1 has a meaningful response commitment.

2. Specify the coverage hours for each severity level in your base support offering, including whether weekends and public holidays are covered.

Why it matters. Off-hours and weekend coverage are where a 'global 24x7' marketing claim hides regional limitations or business-hours-only coverage for anything below P1. Buyers need explicit per-severity coverage hours.

Good answer
  • 24x7x365 coverage for P1 incidents in the base tier.
  • A clear policy for P2-P4 incidents submitted outside of business hours.
  • Holiday coverage is explicitly addressed, and the holiday calendars that apply are named.
Red flags
  • Weekend coverage is offered only at a premium tier.
  • Vague reference to 'follow-the-sun' without concrete regional staffing.
  • Public holiday coverage is based on a single region's calendar, leaving gaps for global customers.

3. List the support channels available to customers (e.g., portal, email, chat, phone, Slack/Teams Connect) and indicate which severity levels each channel is appropriate for.

Why it matters. Channel availability shapes how quickly a customer can reach a human during an incident. Some vendors restrict phone or chat to top severities or premium tiers, which can delay the declaration of a P1 incident.

Good answer
  • Multiple channels are available, including a real-time channel for P1 incidents.
  • Phone or live chat is available for P1/P2 incidents in the base tier.
  • The channel-to-severity mapping is clearly documented.
Red flags
  • Only email or a ticket portal is available, with no real-time channel for emergencies.
  • Phone support is reserved for the highest premium tier.
  • Live chat is available only during the business hours of a single region.

4. Describe your formal escalation path for production incidents, including the trigger conditions and named roles at each escalation level up to an executive sponsor.

Why it matters. An informal escalation process — 'just call your CSM' — collapses under pressure. Buyers need to see a written escalation matrix with concrete triggers (e.g., time-based) and specific, accountable roles up to an executive level.

Good answer
  • A written escalation matrix is shared with the customer at onboarding.
  • Clear time-based or impact-based triggers are defined for each escalation step.
  • A named executive sponsor (VP/SVP level) is identified at the top of the escalation path.
Red flags
  • Escalation is 'on request' only, with no documented path or triggers.
  • The escalation path leads to a generic group inbox rather than a named role.
  • No executive sponsor is named or available.

5. State your contractual initial response time SLA for each severity level included in the standard (base-tier) support offering.

Why it matters. Response-time SLAs state how quickly the vendor commits to engage on a production incident. Buyers need clarity on whether commitments are contractual or aspirational, and whether the base tier includes meaningful response for non-P1 incidents.

Good answer
  • Provides numeric response times per severity in the base tier.
  • Distinguishes initial response from acknowledgment or auto-reply.
  • Commitments appear in the MSA or support exhibit, not just marketing.
Red flags
  • Only P1 has a stated response time.
  • Auto-acknowledgment counted as 'response'.
  • Response times available only at premium tier.

6. List the geographic regions in which you operate support centers and the languages supported in each region.

Why it matters. Genuine follow-the-sun coverage requires staffed regions, not just an on-call rotation in one timezone. Language support is critical for global customers whose operators may not be native English speakers.

Good answer
  • Staffed regions cover the buyer's operating regions and, together, the full day.
  • Specific cities or countries are named.
  • Languages are listed per region, specifying business or fluent proficiency.
Red flags
  • A single-region support center is presented as 'global'.
  • Off-hours coverage relies on a single on-call engineer, creating a bottleneck.
  • English-only support is offered despite an enterprise customer base across multiple continents.

7. Do you provide a public status page covering platform availability and ongoing incidents, and what is its update cadence during an active incident?

Why it matters. Status pages allow customer operations teams to confirm whether an issue is vendor-side without needing to open a ticket. The update cadence during active incidents sets how quickly the buyer learns of changes.

Good answer
  • A public status page exists with per-component status details.
  • The page offers subscribable updates (e.g., email, RSS, webhook).
  • A committed update interval for active P1 incidents is stated.
Red flags
  • The status page is private or accessible only to logged-in customers.
  • The status page is rarely updated during actual incidents, lagging hours behind reality.
  • The page only shows a single 'all systems operational' status with no component-level granularity.

8. Is a named customer success manager or technical account manager included in the base subscription? If so, what are their guaranteed response times for non-incident related queries?

Why it matters. Dedicated relationship contacts can be gated behind premium tiers, leaving base-tier customers with shared mailboxes. Buyers need to know whether a named contact for strategic and non-urgent issues exists at their planned spending level.

Good answer
  • A named CSM/TAM is included at the planned subscription level.
  • A response time for non-incident requests is committed in writing.
  • A backup contact is identified for when the primary CSM/TAM is unavailable.
Red flags
  • A named contact is available only at the highest premium tier.
  • The assigned CSM rotates frequently, or the vendor does not state how many accounts each CSM supports.
  • There is no response-time expectation for non-incident queries.

The full set: 92 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 141 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Severity definitions & SLA commitments (31)
  • Coverage hours & on-call regions (15)
  • Support channels & ticket-routing (24)
  • Escalation paths & executive sponsorship (22)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 141 changes. Three examples:

Wrong or outdated citation

Draft: Maintenance windows scheduled during a customer's business day are functionally outages.

Now: Maintenance that causes downtime during a customer's business day is an outage for that customer.

Maintenance with no downtime, such as the zero-downtime deployments this question lists as a good signal, is not an outage.

Wrong or outdated citation

Draft: Status page hosted on the vendor's primary domain (affected by same outages)

Now: Status page hosted on the same infrastructure as the service (affected by the same outages)

The risk is shared failure, not the domain name as such. A status page fails with the service when it shares the service's hosting or its DNS; a status subdomain can be served by a third-party status provider but still resolves through the primary domain's DNS. Pass 2 adds DNS to the item. (Source corrected in pass 2.)

Wrong or outdated citation

Draft: Status page hosted on the same infrastructure as the service (affected by the same outages)

Now: Status page shares hosting or DNS with the service (affected by the same outages)

A status subdomain depends on the primary domain's DNS; a DNS failure takes down both the service and its status page.

Questions about this module

How many support & operations questions are there?

92: 30 for the RFI stage, 31 for the RFP and 31 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 141 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related