CIOPages
All RFP question modules

AI Governance

Human oversight & escalation questions to ask a software vendor

Questions on keeping a person in the loop: approval steps for automated decisions, escalation, manual override, and an audit trail that shows what the AI did and why.

94
questions
28
RFI
34
RFP
32
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Describe the human-in-the-loop (HITL) controls your platform provides, including the granularity at which a customer administrator can require human review (e.g. per-tenant, per-workflow, per-action, per-decision-type).

Why it matters. HITL that can only be set per session or per tenant cannot require review of a single high-risk action. Granularity determines whether the buyer can actually deploy in higher-risk contexts. The NIST AI RMF asks that processes for human oversight be defined, assessed and documented (MAP 3.5).

Good answer
  • Names specific HITL configuration scopes (workflow, action, decision type, risk tier)
  • Describes admin UI or API to configure HITL without code changes
  • Differentiates HITL for read-only vs. state-changing actions
Red flags
  • HITL only available at session or tenant level
  • HITL described as a roadmap item rather than current capability
  • No way to require review for specific high-risk actions

2. Describe the built-in escalation and approval workflows your platform supports, including how reviewers are assigned, notified, and how multi-step or multi-approver chains are configured.

Why it matters. Enterprise oversight can require multi-party review and segregation of duties. Vendors without configurable workflow primitives push that burden onto customer integrations, where an escalation path may go unenforced.

Good answer
  • Native workflow engine with configurable approver chains
  • Integration with identity provider groups and roles
  • Notifications via email, chat (Slack/Teams), and webhooks
Red flags
  • Escalation described as a customer integration responsibility only
  • No native notification mechanism
  • Approver assignment hard-coded per tenant

3. Describe the manual override capabilities available to administrators and authorized users, including the ability to pause, stop, or terminate an in-progress automated workflow or agent run.

Why it matters. EU AI Act Article 14(4)(e) requires a high-risk AI system to let the people overseeing it intervene or interrupt it through a 'stop' button or similar procedure; for Annex III systems this applies from December 2, 2027. NIST AI RMF MANAGE 2.4 calls for mechanisms to supersede, disengage or deactivate an AI system.

Good answer
  • Per-run pause/stop controls available in UI and API
  • Tenant-wide emergency stop with documented blast radius
  • Controls available to designated roles with audit logging
Red flags
  • No way to stop an in-progress run
  • Stop requires a support ticket
  • Stop has undocumented side effects

4. Describe the audit trail your platform produces for automated AI decisions, including the fields captured (input, output, model/version, prompt or policy in effect, timestamp, identity).

Why it matters. Reconstructing what the AI did and why is essential for audit, dispute, and post-incident analysis. Sparse or missing fields make this impossible after the fact.

Good answer
  • Captures input, output, model version, prompt template/version, timestamp
  • Includes acting identity (user, service, agent) and tenant
  • Includes policy or guardrail decisions applied
Red flags
  • Audit log lacks model version or prompt version
  • Inputs and outputs not retained
  • No identity linkage

5. Can a customer require human approval before the system executes irreversible or state-changing operations (e.g. external API writes, financial transactions, communications sent on behalf of a user)?

Why it matters. Agentic AI systems can take real-world actions. Buyers need to know whether they can gate irreversible side effects behind human approval without sacrificing the rest of the workflow.

Good answer
  • Explicit support for approval gates on tool calls or actions
  • Distinguishes reversible from irreversible actions
  • Allows policy-based selection of which actions require approval
Red flags
  • All-or-nothing automation with no per-action gating
  • Approval only possible by disabling automation entirely
  • No distinction between read and write tool calls

6. How does your platform enforce segregation of duties between the user who initiates an action, the reviewer who approves it, and the administrator who configures the policy?

Why it matters. Segregation of duties is a fundamental control for regulated industries and is a control in ISO/IEC 27001:2022 (Annex A 5.3), and auditors test for it in SOX Section 404 and SOC 2 examinations. Without platform-level enforcement, buyers cannot rely on policy alone.

Good answer
  • RBAC distinguishes initiator, reviewer, and administrator roles
  • Platform prevents self-approval by default
  • Configuration changes logged separately from operational actions
Red flags
  • Same role can both initiate and approve actions
  • Self-approval prevention requires customer-side enforcement
  • Configuration audit log not separated from data plane

7. Can a human reviewer override or edit an automated decision after it has been issued, and how is the override propagated to downstream systems that may have consumed the original output?

Why it matters. Override that is not propagated downstream creates inconsistent state and undermines correction. Buyers need to understand the full reversal path, not just the UI affordance.

Good answer
  • Override mechanism with linkage to original decision
  • Webhooks or events emitted on override for downstream sync
  • Documented behavior for already-actioned outputs
Red flags
  • Override changes only the UI state, not downstream consumers
  • No event emitted on override
  • Override and original are separate uncorrelated records

8. How does the audit trail link human review, approval, and override actions to the underlying automated decisions they relate to?

Why it matters. An audit trail of automated decisions and a separate log of human actions are not sufficient — the linkage is what enables reconstructing accountability.

Good answer
  • Each human action references the originating decision ID
  • Reviewer identity, role, action, and rationale captured
  • Chain of decisions queryable end-to-end
Red flags
  • Human action log separate with no correlation key
  • Reviewer rationale stored only as free-text email
  • No way to query 'what happened to decision X'

The full set: 94 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 94 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Human-in-the-loop mechanisms & granularity (28)
  • Escalation, review & approval workflows (19)
  • Manual override & decision reversal (19)
  • Audit trail of automated & human decisions (28)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 94 changes. Three examples:

Wrong or outdated citation

Draft: The NIST AI RMF treats human oversight as a core function that must be operationally realisable.

Now: The NIST AI RMF asks that processes for human oversight be defined, assessed and documented (MAP 3.5).

The AI RMF 1.0 core functions are Govern, Map, Measure and Manage; human oversight is not one of them. MAP 3.5 and GOVERN 3.2 address it (https://airc.nist.gov/airmf-resources/airmf/5-sec-core/). Also removes the British spelling 'realisable'.

Wrong or outdated citation

Draft: is required by frameworks such as SOX, SOC 2, and ISO 27001

Now: is a control in ISO/IEC 27001:2022 (Annex A 5.3), and auditors test for it in SOX Section 404 and SOC 2 examinations

SOX does not name segregation of duties; it comes in through internal-control assessments under Section 404. SOC 2 is an attestation report against the Trust Services Criteria, not a requirements framework; CC5.1 has the point of focus 'Addresses Segregation of Duties' (https://assets.ctfassets.net/rb9cdnjh59cm/5jT1narHNQNzt4JGlkd1gr/248661d08e42531329d147782a6f8854/Trust-services-criteria.pdf). ISO/IEC 27001:2022 Annex A 5.3 is 'Segregation of duties', confirmed by two secondary sources because the standard is paywalled: https://iseoblue.com/iso-27001/annex-a/control-5-3/ and https://www.isms.online/iso-27001/annex-a-2022/5-3-segregation-of-duties-2022/. (Source corrected in pass 3.)

Wrong or outdated citation

Draft: Customer-side verification or BYOK retention options

Now: Customer-side integrity verification, or export to customer-controlled storage

BYOK (bring your own key) is encryption-key management; it is not a retention or integrity control.

Questions about this module

How many human oversight & escalation questions are there?

94: 28 for the RFI stage, 34 for the RFP and 32 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 94 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related