CIOPages
All RFP question modules

AI Governance

AI safety & responsible AI questions to ask a software vendor

Questions on the vendor's AI safety controls: defenses against prompt injection and unsafe tool use, red-teaming, vulnerability disclosure, harm reporting, abuse mitigation and guidance for safe deployment. Agent permissions are in Agentic safety & autonomy controls.

89
questions
22
RFI
35
RFP
32
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Describe the platform-level controls your product provides to mitigate direct prompt injection (malicious instructions in the user-supplied prompt).

Why it matters. OWASP lists prompt injection first (LLM01:2025) in its Top 10 for LLM Applications 2025. A vendor that pushes the entire defense burden to the customer leaves the buyer to build it. Understanding the platform-side mitigations sets a baseline for the shared-responsibility model.

Good answer
  • Names specific platform mechanisms (e.g., system-prompt isolation, instruction hierarchies, classifier-based detection).
  • Acknowledges that no defense is complete and references known residual risk.
  • Differentiates between platform controls and customer-configurable controls.
Red flags
  • Claims complete or comprehensive prevention with no caveats.
  • Pushes the entire defense to the customer's system prompt.
  • Cannot name any specific mitigation mechanism.

2. Describe your internal red-teaming program for AI safety, including team composition, cadence, and scope of attacks tested.

Why it matters. Red-teaming is how a vendor looks for AI-specific failure modes before release. The buyer needs to know who does it, how often, and what it covers.

Good answer
  • Dedicated red team or named external partners.
  • Regular cadence (e.g., pre-release plus ongoing).
  • Coverage spans prompt injection, jailbreaks, harmful content, and agentic misuse.
Red flags
  • No internal red team and no external partners.
  • Red-teaming done only at launch with no ongoing program.
  • Scope limited to content-moderation evaluations.

3. Do you operate a public vulnerability-disclosure program that explicitly accepts AI-specific vulnerability reports (e.g., prompt injection, jailbreaks, harmful outputs)?

Why it matters. A security VDP may accept only security vulnerabilities and turn away other AI-specific reports. OpenAI, for example, opened a separate Safety Bug Bounty in March 2026 for abuse and safety reports that do not meet its security program's criteria. A public channel that names AI issues in scope tells researchers and customers where to report them.

Good answer
  • Public VDP page that explicitly lists AI-specific issues in scope.
  • Defined safe-harbor language for good-faith research.
  • Acknowledgment and remediation SLAs.
Red flags
  • No public disclosure channel for AI issues.
  • AI issues explicitly out-of-scope of the security VDP.
  • No safe-harbor for researchers.

4. Describe the platform-level mitigations you provide against misuse categories such as CSAM, weapons-of-mass-destruction uplift, malware generation, fraud, and targeted harassment.

Why it matters. Enterprise buyers carry reputational and legal risk if a vendor's product is used to produce illegal or seriously harmful content. Platform-level mitigations apply whatever the customer configures; per-customer policy depends on each customer.

Good answer
  • Names specific mitigations per category, not just generic moderation.
  • Distinguishes hard refusals from soft mitigations.
  • References recognized harm taxonomies (e.g., MLCommons, NIST AI 600-1).
Red flags
  • Single undifferentiated content filter for all categories.
  • Mitigations described only as customer responsibility.
  • No mention of CSAM detection or reporting (e.g., reports to NCMEC's CyberTipline, required of US electronic service providers by 18 U.S.C. 2258A).

5. Describe how your product defends against indirect prompt injection, where malicious instructions arrive via retrieved documents, tool outputs, web content, or other untrusted channels.

Why it matters. Indirect prompt injection reaches the model through retrieved documents, web content and tool outputs, which agentic and RAG-enabled systems process by design. NIST AI 100-2 E2025 and MITRE ATLAS (AML.T0051.000 direct, AML.T0051.001 indirect) treat it as a separate attack from direct injection.

Good answer
  • Explicitly distinguishes indirect from direct prompt injection.
  • Describes provenance tracking, content sandboxing, or untrusted-input tagging.
  • Mentions defenses applied to tool outputs and retrieved content specifically.
Red flags
  • Treats prompt injection as a single undifferentiated category.
  • No defenses applied to retrieved or tool-returned content.
  • Relies solely on a classifier applied to the user's initial prompt.

6. Do you publish system cards, model cards, or safety evaluation reports for your AI products?

Why it matters. Public safety documentation lets buyers independently assess the rigor of evaluation and the residual risk profile. Without it, the buyer has only the vendor's verbal claims.

Good answer
  • Provides links to current model cards or system cards.
  • Reports include quantitative evaluation results, not just narrative.
  • Updated for material model or product changes.
Red flags
  • No public safety documentation.
  • Documentation is marketing material with no evaluation data.
  • Reports are years out of date relative to current products.

7. Do you publish security and safety advisories for AI-specific issues, and where can customers subscribe to them?

Why it matters. Customers must be able to learn promptly of AI-specific vulnerabilities affecting products they depend on. A vendor with no advisory history should explain why.

Good answer
  • Public advisory page with historical AI-specific entries.
  • RSS, email, or API subscription channel.
  • Advisories include severity, affected versions, and mitigation guidance.
Red flags
  • No advisory history at all.
  • Advisories only for traditional CVEs, none AI-specific.
  • No subscription mechanism.

8. How do customers report harmful or unsafe outputs they encounter in your product, and what is your response process?

Why it matters. A working harm-reporting channel is essential for closed-loop safety improvement and for customers to fulfill their own incident-response obligations.

Good answer
  • In-product reporting mechanism plus dedicated email or portal.
  • Triage and acknowledgment SLAs defined.
  • Customer receives status updates and resolution.
Red flags
  • No dedicated channel; harm reports routed through generic support.
  • No acknowledgment or feedback to the reporter.
  • Vendor cannot describe what happens after a report is filed.

The full set: 89 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 101 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Prompt injection & untrusted input (17)
  • Red-teaming & adversarial evaluation (14)
  • Vulnerability disclosure & advisories (14)
  • Abuse and misuse mitigations (24)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 101 changes. Three examples:

Wrong or outdated citation

Draft: Prompt injection is the most prevalent AI-application vulnerability class identified by OWASP.

Now: OWASP lists prompt injection first (LLM01:2025) in its Top 10 for LLM Applications 2025.

OWASP ranks risks; it does not publish prevalence figures for this list. Source: OWASP Top 10 for LLM Applications 2025, LLM01:2025 Prompt Injection (https://genai.owasp.org/llmrisk/llm01-prompt-injection/).

Wrong or outdated citation

Draft: Names recognized external evaluators or AI safety institutes.

Now: Names recognized external evaluators or government AI evaluation bodies (e.g., UK AI Security Institute).

The UK AI Safety Institute was renamed the AI Security Institute on 14 February 2025, and the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI) on 3 June 2025, so 'AI safety institutes' is out of date. Sources: https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver-plan-for-change; https://www.commerce.gov/news/press-releases/2025/06/statement-us-secretary-commerce-howard-lutnick-transforming-us-ai. (Source corrected in pass 2.)

Wrong or outdated citation

Draft: harm-bench

Now: HarmBench

The benchmark's name is HarmBench (Mazeika et al., 2024). Source: https://arxiv.org/abs/2402.04249.

Questions about this module

How many ai safety & responsible ai questions are there?

89: 22 for the RFI stage, 35 for the RFP and 32 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 101 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related