8 questions from the RFI stage, free
These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.
1. Describe the platform-level controls your product provides to mitigate direct prompt injection (malicious instructions in the user-supplied prompt).
Why it matters. OWASP lists prompt injection first (LLM01:2025) in its Top 10 for LLM Applications 2025. A vendor that pushes the entire defense burden to the customer leaves the buyer to build it. Understanding the platform-side mitigations sets a baseline for the shared-responsibility model.
- Names specific platform mechanisms (e.g., system-prompt isolation, instruction hierarchies, classifier-based detection).
- Acknowledges that no defense is complete and references known residual risk.
- Differentiates between platform controls and customer-configurable controls.
- Claims complete or comprehensive prevention with no caveats.
- Pushes the entire defense to the customer's system prompt.
- Cannot name any specific mitigation mechanism.
2. Describe your internal red-teaming program for AI safety, including team composition, cadence, and scope of attacks tested.
Why it matters. Red-teaming is how a vendor looks for AI-specific failure modes before release. The buyer needs to know who does it, how often, and what it covers.
- Dedicated red team or named external partners.
- Regular cadence (e.g., pre-release plus ongoing).
- Coverage spans prompt injection, jailbreaks, harmful content, and agentic misuse.
- No internal red team and no external partners.
- Red-teaming done only at launch with no ongoing program.
- Scope limited to content-moderation evaluations.
3. Do you operate a public vulnerability-disclosure program that explicitly accepts AI-specific vulnerability reports (e.g., prompt injection, jailbreaks, harmful outputs)?
Why it matters. A security VDP may accept only security vulnerabilities and turn away other AI-specific reports. OpenAI, for example, opened a separate Safety Bug Bounty in March 2026 for abuse and safety reports that do not meet its security program's criteria. A public channel that names AI issues in scope tells researchers and customers where to report them.
- Public VDP page that explicitly lists AI-specific issues in scope.
- Defined safe-harbor language for good-faith research.
- Acknowledgment and remediation SLAs.
- No public disclosure channel for AI issues.
- AI issues explicitly out-of-scope of the security VDP.
- No safe-harbor for researchers.
4. Describe the platform-level mitigations you provide against misuse categories such as CSAM, weapons-of-mass-destruction uplift, malware generation, fraud, and targeted harassment.
Why it matters. Enterprise buyers carry reputational and legal risk if a vendor's product is used to produce illegal or seriously harmful content. Platform-level mitigations apply whatever the customer configures; per-customer policy depends on each customer.
- Names specific mitigations per category, not just generic moderation.
- Distinguishes hard refusals from soft mitigations.
- References recognized harm taxonomies (e.g., MLCommons, NIST AI 600-1).
- Single undifferentiated content filter for all categories.
- Mitigations described only as customer responsibility.
- No mention of CSAM detection or reporting (e.g., reports to NCMEC's CyberTipline, required of US electronic service providers by 18 U.S.C. 2258A).
5. Describe how your product defends against indirect prompt injection, where malicious instructions arrive via retrieved documents, tool outputs, web content, or other untrusted channels.
Why it matters. Indirect prompt injection reaches the model through retrieved documents, web content and tool outputs, which agentic and RAG-enabled systems process by design. NIST AI 100-2 E2025 and MITRE ATLAS (AML.T0051.000 direct, AML.T0051.001 indirect) treat it as a separate attack from direct injection.
- Explicitly distinguishes indirect from direct prompt injection.
- Describes provenance tracking, content sandboxing, or untrusted-input tagging.
- Mentions defenses applied to tool outputs and retrieved content specifically.
- Treats prompt injection as a single undifferentiated category.
- No defenses applied to retrieved or tool-returned content.
- Relies solely on a classifier applied to the user's initial prompt.
6. Do you publish system cards, model cards, or safety evaluation reports for your AI products?
Why it matters. Public safety documentation lets buyers independently assess the rigor of evaluation and the residual risk profile. Without it, the buyer has only the vendor's verbal claims.
- Provides links to current model cards or system cards.
- Reports include quantitative evaluation results, not just narrative.
- Updated for material model or product changes.
- No public safety documentation.
- Documentation is marketing material with no evaluation data.
- Reports are years out of date relative to current products.
7. Do you publish security and safety advisories for AI-specific issues, and where can customers subscribe to them?
Why it matters. Customers must be able to learn promptly of AI-specific vulnerabilities affecting products they depend on. A vendor with no advisory history should explain why.
- Public advisory page with historical AI-specific entries.
- RSS, email, or API subscription channel.
- Advisories include severity, affected versions, and mitigation guidance.
- No advisory history at all.
- Advisories only for traditional CVEs, none AI-specific.
- No subscription mechanism.
8. How do customers report harmful or unsafe outputs they encounter in your product, and what is your response process?
Why it matters. A working harm-reporting channel is essential for closed-loop safety improvement and for customers to fulfill their own incident-response obligations.
- In-product reporting mechanism plus dedicated email or portal.
- Triage and acknowledgment SLAs defined.
- Customer receives status updates and resolution.
- No dedicated channel; harm reports routed through generic support.
- No acknowledgment or feedback to the reporter.
- Vendor cannot describe what happens after a report is filed.
What the audit changed
A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 101 changes. Three examples:
Wrong or outdated citation
Draft: Prompt injection is the most prevalent AI-application vulnerability class identified by OWASP.
Now: OWASP lists prompt injection first (LLM01:2025) in its Top 10 for LLM Applications 2025.
OWASP ranks risks; it does not publish prevalence figures for this list. Source: OWASP Top 10 for LLM Applications 2025, LLM01:2025 Prompt Injection (https://genai.owasp.org/llmrisk/llm01-prompt-injection/).
Wrong or outdated citation
Draft: Names recognized external evaluators or AI safety institutes.
Now: Names recognized external evaluators or government AI evaluation bodies (e.g., UK AI Security Institute).
The UK AI Safety Institute was renamed the AI Security Institute on 14 February 2025, and the US AI Safety Institute was renamed the Center for AI Standards and Innovation (CAISI) on 3 June 2025, so 'AI safety institutes' is out of date. Sources: https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver-plan-for-change; https://www.commerce.gov/news/press-releases/2025/06/statement-us-secretary-commerce-howard-lutnick-transforming-us-ai. (Source corrected in pass 2.)
Wrong or outdated citation
Draft: harm-bench
Now: HarmBench
The benchmark's name is HarmBench (Mazeika et al., 2024). Source: https://arxiv.org/abs/2402.04249.