CIOPages
All RFP question modules

AI Governance

AI performance, evaluation & monitoring questions to ask a software vendor

Questions on how the vendor measures AI quality: evaluation methods, published benchmarks, support for testing on your own use cases, and how regressions are caught after release.

102
questions
27
RFI
36
RFP
39
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Describe your evaluation methodology for the AI capabilities offered, including which evaluations are run pre-release versus continuously, and who designs them.

Why it matters. Buyers need to know whether quality claims rest on a disciplined, repeatable evaluation process or on ad-hoc internal testing.

Good answer
  • Distinguishes pre-release, release-gate, and continuous evaluations
  • Names specific evaluation suites (internal and external) and describes their cadence
  • Identifies a dedicated evaluation team or function separate from model development
Red flags
  • Conflates QA testing with model evaluation
  • No standing evaluation suite; testing is per-release and ad hoc
  • Evaluation design owned solely by the team that ships the model

2. List the public and proprietary benchmarks you report against for each AI capability, including the date of the most recent run and the model version used.

Why it matters. Buyers need a clear inventory of benchmark claims tied to a specific model version. Reporting only favorable benchmarks makes a model look stronger than a full set of results would.

Good answer
  • Provides a table mapping capability → benchmark → version → date → score
  • Reports both strengths and weaknesses across a balanced set of benchmarks
  • References widely used public benchmarks where relevant (e.g., MMLU-Pro, GPQA, SWE-bench) or third-party evaluation frameworks (e.g., Stanford HELM)
Red flags
  • Only cites benchmarks where the vendor's model leads
  • No dates or model versions are attached to scores
  • Relies on proprietary-only benchmarks with no methodology disclosure

3. Do you provide an evaluation harness or tooling that allows customers to measure model or product quality on their own data and use cases? Describe what is provided.

Why it matters. Vendor benchmarks are not built from the buyer's workload. A customer-side evaluation capability is essential for evidence-based selection and ongoing monitoring of fitness-for-purpose.

Good answer
  • Provides a documented evaluation SDK, CLI, or hosted UI
  • Supports custom datasets, custom scorers, and custom prompts
  • Returns per-item results for error analysis, not just aggregate scores
Red flags
  • Customer evaluation requires a professional services engagement to set up
  • No documented method to run evaluations on the customer's own data
  • Tooling exists but is undocumented, in private beta, or deprecated

4. When you update an underlying model or component that affects output behavior, how do you detect quality regressions before they impact customers?

Why it matters. Hosted model behavior can change between versions: Chen, Zaharia and Zou (2023) measured large changes in GPT-4 and GPT-3.5 on the same tasks between March and June 2023. The vendor's regression detection decides whether the vendor or the customer finds such a change first.

Good answer
  • A standing regression evaluation suite is run on every candidate build
  • Pre-deployment shadow traffic comparison is used to vet changes
  • Defined quality and performance thresholds must be met to approve a release
Red flags
  • Relies primarily on customer reports to detect regressions
  • No standing regression suite exists
  • Regression thresholds are set on an ad-hoc basis or can be easily overridden

5. Are your evaluation results reproducible by a third party given the same inputs, model version, and configuration? Describe what is required to reproduce them.

Why it matters. Non-reproducible evaluations cannot be independently verified. Reproducibility forces the vendor to pin model versions, sampling parameters, and prompt scaffolding.

Good answer
  • Confirms reproducibility and details required parameters (model version, decoding settings, seeds)
  • Publishes or offers to share evaluation harness code and prompt templates
  • Acknowledges sources of non-determinism and describes how they are bounded or measured
Red flags
  • Claims reproducibility but cannot describe how to achieve it
  • Refuses to share evaluation prompts or scoring rubrics under NDA
  • Dismisses non-determinism as a reason not to pursue reproducibility

6. Explain your selection criteria for which benchmarks to report publicly, and how you handle results from benchmarks where your performance is weak.

Why it matters. A vendor that publishes only favorable results hides the weak areas a buyer needs to plan around.

Good answer
  • Documents selection criteria independent of the expected score
  • Publishes weak results alongside strong ones for transparency
  • Distinguishes internal regression tracking benchmarks from external marketing benchmarks
Red flags
  • Benchmark selection is driven by the marketing team
  • Cannot name a single industry-standard benchmark where their product underperforms competitors
  • Withholds results pending 'improvement'

7. How can customers integrate their own ground-truth datasets and scoring rubrics into your platform's evaluation runs?

Why it matters. Customer ground truth measures fitness for the buyer's own use case. Vendors that make ingestion and rubric definition easy enable buyers to evaluate against the use case that actually matters.

Good answer
  • Documents supported import formats (e.g., JSONL, CSV)
  • Supports custom rubrics and scoring functions, either in UI or via code
  • Provides for version control of datasets and rubrics
Red flags
  • Only fixed, vendor-defined rubric categories are supported
  • No contractual guarantees for the isolation of customer evaluation data
  • Requires manual data loading via support tickets or professional services

8. Do you offer customers the ability to pin a specific model version for a defined support window?

Why it matters. Version pinning protects buyers from involuntary, breaking changes in model behavior. Without it, even passing customer evaluations have a short shelf life.

Good answer
  • Offers named or dated versions with a documented, committed support window
  • States a committed deprecation notice period
  • Offers migration tooling and parallel-run capabilities to ease transitions
Red flags
  • Only a 'latest' version is available
  • No deprecation notice period is stated, or it is shorter than the time the buyer needs to re-validate
  • No support for validating a new version before migrating

The full set: 102 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 131 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Evaluation methodology & reproducibility (29)
  • Benchmark coverage & selection discipline (23)
  • Customer-side evaluation harness & tooling (26)
  • Regression detection across model updates (24)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 131 changes. Three examples:

Wrong or outdated citation

Draft: References widely accepted public benchmarks where relevant (e.g., HELM, MMLU, HumanEval)

Now: References widely used public benchmarks where relevant (e.g., MMLU-Pro, GPQA, SWE-bench) or third-party evaluation frameworks (e.g., Stanford HELM)

HELM (Holistic Evaluation of Language Models, Stanford CRFM) is an evaluation framework and leaderboard, not a benchmark (https://crfm.stanford.edu/helm/). Hugging Face replaced MMLU with MMLU-Pro in its Open LLM Leaderboard v2 (June 2024) because MMLU had proved noisy and too easy (https://huggingface.co/spaces/open-llm-leaderboard/blog). The benchmarks named are examples, not recommendations. (Source corrected in pass 3.)

Wrong or outdated citation

Draft: (e.g., Stanford HELM, Hugging Face Open LLM Leaderboard, Chatbot Arena)

Now: (e.g., Stanford HELM, Arena, formerly LMSYS Chatbot Arena)

The Hugging Face Open LLM Leaderboard was retired on March 13, 2025 and accepts no new models (https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135). Chatbot Arena moved from LMSYS to its own site, lmarena.ai, on September 20, 2024 (https://www.lmsys.org/blog/2024-09-20-arena-new-site/), operating as LMArena, and renamed itself Arena (arena.ai) on January 28, 2026 (https://arena.ai/blog/lmarena-is-now-arena). (Source corrected in pass 2.)

Wrong or outdated citation

Draft: Engages with government-affiliated AI safety institutes where applicable.

Now: Engages with government AI evaluation bodies where applicable (e.g., UK AI Security Institute, US Center for AI Standards and Innovation).

The US Center for AI Standards and Innovation (CAISI) replaced the US AI Safety Institute in June 2025 (https://www.commerce.gov/news/press-releases/2025/06/statement-us-secretary-commerce-howard-lutnick-transforming-us-ai). After Executive Order 14434 of September 29, 2026, NIST renamed it the Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) (https://www.nist.gov/caissi, https://www.nist.gov/super-intelligence). The UK AI Safety Institute was renamed the AI Security Institute on February 14, 2025 (https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver-plan-for-change). Neither body is a regulator. (Source corrected in pass 2.)

Questions about this module

How many ai performance, evaluation & monitoring questions are there?

102: 27 for the RFI stage, 36 for the RFP and 39 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 131 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related