8 questions from the RFI stage, free
These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.
1. Describe your evaluation methodology for the AI capabilities offered, including which evaluations are run pre-release versus continuously, and who designs them.
Why it matters. Buyers need to know whether quality claims rest on a disciplined, repeatable evaluation process or on ad-hoc internal testing.
- Distinguishes pre-release, release-gate, and continuous evaluations
- Names specific evaluation suites (internal and external) and describes their cadence
- Identifies a dedicated evaluation team or function separate from model development
- Conflates QA testing with model evaluation
- No standing evaluation suite; testing is per-release and ad hoc
- Evaluation design owned solely by the team that ships the model
2. List the public and proprietary benchmarks you report against for each AI capability, including the date of the most recent run and the model version used.
Why it matters. Buyers need a clear inventory of benchmark claims tied to a specific model version. Reporting only favorable benchmarks makes a model look stronger than a full set of results would.
- Provides a table mapping capability → benchmark → version → date → score
- Reports both strengths and weaknesses across a balanced set of benchmarks
- References widely used public benchmarks where relevant (e.g., MMLU-Pro, GPQA, SWE-bench) or third-party evaluation frameworks (e.g., Stanford HELM)
- Only cites benchmarks where the vendor's model leads
- No dates or model versions are attached to scores
- Relies on proprietary-only benchmarks with no methodology disclosure
3. Do you provide an evaluation harness or tooling that allows customers to measure model or product quality on their own data and use cases? Describe what is provided.
Why it matters. Vendor benchmarks are not built from the buyer's workload. A customer-side evaluation capability is essential for evidence-based selection and ongoing monitoring of fitness-for-purpose.
- Provides a documented evaluation SDK, CLI, or hosted UI
- Supports custom datasets, custom scorers, and custom prompts
- Returns per-item results for error analysis, not just aggregate scores
- Customer evaluation requires a professional services engagement to set up
- No documented method to run evaluations on the customer's own data
- Tooling exists but is undocumented, in private beta, or deprecated
4. When you update an underlying model or component that affects output behavior, how do you detect quality regressions before they impact customers?
Why it matters. Hosted model behavior can change between versions: Chen, Zaharia and Zou (2023) measured large changes in GPT-4 and GPT-3.5 on the same tasks between March and June 2023. The vendor's regression detection decides whether the vendor or the customer finds such a change first.
- A standing regression evaluation suite is run on every candidate build
- Pre-deployment shadow traffic comparison is used to vet changes
- Defined quality and performance thresholds must be met to approve a release
- Relies primarily on customer reports to detect regressions
- No standing regression suite exists
- Regression thresholds are set on an ad-hoc basis or can be easily overridden
5. Are your evaluation results reproducible by a third party given the same inputs, model version, and configuration? Describe what is required to reproduce them.
Why it matters. Non-reproducible evaluations cannot be independently verified. Reproducibility forces the vendor to pin model versions, sampling parameters, and prompt scaffolding.
- Confirms reproducibility and details required parameters (model version, decoding settings, seeds)
- Publishes or offers to share evaluation harness code and prompt templates
- Acknowledges sources of non-determinism and describes how they are bounded or measured
- Claims reproducibility but cannot describe how to achieve it
- Refuses to share evaluation prompts or scoring rubrics under NDA
- Dismisses non-determinism as a reason not to pursue reproducibility
6. Explain your selection criteria for which benchmarks to report publicly, and how you handle results from benchmarks where your performance is weak.
Why it matters. A vendor that publishes only favorable results hides the weak areas a buyer needs to plan around.
- Documents selection criteria independent of the expected score
- Publishes weak results alongside strong ones for transparency
- Distinguishes internal regression tracking benchmarks from external marketing benchmarks
- Benchmark selection is driven by the marketing team
- Cannot name a single industry-standard benchmark where their product underperforms competitors
- Withholds results pending 'improvement'
7. How can customers integrate their own ground-truth datasets and scoring rubrics into your platform's evaluation runs?
Why it matters. Customer ground truth measures fitness for the buyer's own use case. Vendors that make ingestion and rubric definition easy enable buyers to evaluate against the use case that actually matters.
- Documents supported import formats (e.g., JSONL, CSV)
- Supports custom rubrics and scoring functions, either in UI or via code
- Provides for version control of datasets and rubrics
- Only fixed, vendor-defined rubric categories are supported
- No contractual guarantees for the isolation of customer evaluation data
- Requires manual data loading via support tickets or professional services
8. Do you offer customers the ability to pin a specific model version for a defined support window?
Why it matters. Version pinning protects buyers from involuntary, breaking changes in model behavior. Without it, even passing customer evaluations have a short shelf life.
- Offers named or dated versions with a documented, committed support window
- States a committed deprecation notice period
- Offers migration tooling and parallel-run capabilities to ease transitions
- Only a 'latest' version is available
- No deprecation notice period is stated, or it is shorter than the time the buyer needs to re-validate
- No support for validating a new version before migrating
What the audit changed
A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 131 changes. Three examples:
Wrong or outdated citation
Draft: References widely accepted public benchmarks where relevant (e.g., HELM, MMLU, HumanEval)
Now: References widely used public benchmarks where relevant (e.g., MMLU-Pro, GPQA, SWE-bench) or third-party evaluation frameworks (e.g., Stanford HELM)
HELM (Holistic Evaluation of Language Models, Stanford CRFM) is an evaluation framework and leaderboard, not a benchmark (https://crfm.stanford.edu/helm/). Hugging Face replaced MMLU with MMLU-Pro in its Open LLM Leaderboard v2 (June 2024) because MMLU had proved noisy and too easy (https://huggingface.co/spaces/open-llm-leaderboard/blog). The benchmarks named are examples, not recommendations. (Source corrected in pass 3.)
Wrong or outdated citation
Draft: (e.g., Stanford HELM, Hugging Face Open LLM Leaderboard, Chatbot Arena)
Now: (e.g., Stanford HELM, Arena, formerly LMSYS Chatbot Arena)
The Hugging Face Open LLM Leaderboard was retired on March 13, 2025 and accepts no new models (https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135). Chatbot Arena moved from LMSYS to its own site, lmarena.ai, on September 20, 2024 (https://www.lmsys.org/blog/2024-09-20-arena-new-site/), operating as LMArena, and renamed itself Arena (arena.ai) on January 28, 2026 (https://arena.ai/blog/lmarena-is-now-arena). (Source corrected in pass 2.)
Wrong or outdated citation
Draft: Engages with government-affiliated AI safety institutes where applicable.
Now: Engages with government AI evaluation bodies where applicable (e.g., UK AI Security Institute, US Center for AI Standards and Innovation).
The US Center for AI Standards and Innovation (CAISI) replaced the US AI Safety Institute in June 2025 (https://www.commerce.gov/news/press-releases/2025/06/statement-us-secretary-commerce-howard-lutnick-transforming-us-ai). After Executive Order 14434 of September 29, 2026, NIST renamed it the Center for Advancing Innovation and Standards for Super Intelligence (CAISSI) (https://www.nist.gov/caissi, https://www.nist.gov/super-intelligence). The UK AI Safety Institute was renamed the AI Security Institute on February 14, 2025 (https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver-plan-for-change). Neither body is a regulator. (Source corrected in pass 2.)