CIOPages
All RFP question modules

AI Governance

Bias, fairness & content provenance questions to ask a software vendor

Questions on bias and fairness testing, how results are measured across demographic groups, where training data came from, content provenance, and what the vendor does when you find bias after deployment.

105
questions
25
RFI
41
RFP
39
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Describe your methodology for evaluating bias in the model(s) powering your product. Include the benchmarks and datasets used, the cadence of re-evaluation, and triggers for out-of-cycle reviews.

Why it matters. Bias claims without documented methodology cannot be validated or compared across vendors. Buyers need to know whether the vendor treats bias evaluation as a continuous process tied to model changes, not a one-off exercise.

Good answer
  • Names specific public or internal benchmarks used (e.g. BBQ, BOLD, CrowS-Pairs, WinoBias).
  • Describes statistical tests and significance thresholds used.
  • Defines a regular cadence for re-evaluation (e.g., per model release, quarterly).
Red flags
  • Generic claims of 'rigorous testing' with no named benchmarks or metrics.
  • Treats bias evaluation as a one-time pre-launch activity with no defined re-evaluation cadence.
  • Cadence is described as 'as needed' with no objective triggers.

2. Which fairness metrics (e.g. demographic parity, equalized odds, predictive parity, calibration) do you report for your model(s), and why were those metrics chosen?

Why it matters. Common fairness metrics cannot all be satisfied at once when base rates differ between groups and prediction is imperfect (Kleinberg, Mullainathan and Raghavan, 2016; Chouldechova, 2017). The vendor's choice of metric is a trade-off the buyer should see.

Good answer
  • Names specific metrics with definitions
  • Explains the trade-off accepted and why it fits the product's use case
  • Acknowledges incompatibility theorems (e.g. Chouldechova, Kleinberg)
Red flags
  • Generic 'fairness' claims without naming a metric
  • Claims to satisfy all fairness criteria simultaneously
  • No justification for the chosen metric

3. Describe the sources and categories of training data used for the model(s) powering your product (e.g. licensed datasets, web crawl, customer data, synthetic data, partner data).

Why it matters. Training data provenance bears on bias, IP, and privacy risk. Without a description of data sources, the buyer cannot assess bias or copyright exposure.

Good answer
  • Names categories of data with approximate proportions
  • Distinguishes pretraining, fine-tuning, and RLHF/DPO/post-training data
  • Discloses whether customer data is used for training (and opt-in/opt-out)
Red flags
  • Characterizes data only as 'publicly available' with no detail
  • Cannot distinguish pretraining from post-training data sources
  • Will not disclose whether customer data is used

4. Describe the customer-facing process for reporting suspected bias or fairness incidents observed in production, including intake channel, SLA, and escalation path.

Why it matters. Some bias issues surface only after deployment in real customer contexts. Without a defined intake and SLA, customers have no recourse and the vendor has no feedback loop.

Good answer
  • Defined intake channel (portal, ticket type, contact)
  • Stated SLA for acknowledgment and triage
  • Documented escalation path to engineering and policy teams
Red flags
  • No dedicated channel for bias reports
  • Treats bias issues as ordinary support tickets
  • No SLA for acknowledgment

5. Do you publish a model card or system card for each model that customers can consume through your product?

Why it matters. Model cards (Mitchell et al., 'Model Cards for Model Reporting', 2019) and system cards disclose intended use, limitations, evaluation results, and known biases. Without them, buyers cannot compare vendors on the same fields.

Good answer
  • Public model or system card available per model version
  • Card includes bias evaluation results, intended use, and known limitations
  • Card is versioned and updated with each model release
Red flags
  • No model card or system card is published
  • Card exists but contains no quantitative evaluation results
  • Card is not versioned or has not been updated for the current model

6. Across which demographic slices (e.g. gender, race/ethnicity, age, language, disability, geography) do you evaluate model performance, and how are those slices defined?

Why it matters. Aggregate accuracy can hide significant disparities across populations. Buyers need to know which slices the vendor evaluates and whether those slices match the populations the customer serves.

Good answer
  • Lists specific demographic dimensions evaluated
  • Describes how slice membership is determined (self-reported, inferred, synthetic personas)
  • Covers intersectional slices, not just marginal
Red flags
  • Only evaluates aggregate accuracy
  • Cannot list demographic dimensions
  • Slices are limited to US-centric categories despite global deployment

7. If you build on third-party foundation models (open-weight or closed), name the model(s), the providers, and the version pinning customers can rely on.

Why it matters. A product vendor that integrates an upstream foundation model did not control that model's bias and provenance profile. Buyers need to know which upstream models are in scope so they can pull the relevant model cards and assess inherited risk.

Good answer
  • Names each upstream model and provider explicitly
  • Specifies version pinning and update policy
  • Identifies which product features use which model
Red flags
  • Refuses to name upstream models
  • Cannot specify version pinning
  • Routes silently between models with no customer visibility

8. What remediation options are available to customers when bias is detected post-deployment (e.g. prompt-layer mitigations, fine-tuning, model swap, configuration controls)?

Why it matters. Buyers need to know what concrete controls they have to reduce harm while the underlying model is updated.

Good answer
  • Lists concrete remediation levers available to customers
  • Distinguishes vendor-side fixes from customer-configurable controls
  • References published guidance for each lever
Red flags
  • Only remediation is to wait for the next model version
  • No customer-configurable controls
  • Cannot give timelines

The full set: 105 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 154 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Bias evaluation methodology & cadence (22)
  • Fairness metrics reported & demographic slices (24)
  • Training & evaluation data provenance (32)
  • Customer-facing remediation & escalation (27)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 154 changes. Three examples:

Wrong or outdated citation

Draft: Sector-specific regulators impose distinct fairness standards (e.g. four-fifths rule in employment, ECOA in credit).

Now: Sector-specific laws and guidelines set distinct fairness tests (e.g. the four-fifths rule of thumb for adverse impact in the Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4(D); the Equal Credit Opportunity Act in credit).

The four-fifths rule is a rule of thumb in the Uniform Guidelines, not a regulator's standard, and ECOA is a statute. Source: https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607

Wrong or outdated citation

Draft: Distinguishes high-risk from low-risk use under the EU AI Act

Now: Distinguishes high-risk use under the EU AI Act from other use

The EU AI Act has no 'low-risk' category; it sets prohibited practices, high-risk systems and transparency obligations. Source: Regulation (EU) 2024/1689, https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng

Wrong or outdated citation

Draft: Fairness metrics encode incompatible normative trade-offs and cannot all be satisfied simultaneously. A vendor that can name its chosen metrics and justify the trade-off has thought through the problem; one that cannot has not.

Now: Common fairness metrics cannot all be satisfied at once when base rates differ between groups and prediction is imperfect (Kleinberg, Mullainathan and Raghavan, 2016; Chouldechova, 2017). The vendor's choice of metric is a trade-off the buyer should see.

The incompatibility holds only when base rates differ and the predictor is imperfect (Kleinberg et al., 'Inherent Trade-Offs in the Fair Determination of Risk Scores', 2016; Chouldechova, 'Fair prediction with disparate impact', 2017). The last sentence was a judgment stated as fact.

Questions about this module

How many bias, fairness & content provenance questions are there?

105: 25 for the RFI stage, 41 for the RFP and 39 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 154 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related