8 questions from the RFI stage, free
These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.
1. Describe your methodology for evaluating bias in the model(s) powering your product. Include the benchmarks and datasets used, the cadence of re-evaluation, and triggers for out-of-cycle reviews.
Why it matters. Bias claims without documented methodology cannot be validated or compared across vendors. Buyers need to know whether the vendor treats bias evaluation as a continuous process tied to model changes, not a one-off exercise.
- Names specific public or internal benchmarks used (e.g. BBQ, BOLD, CrowS-Pairs, WinoBias).
- Describes statistical tests and significance thresholds used.
- Defines a regular cadence for re-evaluation (e.g., per model release, quarterly).
- Generic claims of 'rigorous testing' with no named benchmarks or metrics.
- Treats bias evaluation as a one-time pre-launch activity with no defined re-evaluation cadence.
- Cadence is described as 'as needed' with no objective triggers.
2. Which fairness metrics (e.g. demographic parity, equalized odds, predictive parity, calibration) do you report for your model(s), and why were those metrics chosen?
Why it matters. Common fairness metrics cannot all be satisfied at once when base rates differ between groups and prediction is imperfect (Kleinberg, Mullainathan and Raghavan, 2016; Chouldechova, 2017). The vendor's choice of metric is a trade-off the buyer should see.
- Names specific metrics with definitions
- Explains the trade-off accepted and why it fits the product's use case
- Acknowledges incompatibility theorems (e.g. Chouldechova, Kleinberg)
- Generic 'fairness' claims without naming a metric
- Claims to satisfy all fairness criteria simultaneously
- No justification for the chosen metric
3. Describe the sources and categories of training data used for the model(s) powering your product (e.g. licensed datasets, web crawl, customer data, synthetic data, partner data).
Why it matters. Training data provenance bears on bias, IP, and privacy risk. Without a description of data sources, the buyer cannot assess bias or copyright exposure.
- Names categories of data with approximate proportions
- Distinguishes pretraining, fine-tuning, and RLHF/DPO/post-training data
- Discloses whether customer data is used for training (and opt-in/opt-out)
- Characterizes data only as 'publicly available' with no detail
- Cannot distinguish pretraining from post-training data sources
- Will not disclose whether customer data is used
4. Describe the customer-facing process for reporting suspected bias or fairness incidents observed in production, including intake channel, SLA, and escalation path.
Why it matters. Some bias issues surface only after deployment in real customer contexts. Without a defined intake and SLA, customers have no recourse and the vendor has no feedback loop.
- Defined intake channel (portal, ticket type, contact)
- Stated SLA for acknowledgment and triage
- Documented escalation path to engineering and policy teams
- No dedicated channel for bias reports
- Treats bias issues as ordinary support tickets
- No SLA for acknowledgment
5. Do you publish a model card or system card for each model that customers can consume through your product?
Why it matters. Model cards (Mitchell et al., 'Model Cards for Model Reporting', 2019) and system cards disclose intended use, limitations, evaluation results, and known biases. Without them, buyers cannot compare vendors on the same fields.
- Public model or system card available per model version
- Card includes bias evaluation results, intended use, and known limitations
- Card is versioned and updated with each model release
- No model card or system card is published
- Card exists but contains no quantitative evaluation results
- Card is not versioned or has not been updated for the current model
6. Across which demographic slices (e.g. gender, race/ethnicity, age, language, disability, geography) do you evaluate model performance, and how are those slices defined?
Why it matters. Aggregate accuracy can hide significant disparities across populations. Buyers need to know which slices the vendor evaluates and whether those slices match the populations the customer serves.
- Lists specific demographic dimensions evaluated
- Describes how slice membership is determined (self-reported, inferred, synthetic personas)
- Covers intersectional slices, not just marginal
- Only evaluates aggregate accuracy
- Cannot list demographic dimensions
- Slices are limited to US-centric categories despite global deployment
7. If you build on third-party foundation models (open-weight or closed), name the model(s), the providers, and the version pinning customers can rely on.
Why it matters. A product vendor that integrates an upstream foundation model did not control that model's bias and provenance profile. Buyers need to know which upstream models are in scope so they can pull the relevant model cards and assess inherited risk.
- Names each upstream model and provider explicitly
- Specifies version pinning and update policy
- Identifies which product features use which model
- Refuses to name upstream models
- Cannot specify version pinning
- Routes silently between models with no customer visibility
8. What remediation options are available to customers when bias is detected post-deployment (e.g. prompt-layer mitigations, fine-tuning, model swap, configuration controls)?
Why it matters. Buyers need to know what concrete controls they have to reduce harm while the underlying model is updated.
- Lists concrete remediation levers available to customers
- Distinguishes vendor-side fixes from customer-configurable controls
- References published guidance for each lever
- Only remediation is to wait for the next model version
- No customer-configurable controls
- Cannot give timelines
What the audit changed
A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 154 changes. Three examples:
Wrong or outdated citation
Draft: Sector-specific regulators impose distinct fairness standards (e.g. four-fifths rule in employment, ECOA in credit).
Now: Sector-specific laws and guidelines set distinct fairness tests (e.g. the four-fifths rule of thumb for adverse impact in the Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4(D); the Equal Credit Opportunity Act in credit).
The four-fifths rule is a rule of thumb in the Uniform Guidelines, not a regulator's standard, and ECOA is a statute. Source: https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607
Wrong or outdated citation
Draft: Distinguishes high-risk from low-risk use under the EU AI Act
Now: Distinguishes high-risk use under the EU AI Act from other use
The EU AI Act has no 'low-risk' category; it sets prohibited practices, high-risk systems and transparency obligations. Source: Regulation (EU) 2024/1689, https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Wrong or outdated citation
Draft: Fairness metrics encode incompatible normative trade-offs and cannot all be satisfied simultaneously. A vendor that can name its chosen metrics and justify the trade-off has thought through the problem; one that cannot has not.
Now: Common fairness metrics cannot all be satisfied at once when base rates differ between groups and prediction is imperfect (Kleinberg, Mullainathan and Raghavan, 2016; Chouldechova, 2017). The vendor's choice of metric is a trade-off the buyer should see.
The incompatibility holds only when base rates differ and the predictor is imperfect (Kleinberg et al., 'Inherent Trade-Offs in the Fair Determination of Risk Scores', 2016; Chouldechova, 'Fair prediction with disparate impact', 2017). The last sentence was a judgment stated as fact.