CIOPages
All RFP question modules

Security & Compliance

Data protection & privacy questions to ask a software vendor

Questions on how the vendor handles your data: personal data, consent, residency, retention, deletion, the data processing agreement and sub-processors. Written against GDPR and CCPA obligations, with signals specific enough to check.

111
questions
35
RFI
43
RFP
33
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Describe how your platform identifies, classifies, and minimizes personal data (including special-category data) that customers submit as prompts, files, or context to your AI features.

Why it matters. Data minimization is a GDPR principle (Article 5(1)(c)) and a baseline expectation for any AI tool that ingests customer content. Personal data the vendor does not classify cannot be minimized, and every unminimized copy adds to the buyer's breach exposure and DSAR workload.

Good answer
  • Names specific classifiers or DLP integrations applied to inbound content
  • Distinguishes handling of special-category data (health, biometric, etc.) from ordinary PII
  • Cites a published data classification policy or schema
Red flags
  • Claims 'we don't see customer data' without explaining how that is enforced
  • No documented classification scheme
  • Treats all customer input as a single undifferentiated category

2. List the regions in which customer data can be stored at rest, and indicate whether residency commitments cover backups, replicas, and disaster-recovery copies in addition to primary storage.

Why it matters. A residency commitment may cover the application layer while backups, snapshots, or DR replicas are stored elsewhere. Buyers in regulated industries need residency guarantees that extend to every copy of the data, not just the primary store.

Good answer
  • Names specific cloud regions per data class
  • Explicitly covers backups, replicas, and DR within the same residency boundary
  • Cites contractual residency language
Red flags
  • Residency stated only for primary storage
  • Backups described as 'in a secure location' without region disclosure
  • DR region undisclosed or outside the stated residency boundary

3. Provide your standard Data Processing Addendum and indicate which clauses are negotiable for enterprise customers.

Why it matters. Buyers need to see the paper up front and understand the negotiation surface before investing legal review cycles.

Good answer
  • Standard DPA published or readily available
  • Identifies typically-negotiated clauses (liability caps, audit rights, sub-processor approval)
  • Includes EU SCCs and UK addendum by reference or attachment
Red flags
  • No standard DPA exists
  • DPA only available after contract signature
  • No SCCs incorporated

4. Describe the default retention periods for each class of customer data (e.g. content, prompts/completions, embeddings, logs, telemetry, backups) and whether customers can configure these.

Why it matters. Storage limitation is a GDPR principle (Article 5(1)(e)), and shorter retention reduces what a breach can expose. Undisclosed retention of prompts or embeddings leaves the buyer unable to state its own retention periods.

Good answer
  • Itemized retention table per data class
  • Customer-configurable retention with explicit floors and ceilings
  • Embeddings and derivative artifacts have stated retention
Red flags
  • Single retention period stated for 'all data'
  • Backups described as indefinite
  • Embeddings or fine-tunes outlive source content with no disclosure

5. Confirm whether customer data (including prompts, completions, embeddings, fine-tuning data, and telemetry) is ever used to train, fine-tune, or evaluate your or any third party's models, and describe the consent mechanism that governs this.

Why it matters. Training on customer data without the customer's consent can expose the buyer's data to other model users and breach the buyer's own purpose limits. Buyers need an unambiguous yes/no plus the mechanics of consent — opt-in vs opt-out, default state, and scope across all data types including embeddings and telemetry.

Good answer
  • States training-use posture clearly for each data type (prompts, completions, embeddings, fine-tunes, logs)
  • Default is no-training without explicit opt-in
  • Cites contractual language in the DPA that binds this commitment
Red flags
  • Vague 'we may use aggregated or anonymized data' language
  • Opt-out required rather than opt-in
  • Silent on telemetry, embeddings, or evaluation data

6. Describe the residency posture for metadata, logs, telemetry, and support-access data (e.g. session recordings, debug bundles).

Why it matters. Metadata and operational telemetry can be stored or processed in other regions even when customer content is kept in-region.

Good answer
  • Itemizes each operational data class and its residency
  • Logs and telemetry stay within the same boundary as content
  • Customer support data flows are documented
Red flags
  • Conflates 'customer data' residency with all data residency
  • Logs admitted to flow to a central region with no controls
  • Cannot describe support-engineer access paths

7. State whether you act as a data processor, controller, or joint controller for each category of personal data processed, and describe any scenarios in which your role changes.

Why it matters. An AI vendor may act as a controller for telemetry, analytics, or model-improvement data while acting as a processor for customer content. The buyer must know which role applies to each category, because GDPR Article 28 obligations apply only where the vendor is a processor.

Good answer
  • Provides a role matrix by data category
  • Identifies any controller activities clearly
  • Explains the lawful basis used when acting as controller
Red flags
  • Claims processor-only without acknowledging telemetry or marketing data
  • Cannot articulate lawful basis for controller activities
  • Conflates joint controllership with processor relationship

8. Describe the process and SLA for data deletion upon contract termination, including how derivative data (e.g., embeddings, fine-tuned models, caches, logs, backups) is handled.

Why it matters. Deleting source content while retaining embeddings or fine-tuned weights preserves information from the data in a different form. Buyers need explicit assurance that all derivatives are deleted.

Good answer
  • Lists every data location wiped on deletion (primary, replicas, backups, embeddings, fine-tunes, caches, logs)
  • States the SLA for completion of deletion across all locations
  • Documents handling of fine-tuned model weights derived from customer data
Red flags
  • Deletion covers only the primary store
  • Backups retained for the full backup cycle with no acceleration option
  • Fine-tuned weights retained after source deletion

The full set: 111 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 193 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • PII handling, minimization & classification (30)
  • Data residency at the data layer (17)
  • DPA, controller/processor terms & consent (40)
  • Retention policy & verified deletion (24)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 193 changes. Three examples:

Wrong or outdated citation

Draft: PII minimization is a core GDPR principle

Now: Data minimization is a GDPR principle (Article 5(1)(c))

GDPR Art. 5(1)(c) names data minimization; it applies to personal data, not a PII category. Regulation (EU) 2016/679, https://eur-lex.europa.eu/eli/reg/2016/679/oj

Wrong or outdated citation

Draft: Explain how children's data and other special-category data are handled

Now: Explain how children's data is handled

Children's data is not a GDPR Article 9 special category, so "other special-category data" is wrong; Art. 9(1) lists racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data, biometric data for unique identification, health, sex life and sexual orientation. Special-category handling moves to rfi.004b. Regulation (EU) 2016/679, https://eur-lex.europa.eu/eli/reg/2016/679/oj

Wrong or outdated citation

Draft: States whether the service is intended for under-16 users

Now: States whether the service is intended for users under 13 (COPPA) or under the applicable GDPR Article 8 age (16 by default)

COPPA applies under 13; GDPR Art. 8(1) sets 16 with member-state option down to 13.

Questions about this module

How many data protection & privacy questions are there?

111: 35 for the RFI stage, 43 for the RFP and 33 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 193 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related