CIOPages
All RFP packages

RFP Package · AI & Automation

Generative AI & LLM Platforms RFP questions and template

131 questions, 10 demo scenarios and a five-vendor scorecard for choosing Generative AI & LLM Platforms software, in one Excel workbook.

What this package is for

Use it to run a Generative AI & LLM Platforms software selection, from the first long list to the final scorecard.

What the category covers. Platforms for building on large language models: the model catalog, evaluation, retrieval, fine-tuning, agents, portability across providers, model lifecycle, guardrails, deployment boundary, serving and observability. Bought by CIOs, CTOs, platform engineering and AI leaders choosing where their models, data and agents run.

A selection usually runs in three rounds. The package has questions for each:

  • RFI, to the long list. 27 questions screen out products that lack something you need.
  • RFP, to the shortlist. 68 questions ask how each product does the work.
  • Deep dive, to the finalists. 36 questions ask for proof on your own data.

10 demo scenarios tell each vendor what to load and what to show, so every product does the same work in front of you. 100 due-diligence questions cover security, integration, implementation and exit. The scorecard weights the answers and ranks up to five vendors.

Each question comes with why it matters, what a good answer looks like and the red flags, so the people scoring the replies know what to look for.

3 questions from the package

From the RFI round. The first shows part of the guide each question carries; the workbook adds follow-ups, how to verify the answer, a priority and a weight.

1. Provide the list of models callable through your platform today, with each model's provider, size tier (frontier, mid-size or small) and whether its weights are openly available.

Why it matters. If the catalog lacks small or open-weight options, the buyer cannot route bounded tasks to lower-cost models or meet self-hosting requirements without adding a second platform.

Good answer
  • The list names individual models with versions, not just model families.
  • It includes models from providers other than the vendor itself.
  • Each entry marks open-weight status and names the license that applies.
Red flags
  • A model count or a marketing page is offered in place of an itemized list.
  • Open-weight models are listed but can only be reached through a separate product or contract.
  • The list carries no date or version.

2. Does your platform store our own evaluation sets, made up of prompts with reference answers, as versioned datasets inside the product?

Why it matters. Without a stored, versioned evaluation set inside the product, the buyer has to keep test prompts in spreadsheets or a separate tool. Scores from different months then cannot be tied to the same set of items.

3. List the document file formats your retrieval pipeline can ingest and parse with no preprocessing on our side.

Why it matters. If the formats the buyer holds are not parsed natively, the buyer must build and maintain conversion steps. Content lost during conversion never reaches retrieval.

Capability areas

Model Catalog & Task Fit (10)

Covers the breadth of models reachable on the platform (frontier, mid-size, small and open-weight), modality coverage, context-window limits, how the buyer matches a model to a task, and how an administrator controls which models each project or team may call. The vendor's published benchmark claims are out; the buyer's own evaluation is the EVL area.

Evaluation & Regression Harness (13)

Covers product tooling the buyer uses to build, version and run its own evaluation sets, configure graders (reference, rubric, model-as-judge), compare models side by side, and gate releases on regression results. How the vendor evaluates its own models is out.

Retrieval & Grounding (13)

Covers document ingestion, parsing and chunking, embedding and vector or hybrid search, permission-aware retrieval, source citation in responses, and index refresh. Generic prebuilt connectors to enterprise systems are out; they belong to the integration module.

Fine-Tuning & Model Customization (10)

Covers managed fine-tuning methods, continued pre-training, distillation into smaller models, training-data preparation, hosting of tuned models, and export of tuned weights or adapters. Provenance of the training data behind the vendor's base models is out.

Tool Calling & Agent Runtime (13)

Covers function calling and structured-output reliability, open agent protocol support, a managed runtime with memory and state, multi-step and multi-agent orchestration, and step-level tracing of agent runs. Action scoping, tool permissions, sandboxing and kill switches are out; they belong to the ai-agentic-autonomy module.

Model Abstraction & Portability (11)

Covers a single interface across model families and providers, routing and fallback between models, translation of prompt and tool schemas across providers, and the measurable rework needed to swap a model. Contract exit terms and bulk data export are out; they belong to the migration-exit module.

Model Versioning & Lifecycle (9)

Covers version pinning per deployment, model alias behavior, deprecation notice and retirement mechanics for hosted models, detection of behavior changes behind an unchanged model name, and migration support between versions. The vendor's organization-wide model risk management program is out.

Guardrails & Content Controls (12)

Covers configurable input and output filters, prompt-injection and jailbreak detection settings, in-line PII detection and redaction, topic and policy rules per application, and how guardrail decisions surface to the calling application. The vendor's red-teaming program and abuse-reporting posture are out; they belong to the ai-safety module.

Inference Deployment & Data Path (11)

Covers where inference runs for each model (multi-tenant endpoint, private network endpoint, customer cloud account, self-hosted open weights, disconnected environment), the path prompts and outputs take, whether inference stays in a chosen region or can be routed across regions, and per-deployment controls for prompt retention and abuse-monitoring storage. Generic hosting architecture, the DPA and sub-processor lists are out.

Serving Throughput & Performance (10)

Covers rate limits and quota increases, provisioned or reserved throughput, batch inference, prompt caching, streaming, latency behavior under load, and serving options for self-hosted models such as quantization and autoscaling. Pricing and spend controls are out; they belong to the ai-cost-finops module.

Observability & Audit Logging (10)

Covers capture of prompts, responses, retrieved context and tool calls, request-level tracing, token usage attribution, production quality monitoring and drift alerts, and export of logs to the buyer's monitoring and security tools. Human-in-the-loop decision trails for automated business decisions are out.

Prompt & Release Management (9)

Covers prompt templates and versioning, environment separation from development to production, promotion and rollback of prompt and model configuration changes, and moving work from playground to deployed endpoint. Source-code CI/CD outside the platform is out.

Demo scenarios

Each scenario lists the data to load before the demo, then the steps to show, and the questions it scores.

  1. Compare three models on our evaluation set
  2. Swap the model behind a working application
  3. Two users ask the same question
  4. Agent calls our tools and waits for approval
  5. Injected instruction and personal data meet the guardrails
  6. Fine-tune a small model on our examples
  7. Burst traffic past the quota
  8. Private endpoint with no public egress
  9. Respond to a model deprecation notice
  10. Production quality drops and an alert fires

Due diligence

The workbook carries the screening questions from these modules. Each module is also sold on its own.

Questions about this package

How many Generative AI & LLM Platforms RFP questions are there?

131 solution questions in 12 capability areas: 27 for the RFI, 68 for the RFP and 36 deep-dive questions for the finalists. The workbook adds 100 due-diligence questions on security, integration, implementation and exit.

What comes with each question?

Why it matters, good-answer signals, red flags, follow-up questions, how to verify the answer (a demo step, a test or a document), and a suggested priority and weight for scoring.

Can I edit the questions?

Yes. The workbook is an ordinary Excel file. Change, add or remove questions, and change the weights; the scorecard recalculates.

Which license do I need?

The Enterprise License covers any number of evaluations inside one organization. The Consultancy License covers use with any number of clients. Neither allows reselling or republishing the questions.

Before you shortlist

The buyer guide compares the products in this category and what decides between them.

Buyer Guide
Generative AI & LLM Platforms

For the business side of the same change: