3 questions from the package
From the RFI round. The first shows part of the guide each question carries; the workbook adds follow-ups, how to verify the answer, a priority and a weight.
1. Provide the list of models callable through your platform today, with each model's provider, size tier (frontier, mid-size or small) and whether its weights are openly available.
Why it matters. If the catalog lacks small or open-weight options, the buyer cannot route bounded tasks to lower-cost models or meet self-hosting requirements without adding a second platform.
Good answer
- The list names individual models with versions, not just model families.
- It includes models from providers other than the vendor itself.
- Each entry marks open-weight status and names the license that applies.
Red flags
- A model count or a marketing page is offered in place of an itemized list.
- Open-weight models are listed but can only be reached through a separate product or contract.
- The list carries no date or version.
2. Does your platform store our own evaluation sets, made up of prompts with reference answers, as versioned datasets inside the product?
Why it matters. Without a stored, versioned evaluation set inside the product, the buyer has to keep test prompts in spreadsheets or a separate tool. Scores from different months then cannot be tied to the same set of items.
3. List the document file formats your retrieval pipeline can ingest and parse with no preprocessing on our side.
Why it matters. If the formats the buyer holds are not parsed natively, the buyer must build and maintain conversion steps. Content lost during conversion never reaches retrieval.
Capability areas
Model Catalog & Task Fit (10)
Covers the breadth of models reachable on the platform (frontier, mid-size, small and open-weight), modality coverage, context-window limits, how the buyer matches a model to a task, and how an administrator controls which models each project or team may call. The vendor's published benchmark claims are out; the buyer's own evaluation is the EVL area.
Evaluation & Regression Harness (13)
Covers product tooling the buyer uses to build, version and run its own evaluation sets, configure graders (reference, rubric, model-as-judge), compare models side by side, and gate releases on regression results. How the vendor evaluates its own models is out.
Retrieval & Grounding (13)
Covers document ingestion, parsing and chunking, embedding and vector or hybrid search, permission-aware retrieval, source citation in responses, and index refresh. Generic prebuilt connectors to enterprise systems are out; they belong to the integration module.
Fine-Tuning & Model Customization (10)
Covers managed fine-tuning methods, continued pre-training, distillation into smaller models, training-data preparation, hosting of tuned models, and export of tuned weights or adapters. Provenance of the training data behind the vendor's base models is out.
Tool Calling & Agent Runtime (13)
Covers function calling and structured-output reliability, open agent protocol support, a managed runtime with memory and state, multi-step and multi-agent orchestration, and step-level tracing of agent runs. Action scoping, tool permissions, sandboxing and kill switches are out; they belong to the ai-agentic-autonomy module.
Model Abstraction & Portability (11)
Covers a single interface across model families and providers, routing and fallback between models, translation of prompt and tool schemas across providers, and the measurable rework needed to swap a model. Contract exit terms and bulk data export are out; they belong to the migration-exit module.
Model Versioning & Lifecycle (9)
Covers version pinning per deployment, model alias behavior, deprecation notice and retirement mechanics for hosted models, detection of behavior changes behind an unchanged model name, and migration support between versions. The vendor's organization-wide model risk management program is out.
Guardrails & Content Controls (12)
Covers configurable input and output filters, prompt-injection and jailbreak detection settings, in-line PII detection and redaction, topic and policy rules per application, and how guardrail decisions surface to the calling application. The vendor's red-teaming program and abuse-reporting posture are out; they belong to the ai-safety module.
Inference Deployment & Data Path (11)
Covers where inference runs for each model (multi-tenant endpoint, private network endpoint, customer cloud account, self-hosted open weights, disconnected environment), the path prompts and outputs take, whether inference stays in a chosen region or can be routed across regions, and per-deployment controls for prompt retention and abuse-monitoring storage. Generic hosting architecture, the DPA and sub-processor lists are out.
Serving Throughput & Performance (10)
Covers rate limits and quota increases, provisioned or reserved throughput, batch inference, prompt caching, streaming, latency behavior under load, and serving options for self-hosted models such as quantization and autoscaling. Pricing and spend controls are out; they belong to the ai-cost-finops module.
Observability & Audit Logging (10)
Covers capture of prompts, responses, retrieved context and tool calls, request-level tracing, token usage attribution, production quality monitoring and drift alerts, and export of logs to the buyer's monitoring and security tools. Human-in-the-loop decision trails for automated business decisions are out.
Prompt & Release Management (9)
Covers prompt templates and versioning, environment separation from development to production, promotion and rollback of prompt and model configuration changes, and moving work from playground to deployed endpoint. Source-code CI/CD outside the platform is out.
Questions about this package
How many Generative AI & LLM Platforms RFP questions are there?
131 solution questions in 12 capability areas: 27 for the RFI, 68 for the RFP and 36 deep-dive questions for the finalists. The workbook adds 100 due-diligence questions on security, integration, implementation and exit.
What comes with each question?
Why it matters, good-answer signals, red flags, follow-up questions, how to verify the answer (a demo step, a test or a document), and a suggested priority and weight for scoring.
Can I edit the questions?
Yes. The workbook is an ordinary Excel file. Change, add or remove questions, and change the weights; the scorecard recalculates.
Which license do I need?
The Enterprise License covers any number of evaluations inside one organization. The Consultancy License covers use with any number of clients. Neither allows reselling or republishing the questions.