CIOPages
All RFP question modules

AI Governance

AI cost predictability & FinOps questions to ask a software vendor

Questions on controlling AI spend: token and compute metering, cost forecasting, budgets and caps, attribution to teams and projects, and optimization options.

95
questions
31
RFI
32
RFP
32
deep-dive

8 questions from the RFI stage, free

These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.

1. Provide a comprehensive list of every billable unit your platform meters, and describe the metering mechanism for each. Include units such as input/output/cached tokens, embeddings, image/audio units, fine-tuning, tool calls, retrieval operations, and other agent-step actions. For each unit, specify whether it appears on the invoice or is visible only in usage reports.

Why it matters. AI products bill in heterogeneous units that don't map to traditional license SKUs. Buyers need a complete inventory of billable units to forecast spend. Agentic workflows can multiply costs through chained tool calls and retrievals, so visibility into the cost of each step is critical, not just the top-level request, to avoid significant hidden costs.

Good answer
  • Provides a complete, well-defined list of all billable units.
  • Clearly distinguishes metering for different token types (input, output, cached).
  • Details per-step accounting for agentic workflows (tool calls, retrieval).
Red flags
  • Vague answer referring only to 'tokens' or 'API calls'.
  • Fails to distinguish costs for tool calls or agent steps from primary generation.
  • Only meters the top-level user request in an agentic workflow.

2. What tooling do you provide to forecast monthly or quarterly spend based on a customer's historical usage patterns?

Why it matters. Budget owners need predictive tooling to set guardrails and justify spend internally. Without forecasting, AI budgets become reactive and unpredictable, hindering adoption.

Good answer
  • Provides a forecasting dashboard with trend extrapolation.
  • Allows adjustments for seasonality and expected growth.
  • Displays confidence intervals on cost projections.
Red flags
  • No forecasting capability is offered.
  • Forecast is simply the previous month's spend with no trending.
  • Forecasting is only available as part of a paid professional services engagement.

3. What automated cost-optimization recommendations does your platform provide (e.g. 'this workload could use a cheaper model', 'prompt caching would save X%', 'batch API eligible')?

Why it matters. Vendor-side recommendations capture optimizations the customer might miss. The answer shows whether the vendor recommends options that lower the customer's spend, such as a cheaper model.

Good answer
  • Provides concrete recommendations with estimated monthly savings.
  • Covers a range of optimizations: model selection, caching, batching, context size.
  • Recommendations are refreshed regularly as new models or features emerge.
Red flags
  • No automated optimization recommendations are provided.
  • Recommendations are only available via a paid professional services contract.
  • Recommendations are generic and never suggest using a cheaper model.

4. How does the platform support tagging or labeling of requests for cost attribution (e.g. project, team, business unit, cost center, environment)?

Why it matters. Without robust tagging, finance teams cannot allocate AI spend back to the business units that consumed it. Tagging is the foundation of internal chargeback and showback for any shared service.

Good answer
  • Supports arbitrary key-value tags on requests via API headers or parameters.
  • Tags propagate consistently to invoices and all usage exports.
  • Provides policies to enforce the presence of required tags on requests.
Red flags
  • No tagging support is offered.
  • Only a single, fixed 'user' field is available for attribution.
  • Tags are visible in the dashboard but are not included in data exports.

5. At what granularity is usage data available (per request, per user, per session, per API key, per project)?

Why it matters. The FinOps Foundation's principles say FinOps data should be accessible, timely and accurate, and that cost data should be processed and shared as soon as it becomes available. Daily-only aggregates allow significant cost overruns before anyone notices.

Good answer
  • Per-request granularity is available via an API.
  • Data can be grouped by multiple dimensions (key, project, user, model).
Red flags
  • Only daily or monthly aggregates are available.
  • No per-request detail is available for inspection.
  • Granular data is gated behind a premium support or pricing tier.

6. Do you provide budget thresholds with automated alerts (e.g. notification at 50%, 80%, 100% of monthly budget), and can these be configured per project, team, or API key?

Why it matters. Automated budget alerts tell a budget owner that spend is approaching a limit before the invoice arrives. Applying budgets to specific scopes (like projects or teams) allows each group to manage its own cost behavior.

Good answer
  • Supports multi-threshold alerts (e.g., 50%, 80%, 100% of budget).
  • Delivers alerts to multiple channels (email, webhook, Slack).
  • Budgets can be set at granular scopes (project, team, API key).
Red flags
  • Only a single, account-wide budget is supported.
  • Alerts are not sent until after an invoice is generated.
  • No webhook integration for connecting alerts to ITSM or chat tools.

7. Do you offer a batch processing tier with reduced unit pricing for non-latency-sensitive workloads, and how does a customer route eligible traffic to it?

Why it matters. At the time of writing (October 2026), OpenAI, Anthropic and Google list batch processing at 50% of standard pricing, and AWS lists Amazon Bedrock batch inference at 50% of on-demand pricing for select models. Workloads such as summarization or classification that do not need an immediate response can use a batch tier.

Good answer
  • A batch tier is available with a clearly documented discount.
  • A dedicated API endpoint or parameter is used to submit batch jobs.
  • A documented completion window for batch jobs is provided, and the vendor states whether it is a contractual SLA.
Red flags
  • No batch tier is offered.
  • The batch tier is missing key models available in the synchronous API.
  • No completion window or target is documented for batch jobs.

8. Can the platform produce showback or chargeback reports broken down by tag, project, team, or cost center, on a configurable schedule?

Why it matters. The FinOps Framework covers showback and chargeback in its Invoicing & Chargeback capability (Manage the FinOps Practice domain). Automation of these reports directly determines whether finance teams can sustain monthly allocation cycles without significant manual effort.

Good answer
  • Provides a scheduled reporting feature that groups costs by specified dimensions.
  • Reports can be delivered via email or API.
  • Reports include both consumption units and dollar amounts.
Red flags
  • Reports require manual data extraction and manipulation every month.
  • No breakdown is possible beyond the top-level account.
  • Creating custom reports requires a professional services engagement.

The full set: 95 questions in a scored Excel workbook

  • RFI, RFP and deep-dive sheets, with an evaluator guide on every question
  • A 0–5 score column, suggested weights and a scorecard that totals by depth and section
  • An RFP cover template in Word
  • An audit log of all 115 changes made to the draft

Consultancy License $399, for use with any number of clients.

What the module covers

  • Unit accounting (tokens, embeddings, agent steps) (33)
  • Cost prediction & budgeting tooling (18)
  • Cost-optimization recommendations (18)
  • Per-team, per-project, per-workload attribution (16)

What the audit changed

A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 115 changes. Three examples:

Wrong or outdated citation

Draft: FinOps practice requires near-real-time visibility into spend to detect runaway workloads.

Now: The FinOps Foundation's principles say FinOps data should be accessible, timely and accurate, and that cost data should be processed and shared as soon as it becomes available.

The FinOps principles do not require 'near-real-time' visibility. Source: FinOps Foundation, FinOps Principles: 'FinOps data should be accessible, timely, and accurate' (https://www.finops.org/framework/principles/). The principle's guidance reads 'Process and share cost data as soon as it becomes available.'

Wrong or outdated citation

Draft: The current version is FOCUS 1.4, ratified June 4, 2026.

Now: At the time of writing (October 2026), the current version is FOCUS 1.4, ratified June 4, 2026.

Dated, because FOCUS 1.5 is in progress. Source: https://www.finops.org/insights/introducing-focus-1-4/ (ratified June 4, 2026; work on 1.5 underway); https://focus.finops.org/ lists 1.4 as current, read 2026-10-05.

Wrong or outdated citation

Draft: A clear SLA on batch completion time is provided.

Now: A documented completion window for batch jobs is provided, and the vendor states whether it is a contractual SLA.

Major providers publish a completion window or target (OpenAI 24h window, Google 24h target), not always an SLA. Sources as in the why_it_matters edit.

Questions about this module

How many ai cost predictability & finops questions are there?

95: 31 for the RFI stage, 32 for the RFP and 32 deep-dive questions for the finalists.

What comes with each question?

Why it matters, what a good answer looks like, the red flags, follow-up questions, the response format, whether most buyers treat it as mandatory, and a suggested weight for scoring.

Were the questions checked?

A language model drafted them and a second model critiqued them. Three audit passes followed (2026-10-05) and made 115 changes, each listed in the workbook with the old and new text. No named subject-matter expert wrote them.

Related