8 questions from the RFI stage, free
These come from the module as sold. The workbook adds follow-ups, a response format, a weight and a score column to each.
1. Provide a comprehensive list of every billable unit your platform meters, and describe the metering mechanism for each. Include units such as input/output/cached tokens, embeddings, image/audio units, fine-tuning, tool calls, retrieval operations, and other agent-step actions. For each unit, specify whether it appears on the invoice or is visible only in usage reports.
Why it matters. AI products bill in heterogeneous units that don't map to traditional license SKUs. Buyers need a complete inventory of billable units to forecast spend. Agentic workflows can multiply costs through chained tool calls and retrievals, so visibility into the cost of each step is critical, not just the top-level request, to avoid significant hidden costs.
- Provides a complete, well-defined list of all billable units.
- Clearly distinguishes metering for different token types (input, output, cached).
- Details per-step accounting for agentic workflows (tool calls, retrieval).
- Vague answer referring only to 'tokens' or 'API calls'.
- Fails to distinguish costs for tool calls or agent steps from primary generation.
- Only meters the top-level user request in an agentic workflow.
2. What tooling do you provide to forecast monthly or quarterly spend based on a customer's historical usage patterns?
Why it matters. Budget owners need predictive tooling to set guardrails and justify spend internally. Without forecasting, AI budgets become reactive and unpredictable, hindering adoption.
- Provides a forecasting dashboard with trend extrapolation.
- Allows adjustments for seasonality and expected growth.
- Displays confidence intervals on cost projections.
- No forecasting capability is offered.
- Forecast is simply the previous month's spend with no trending.
- Forecasting is only available as part of a paid professional services engagement.
3. What automated cost-optimization recommendations does your platform provide (e.g. 'this workload could use a cheaper model', 'prompt caching would save X%', 'batch API eligible')?
Why it matters. Vendor-side recommendations capture optimizations the customer might miss. The answer shows whether the vendor recommends options that lower the customer's spend, such as a cheaper model.
- Provides concrete recommendations with estimated monthly savings.
- Covers a range of optimizations: model selection, caching, batching, context size.
- Recommendations are refreshed regularly as new models or features emerge.
- No automated optimization recommendations are provided.
- Recommendations are only available via a paid professional services contract.
- Recommendations are generic and never suggest using a cheaper model.
4. How does the platform support tagging or labeling of requests for cost attribution (e.g. project, team, business unit, cost center, environment)?
Why it matters. Without robust tagging, finance teams cannot allocate AI spend back to the business units that consumed it. Tagging is the foundation of internal chargeback and showback for any shared service.
- Supports arbitrary key-value tags on requests via API headers or parameters.
- Tags propagate consistently to invoices and all usage exports.
- Provides policies to enforce the presence of required tags on requests.
- No tagging support is offered.
- Only a single, fixed 'user' field is available for attribution.
- Tags are visible in the dashboard but are not included in data exports.
5. At what granularity is usage data available (per request, per user, per session, per API key, per project)?
Why it matters. The FinOps Foundation's principles say FinOps data should be accessible, timely and accurate, and that cost data should be processed and shared as soon as it becomes available. Daily-only aggregates allow significant cost overruns before anyone notices.
- Per-request granularity is available via an API.
- Data can be grouped by multiple dimensions (key, project, user, model).
- Only daily or monthly aggregates are available.
- No per-request detail is available for inspection.
- Granular data is gated behind a premium support or pricing tier.
6. Do you provide budget thresholds with automated alerts (e.g. notification at 50%, 80%, 100% of monthly budget), and can these be configured per project, team, or API key?
Why it matters. Automated budget alerts tell a budget owner that spend is approaching a limit before the invoice arrives. Applying budgets to specific scopes (like projects or teams) allows each group to manage its own cost behavior.
- Supports multi-threshold alerts (e.g., 50%, 80%, 100% of budget).
- Delivers alerts to multiple channels (email, webhook, Slack).
- Budgets can be set at granular scopes (project, team, API key).
- Only a single, account-wide budget is supported.
- Alerts are not sent until after an invoice is generated.
- No webhook integration for connecting alerts to ITSM or chat tools.
7. Do you offer a batch processing tier with reduced unit pricing for non-latency-sensitive workloads, and how does a customer route eligible traffic to it?
Why it matters. At the time of writing (October 2026), OpenAI, Anthropic and Google list batch processing at 50% of standard pricing, and AWS lists Amazon Bedrock batch inference at 50% of on-demand pricing for select models. Workloads such as summarization or classification that do not need an immediate response can use a batch tier.
- A batch tier is available with a clearly documented discount.
- A dedicated API endpoint or parameter is used to submit batch jobs.
- A documented completion window for batch jobs is provided, and the vendor states whether it is a contractual SLA.
- No batch tier is offered.
- The batch tier is missing key models available in the synchronous API.
- No completion window or target is documented for batch jobs.
8. Can the platform produce showback or chargeback reports broken down by tag, project, team, or cost center, on a configurable schedule?
Why it matters. The FinOps Framework covers showback and chargeback in its Invoicing & Chargeback capability (Manage the FinOps Practice domain). Automation of these reports directly determines whether finance teams can sustain monthly allocation cycles without significant manual effort.
- Provides a scheduled reporting feature that groups costs by specified dimensions.
- Reports can be delivered via email or API.
- Reports include both consumption units and dollar amounts.
- Reports require manual data extraction and manipulation every month.
- No breakdown is possible beyond the top-level account.
- Creating custom reports requires a professional services engagement.
What the audit changed
A language model drafted these questions and a second model critiqued them. Three audit passes followed and made 115 changes. Three examples:
Wrong or outdated citation
Draft: FinOps practice requires near-real-time visibility into spend to detect runaway workloads.
Now: The FinOps Foundation's principles say FinOps data should be accessible, timely and accurate, and that cost data should be processed and shared as soon as it becomes available.
The FinOps principles do not require 'near-real-time' visibility. Source: FinOps Foundation, FinOps Principles: 'FinOps data should be accessible, timely, and accurate' (https://www.finops.org/framework/principles/). The principle's guidance reads 'Process and share cost data as soon as it becomes available.'
Wrong or outdated citation
Draft: The current version is FOCUS 1.4, ratified June 4, 2026.
Now: At the time of writing (October 2026), the current version is FOCUS 1.4, ratified June 4, 2026.
Dated, because FOCUS 1.5 is in progress. Source: https://www.finops.org/insights/introducing-focus-1-4/ (ratified June 4, 2026; work on 1.5 underway); https://focus.finops.org/ lists 1.4 as current, read 2026-10-05.
Wrong or outdated citation
Draft: A clear SLA on batch completion time is provided.
Now: A documented completion window for batch jobs is provided, and the vendor states whether it is a contractual SLA.
Major providers publish a completion window or target (OpenAI 24h window, Google 24h target), not always an SLA. Sources as in the why_it_matters edit.