3 questions from the package
From the RFI round. The first shows part of the guide each question carries; the workbook adds follow-ups, how to verify the answer, a priority and a weight.
1. Does your product provide a query profile that shows, for each operator in an executed query, the rows processed, bytes read and time spent?
Why it matters. Without per-operator statistics, the buyer's engineers cannot tell whether a slow query is caused by a scan, a join, a spill or network transfer. Tuning then becomes trial and error.
Good answer
- Profile is available for completed queries as well as queries still running
- Each operator shows rows in, rows out, bytes read and elapsed or CPU time
- Spill to local or remote storage is shown at the operator where it happened
Red flags
- Only a pre-execution estimated plan is available, with no runtime statistics
- Operator-level detail requires opening a support ticket
- Profile shows only total query time and total bytes scanned
2. Can separate compute resources serving BI, ETL and ad-hoc workloads read and write the same stored tables at the same time without sharing CPU, memory or local cache?
Why it matters. If workloads share compute, a large ETL job or an ad-hoc scan slows dashboards for every business user. If isolation requires copying data to a second cluster, the buyer pays for duplicate storage and has to manage stale copies.
3. List the file formats and compression codecs your product can bulk load from object storage without converting the files first.
Why it matters. If a format the buyer's source systems produce needs conversion before load, the buyer has to build and run an extra conversion step. That step adds cost, latency and a point of failure to every pipeline.
Capability areas
Query Performance & Optimization (13)
Covers optimizer behavior, result and data caching, materialized views, clustering and partition pruning, and query plan and profile diagnostics. Concurrency and workload isolation are covered in CON.
Concurrency, Elastic Scaling & Workload Isolation (10)
Covers how the platform serves mixed BI, ETL and ad-hoc workloads at the same time: compute isolation, automatic scale-out, queuing behavior, and workload priorities. Cost controls on that compute are covered in CST.
Batch, Streaming & CDC Ingestion (12)
Covers bulk loading from object storage, streaming ingestion from event platforms, change data capture from operational databases, data freshness, error handling, and schema drift during load. Transformation after landing is covered in TRN.
In-Warehouse Transformation & Orchestration (9)
Covers SQL and procedural transformation inside the warehouse, incremental processing, scheduled and dependency-driven tasks, and integration with external transformation frameworks and orchestrators. It also covers deploying schema and pipeline changes across environments.
Storage, Table Formats & Data Types (12)
Covers compute/storage separation, native handling of semi-structured data (JSON, Avro, Parquet), support for open table formats read and written by other engines, time travel, zero-copy cloning, and data recovery from user error. Platform-level disaster recovery is out of scope.
Fine-Grained Access Policies & Masking (11)
Covers row-level and column-level security, dynamic data masking, tag- or attribute-based policies, and how policies are enforced across SQL, Python, sharing and external engines. Generic identity integration and platform security are out of scope.
Catalog, Lineage & Data Auditing (10)
Covers the built-in data catalog, column-level lineage from ingestion through consumption, impact analysis, data quality checks, and queryable history of which users and roles read which tables and columns. General audit logging and third-party attestations are out of scope.
Data Sharing & Collaboration (10)
Covers sharing governed data without copying across accounts, regions and clouds, sharing with external parties that have no account of their own, controlled-computation environments for joint analysis, consuming third-party datasets through a marketplace or exchange, and revoking access.
In-Platform AI & ML (11)
Covers Python and Spark runtimes running next to the data, feature management, model training and batch scoring inside the platform, vector storage and search, and SQL-callable LLM functions. AI safety, bias and model governance are covered by the cross-cutting AI modules.
Compute Cost Controls & Usage Attribution (10)
Covers auto-suspend and auto-resume, resource monitors and budget limits, attribution of cost to individual queries, users, teams and tags, and recommendations for optimizing spend. Pricing and contract terms are covered by the commercial module.
Legacy Warehouse Migration & Coexistence (11)
Covers moving from on-premises warehouses: workload assessment, SQL dialect and stored-procedure translation, bulk historical data transfer, reconciliation of results, and running old and new platforms in parallel. Leaving this vendor is covered by the migration-exit module.
BI & SQL Client Compatibility (8)
Covers JDBC/ODBC and native driver behavior, SQL dialect coverage, semantic layer or metrics definitions, and how BI tools perform in live-query and extract modes. Generic APIs and SSO are covered by the integration module.
Questions about this package
How many Cloud Data Warehouse RFP questions are there?
127 solution questions in 12 capability areas: 26 for the RFI, 67 for the RFP and 34 deep-dive questions for the finalists. The workbook adds 90 due-diligence questions on security, integration, implementation and exit.
What comes with each question?
Why it matters, good-answer signals, red flags, follow-up questions, how to verify the answer (a demo step, a test or a document), and a suggested priority and weight for scoring.
Can I edit the questions?
Yes. The workbook is an ordinary Excel file. Change, add or remove questions, and change the weights; the scorecard recalculates.
Which license do I need?
The Enterprise License covers any number of evaluations inside one organization. The Consultancy License covers use with any number of clients. Neither allows reselling or republishing the questions.