CIOPages
All Buyer Guides
Data & AnalyticsHigh Complexity

Buyer's Guide: Data Lakehouse Platforms

Evaluate Databricks, Snowflake, Microsoft Fabric, BigQuery, AWS, Dremio, Starburst, Onehouse, and Cloudera around the decision that actually outlives the others — the open table format and catalog you standardize on, not the engine you start with.

16 min read 9 vendors evaluated Typical deal: $200K – $3M+ Updated June 2026
Section 1

Executive Summary

Data Lakehouse Platforms offer a unified architecture for data, combining warehouse-grade SQL, ACID transactions, and governance on open formats in object storage. The choice hinges on table format and catalog strategy, workload fit (SQL analytics, AI/ML), governance reach, and consumption-cost control. Key platforms include Databricks, Snowflake, Microsoft Fabric, BigQuery, and open Iceberg-plus-Trino/Dremio stacks.

The lakehouse promise is one copy of data, open on cheap object storage, queried by whatever engine fits the job — which makes the table format and catalog you standardize on a more durable decision than the engine you start with.

Databricks, Snowflake, Microsoft Fabric, BigQuery, and the open Iceberg-plus-Trino/Dremio stack are converging on the same architecture from different origins: warehouse-grade SQL, ACID transactions, and governance applied directly to data in open formats on object storage. The defining contest is no longer lake versus warehouse but open versus proprietary — Databricks built around Delta Lake and Spark for data engineering and AI, Snowflake and BigQuery extending from managed SQL warehouses, and all of them now embracing Apache Iceberg as the interoperability layer that keeps storage portable across engines.

This guide provides a vendor-neutral evaluation framework for 9 leading platforms, weighing table-format and catalog strategy, workload fit across SQL analytics and AI/ML, governance reach, and consumption-cost control so you can keep your data open and your compute swappable rather than locked to one engine.


Section 2

Why Data Lakehouse Platforms Matter for Enterprise Strategy

Data Lakehouse Platforms matter because they establish an open table format and catalog, determining data portability as engines like Databricks, Snowflake, Fabric, or BigQuery evolve. This architecture, often using Apache Iceberg or Delta, has become the system of record for analytics and the feature store for AI, preserving leverage against proprietary lock-in.

The decision that outlives the others is your open table format and the catalog that governs it, because together they determine whether storage stays portable while engines come and go. Selection then balances workload gravity — heavy data engineering and AI pull toward Databricks and Spark, while broad SQL analytics and operational simplicity favor Snowflake, Fabric, or BigQuery — against how aggressively you want to avoid proprietary lock-in.

🎯
Strategic Impact
The lakehouse has quietly become the system of record for analytics and the feature store for AI — which is why the table format and catalog are now architecture decisions, not implementation details. Standardize on an open format (Iceberg or Delta) and an interoperable catalog (Unity, Polaris/Open Catalog, or an Iceberg REST endpoint) and engine choice becomes reversible. Commit to a proprietary format and catalog for early convenience and “one copy of the data” quietly becomes one vendor’s compute — the most expensive form of lock-in in the modern data stack.

Apache Iceberg has become the interoperability lingua franca that lets multiple engines read and write the same governed tables, eroding the moat around any single proprietary format. Both incumbents have responded: Databricks acquired Tabular and now exposes Unity Catalog tables through the Iceberg REST API, while Snowflake created Polaris (donated to the Apache Software Foundation) and added full Iceberg support. Weigh how genuinely open each platform’s Iceberg story is — native managed tables with catalog interoperability and credential vending, versus a thin import path or read-only mirror — because that openness is what preserves your leverage over time.


Section 3

Should you build or buy Data Lakehouse Platforms?

You should buy a data lakehouse platform, as hand-rolling one from raw Parquet and a homegrown metastore is rarely defensible given mature table formats and catalogs. The key decision is architectural: choose an open table format like Iceberg, an interoperable catalog, and whether to anchor on a single managed platform (Databricks, Snowflake, BigQuery, Fabric) or a multi-engine open stack (Dremio, Starburst, Onehouse, Cloudera). Avoid proprietary table formats and catalogs to prevent vendor lock-in.

Almost nobody builds a lakehouse from raw Parquet and a homegrown metastore anymore — the table formats and catalogs are mature enough that hand-rolling is rarely defensible. The real decisions are architectural: which open table format becomes your standard, which catalog governs it, whether you anchor on a single managed platform or a multi-engine open stack, and how much of the table-maintenance toil (compaction, clustering, snapshot expiry) you want the vendor to absorb. Frame the choice around data gravity and lock-in tolerance, not the feature checklist.

Your Situation Recommended Path Rationale
Heavy data engineering and AI/ML on large, varied data Spark-native lakehouse (Databricks) Unified pipelines, notebooks, ML, and SQL on one runtime with Delta/Iceberg interoperability beats stitching a warehouse to a separate ML stack when data engineering is the center of gravity.
Broad SQL analytics and data sharing, lean platform team Managed SQL platform (Snowflake, BigQuery, Fabric) Near-zero infrastructure, fast time-to-value, and governed sharing matter more than engine flexibility when most workloads are BI and SQL; all three now read and write Iceberg to keep storage portable.
Open-format mandate — avoid engine lock-in by policy Open Iceberg stack (Dremio, Starburst, Onehouse) Standardize on Iceberg in your own object storage with a REST catalog, then attach Trino, Spark, Flink, or Dremio per workload — the storage layer is governed and the compute is genuinely swappable.
Already deep in a hyperscaler (Azure, GCP, or AWS) Native cloud lakehouse (Fabric / BigLake / S3 Tables) OneLake, BigLake, and S3 Tables fold the lakehouse into existing identity, billing, and security; the integration and procurement savings often outweigh a marginally stronger standalone engine.
Hybrid or data-sovereignty constraints, large on-prem estate Hybrid open lakehouse (Cloudera) When data must stay on-prem or in a private cloud for regulatory or gravity reasons, a platform that runs the same Iceberg lakehouse across public, private, and on-prem avoids a forced all-in cloud migration.
⚠️
Common Pitfall
The most common lakehouse mistake is committing to a proprietary table format and catalog for early convenience, then discovering that “one copy of the data” is effectively trapped behind a single vendor’s compute. The catalog is the real lock-in surface, not the file format: a platform can write open Iceberg files and still gate every read, write, and grant through a closed catalog. Standardize on an open format and an interoperable catalog from the start, and treat engine choice as the reversible decision it should be.

Section 4

How do you evaluate Data Lakehouse Platforms?

To evaluate Data Lakehouse Platforms, prioritize open format and catalog interoperability (25%), ensuring native read/write to Apache Iceberg or Delta Lake, and Iceberg REST Catalog support. Next, assess query engine performance (20%) on your own data, data engineering/streaming/AI/ML capabilities (20%), and governance/security/lineage (15%). Operational simplicity (10%) and the cost model (10%) are also key. Test interoperability by writing to a governed table from an unaffiliated engine like Trino or open-source Spark.

Weight these domains against your workload mix and lock-in tolerance. For most enterprises, openness of the table format and catalog now outranks raw single-engine benchmark speed, because the format decision is the one that is expensive to reverse. Score engines on your own data and query patterns, not vendor TPC-DS numbers.

Capability Domain Weight What to Evaluate
Open Format & Catalog Interoperability 25% Native read AND write to Apache Iceberg and/or Delta Lake, Iceberg REST Catalog support, credential vending for external engines, cross-format bridges (Delta UniForm, Apache XTable), and whether grants and lineage travel with the table when another engine reads it
Query Engine Performance & Concurrency 20% Vectorized execution (Photon, Arrow/Gandiva, native engines), caching and materialization (reflections, result cache, Warp Speed), high-concurrency BI behavior, autoscaling, and predictable performance on your own data — not vendor benchmarks
Data Engineering, Streaming & AI/ML 20% Batch and streaming ingestion into open tables, incremental/CDC and upsert support, orchestration, ML lifecycle (feature store, training, model serving), notebook and Python/Spark depth, and native LLM/agentic and vector capabilities
Governance, Security & Lineage 15% A unified catalog spanning tables, files, ML models and (increasingly) unstructured data; fine-grained RBAC/ABAC, row/column masking, data sharing, automated lineage, and consistent policy enforcement across every engine that touches the data
Operational Simplicity & Table Maintenance 10% Automated compaction, clustering, snapshot expiry and orphan-file cleanup; serverless vs. cluster sizing; multi-cloud and hybrid/on-prem reach; admin and FinOps tooling; and how much table toil the team must own versus the platform absorbing it
Cost Model & Consumption Control 10% Consumption unit (DBU, credit, capacity unit, bytes/slots), separation of storage and compute, idle-suspend and autoscaling guardrails, egress and cross-region exposure, workload isolation, and the FinOps tooling to attribute and cap spend
💡
Evaluation Tip
Run the interoperability test the vendors hope you skip: in your POC, create a governed table in the platform, then read AND write it from a second, unaffiliated engine (Trino, open-source Spark, or DuckDB) through the Iceberg REST catalog — and confirm that row/column security and lineage still apply on that path. Many platforms demo “open” as read-only export of stale snapshots; the ones that let an outside engine safely write back through a governed catalog are the ones that actually preserve your exit option.

Section 5

Which vendors lead in Data Lakehouse Platforms?

Consider Databricks for unified data engineering and AI, Snowflake for managed SQL analytics and sharing, and Microsoft Fabric for an Azure-native SaaS platform. Google BigQuery offers serverless analytics, while AWS provides a composable Iceberg lakehouse. Dremio specializes in high-performance SQL directly on Iceberg, and hybrid incumbents like Cloudera support on-prem and sovereignty needs.

9 vendors evaluated — positioning and best fit at a glance
Vendor Positioning Best for
Databricks Leader — Spark + AI Data-intensive organizations unifying data engineering, ML, and AI — and teams that want Spark depth without giving up open formats
Snowflake Leader — Managed SQL Analytics-led organizations that prize operational simplicity and governed data sharing, and want Iceberg openness without running infrastructure
Microsoft Fabric Strong — Azure-Native SaaS Microsoft-centric enterprises wanting a single SaaS platform from ingestion through Power BI, with governance via Purview
Google BigQuery / BigLake Strong — Serverless Google Cloud–native organizations wanting serverless analytics with embedded ML and a managed Iceberg catalog
AWS (Athena / EMR + S3 Tables) Strong — Composable Stack AWS-centric teams that want to compose an open Iceberg lakehouse from managed building blocks rather than buy a single platform
Dremio Strong — Open Iceberg Organizations standardizing on Iceberg that want fast, governed self-service SQL and a semantic layer on data they keep in open storage
Starburst Strong — Trino + Federation Enterprises that need to query across many sources today and want a managed Trino-on-Iceberg lakehouse and governed data products
Onehouse Emerging — Universal Lakehouse Teams that want a managed, vendor-neutral lakehouse foundation with best-in-class incremental ingestion and the freedom to bring any engine
Cloudera Niche — Hybrid & On-Prem Regulated and hybrid enterprises needing an Iceberg lakehouse that spans on-prem and cloud under one governance and security model

The market splits into four camps that increasingly overlap. Spark-native platforms (Databricks) lead with data engineering and AI; managed SQL platforms (Snowflake, BigQuery, Microsoft Fabric) lead with analytics simplicity and governed sharing; open-engine specialists (Dremio, Starburst, Onehouse) sell engine-and-format neutrality on storage you control; and hybrid incumbents (Cloudera) carry the on-prem and sovereignty cases. The convergence point is Apache Iceberg: nearly every platform here now reads and writes it, so shortlists increasingly compare governance reach, catalog openness, and cost model rather than whether a vendor “does” the lakehouse at all.

Watch the catalog layer specifically. Databricks open-sourced Unity Catalog and exposes it via the Iceberg REST API; Snowflake created Polaris and donated it to the Apache Software Foundation (shipping it commercially as Open Catalog); AWS, Google, and Cloudera each run their own Iceberg REST endpoints. The format war is effectively settling; the catalog war — over who governs and vends credentials for your one copy of data — is the live front.

Databricks

Leader — Spark + AI

Strengths: Pioneered the lakehouse and remains the strongest unified platform for data engineering, ML, and AI on one runtime; Delta Lake plus the Photon engine for fast SQL; Unity Catalog (now open-sourced) governs tables, models, and files across formats. After acquiring Tabular it added full managed Iceberg tables exposed through the Iceberg REST API, and UniForm lets Delta tables be read by Iceberg and Hudi clients. Considerations: Premium pricing and DBU-plus-cloud-compute costs that escalate quickly without disciplined cluster governance; Spark fluency still helps for advanced work; the deepest value sits inside the Databricks runtime even though the formats and catalog are open. Pure-SQL BI shops may find it heavier than a managed warehouse.

Best for: Data-intensive organizations unifying data engineering, ML, and AI — and teams that want Spark depth without giving up open formats

Snowflake

Leader — Managed SQL

Strengths: The easiest-to-operate cloud data platform, with clean separation of storage and compute, strong governed data sharing and Marketplace, and excellent SQL performance for analytics. Now ships full Apache Iceberg support (including v3) so Snowflake-grade performance, governance, and sharing apply to open tables, and created Polaris — donated to Apache and offered commercially as Open Catalog — as a vendor-neutral Iceberg REST catalog. Considerations: Credit-based consumption pricing can make budgets unpredictable without warehouse-sizing and auto-suspend discipline; data engineering and ML (Snowpark, Cortex) are maturing but trail Databricks for heavy Spark and custom ML; Horizon governance is strongest on Snowflake-managed data.

Best for: Analytics-led organizations that prize operational simplicity and governed data sharing, and want Iceberg openness without running infrastructure

Microsoft Fabric

Strong — Azure-Native SaaS

Strengths: An all-in-one SaaS analytics platform unifying data engineering, warehousing, real-time, and Power BI on OneLake, a single logical data lake. Delta Lake is the default open format, and OneLake metadata virtualization makes Iceberg tables readable as Delta (and vice versa) across Fabric engines; shortcuts reference data in S3, ADLS, and Iceberg sources without copying. Deep Microsoft 365, Purview, and Power BI integration. Considerations: Capacity-unit pricing pools compute across all workloads, so noisy neighbors and capacity sizing need active management; newer than incumbents and still maturing in spots; strongest value assumes a Microsoft-centric estate; Iceberg support is improving but Delta is the first-class citizen.

Best for: Microsoft-centric enterprises wanting a single SaaS platform from ingestion through Power BI, with governance via Purview

Google BigQuery / BigLake

Strong — Serverless

Strengths: Serverless architecture with effectively zero infrastructure to manage, strong embedded ML (BigQuery ML, Vertex AI) and cost-performance for ad-hoc SQL. BigLake (now “Lakehouse for Apache Iceberg”) adds managed BigQuery tables for Iceberg plus BigLake metastore — a fully managed, Spanner-backed Iceberg REST catalog — so Spark, Trino, and BigQuery share one governed copy with credential-vended access. Considerations: Tightest value inside the Google Cloud ecosystem; slot- and bytes-based pricing has its own complexity and needs reservations to stay predictable; data engineering tooling is lighter than Databricks; some advanced Iceberg DML capabilities are still rolling out.

Best for: Google Cloud–native organizations wanting serverless analytics with embedded ML and a managed Iceberg catalog

AWS (Athena / EMR + S3 Tables)

Strong — Composable Stack

Strengths: Not one product but a composable Iceberg lakehouse: S3 Tables provides a managed Iceberg bucket type with automatic compaction and maintenance, the Glue Data Catalog and Glue Iceberg REST endpoint govern tables, and Athena, EMR, and Redshift query them, unified under SageMaker Lakehouse. Lake Formation supplies fine-grained governance, and the whole stack speaks the Iceberg REST spec so PyIceberg and open Spark connect natively. Considerations: Assembling Glue, Lake Formation, S3 Tables, Athena, and EMR into a coherent platform takes more architectural ownership than a turnkey product; governance and cost span multiple services; best leveraged by teams already fluent in AWS data services.

Best for: AWS-centric teams that want to compose an open Iceberg lakehouse from managed building blocks rather than buy a single platform

Dremio

Strong — Open Iceberg

Strengths: A high-performance, Arrow-native SQL lakehouse engine that runs directly on Apache Iceberg without format conversion; Reflections accelerate queries by learning workload patterns and refreshing incrementally; a semantic layer and Open Catalog (built on Apache Polaris and Nessie, with Git-like branching) target self-service analytics directly on your object storage. Increasingly positioned for agentic and AI workloads. Considerations: Narrower than the hyperscalers — an analytics-and-query lakehouse, not a full data-engineering-plus-ML suite, so heavy pipelines and ML often pair it with other tools; smaller ecosystem and support footprint than the incumbents; most differentiated specifically for Iceberg shops.

Best for: Organizations standardizing on Iceberg that want fast, governed self-service SQL and a semantic layer on data they keep in open storage

Starburst

Strong — Trino + Federation

Strengths: The commercial Trino company, offering both massively parallel federated query across many sources and a fully managed open lakehouse (“Icehouse”: Trino plus Iceberg) in Galaxy. Warp Speed adds lakehouse-native indexing and caching; streaming ingest lands data into Iceberg in your own bucket; and governed data products plus an MCP interface expose curated data to analysts and AI agents. Considerations: Federation breadth can mask data-quality and performance issues if used to avoid consolidation; running Trino well at scale takes expertise (Galaxy mitigates this); positioned as engine-and-access layer rather than an end-to-end ML platform.

Best for: Enterprises that need to query across many sources today and want a managed Trino-on-Iceberg lakehouse and governed data products

Onehouse

Emerging — Universal Lakehouse

Strengths: Founded by the creators of Apache Hudi, Onehouse delivers a fully managed, format-agnostic lakehouse: ingestion and table management built on Hudi’s strong incremental/CDC and upsert tooling, Apache XTable for interoperability across Hudi, Delta, and Iceberg, and Open Engines to deploy Flink, Trino, and Ray on its managed Compute Runtime. The pitch is open data foundations you own, with the toil automated. Considerations: Younger and smaller than the established platforms, with a correspondingly smaller ecosystem and reference base; most compelling when low-latency incremental ingestion and genuine multi-format neutrality are first-order requirements; not a BI or semantic-layer destination on its own.

Best for: Teams that want a managed, vendor-neutral lakehouse foundation with best-in-class incremental ingestion and the freedom to bring any engine

Cloudera

Niche — Hybrid & On-Prem

Strengths: An Open Data Lakehouse on Apache Iceberg that runs the same way across public cloud, private cloud, and on-premises — the strongest answer when data must stay in your data center for sovereignty, latency, or cost. Added an Iceberg REST Catalog for zero-copy cross-engine sharing and a Lakehouse Optimizer that automates Iceberg table maintenance, with mature multi-function analytics and security inherited from CDP. Considerations: Operationally heavier and less cloud-native-slick than the SaaS platforms; ecosystem momentum and developer mindshare trail the cloud leaders; best justified by genuine hybrid or on-prem requirements rather than chosen for a greenfield all-cloud build.

Best for: Regulated and hybrid enterprises needing an Iceberg lakehouse that spans on-prem and cloud under one governance and security model
🔎
Market Insight
The table-format war is effectively over — Apache Iceberg won as the interoperability standard, and even Delta and Hudi now interoperate with it through UniForm and XTable. The live battle has moved up the stack to the catalog: whoever governs and vends credentials for your one copy of data holds the real leverage. Databricks (Unity Catalog) and Snowflake (Polaris/Open Catalog) are racing to be that neutral catalog, while every hyperscaler ships an Iceberg REST endpoint. Evaluate the catalog as carefully as the engine — it is where the next decade of lock-in, or openness, is being decided.

Section 6

How much should you budget for Data Lakehouse Platforms?

Budgeting for a Data Lakehouse Platform is primarily consumption-based, with costs varying by unit (DBUs, credits, capacity units, slots, or bytes scanned). Most platforms separate storage from compute, with key cost drivers including query efficiency, idle-suspend discipline, and workload isolation. Databricks, Snowflake, and Microsoft Fabric are Premium to Moderate–Premium, while Google BigQuery, AWS, and Dremio are Moderate.

Lakehouse pricing is overwhelmingly consumption-based, but the unit varies — DBUs, credits, capacity units, slots, or bytes scanned — and that unit, more than any headline rate, determines what you pay as workloads grow. Most platforms separate storage (cheap object storage you often pay your cloud provider for directly) from compute (the expensive, elastic part), so the real cost levers are query efficiency, idle-suspend discipline, and workload isolation. Model spend against your concurrency and refresh patterns, and price in cross-region and egress charges for shared or federated data.

Vendor Pricing Model Relative Tier Key Cost Drivers
Databricks Consumption per DBU + underlying cloud compute/storage Premium DBU consumption by workload type (jobs, all-purpose, SQL), serverless vs. classic compute, separate cloud VM bill, idle clusters, edition tier
Snowflake Consumption by credit (per-second warehouse compute) + storage Premium Warehouse size and uptime, auto-suspend hygiene, edition (Enterprise / Business Critical), serverless features, data-sharing and egress
Microsoft Fabric Capacity-unit subscription (pooled across workloads) + OneLake storage Moderate–Premium Capacity SKU size, workload concurrency against shared capacity, OneLake storage, Power BI seats, pause/scale discipline
Google BigQuery On-demand bytes scanned or slot reservations + storage Moderate Query bytes scanned vs. committed slots, edition (Standard / Enterprise), BigLake/Iceberg storage, streaming, BI Engine and ML usage
AWS (Athena/EMR + S3 Tables) Per-service consumption (Athena bytes / EMR compute / S3 Tables) Moderate Athena data scanned, EMR cluster hours, S3 Tables storage and maintenance, Glue catalog requests, Lake Formation, cross-AZ traffic
Dremio Subscription / consumption (cloud or self-managed) Moderate Compute engine size and uptime, Reflection storage and refresh, edition tier, self-managed vs. Dremio Cloud, supported user count
Starburst Consumption (Galaxy) or subscription (Enterprise) Moderate–Premium Cluster compute and uptime, Warp Speed caching, number and breadth of federated connectors, managed vs. self-hosted, support tier
Onehouse Managed consumption on Compute Runtime + your object storage Moderate Ingestion and table-management volume, Compute Runtime usage, engines deployed via Open Engines, data volume under management
Cloudera Subscription / capacity (CDP) across cloud, private, and on-prem Moderate at scale Compute capacity and node count, deployment mix (public / private / on-prem), data services enabled, support tier, hardware for on-prem
3-Year TCO Formula
TCO = (Compute consumption + Object Storage + Ingestion + Egress/Cross-region) × 36 months + Data Engineering & Platform FTE + Migration + Training − Legacy DW/ETL Savings − Avoided Lock-in / Engine-portability Value

Section 7

How long does implementation take for Data Lakehouse Platforms?

Implementing a Data Lakehouse Platform typically takes 10-14 months to reach optimization and full governance. The initial standardization of format and catalog occurs in months 1-2, followed by foundation and first workloads in months 3-5. Scaling and multi-engine integration, including engines like Trino and Spark, takes place in months 6-9.

Sequence the rollout around the durable decisions first — format, catalog, and governance — then layer engines and workloads on top. The teams that struggle are the ones that migrate workloads before standardizing the table format and catalog, then have to re-platform their “open” data later.

Phase 1
Standardize Format & Catalog (Months 1–2)

Choose your open table format (Iceberg and/or Delta) and the catalog that governs it, define the object-storage layout and naming, and set governance, RBAC, and lineage policies up front. Validate the interoperability path — a second engine reading and writing through the REST catalog — before any workload depends on it.

Phase 2
Foundation & First Workloads (Months 3–5)

Stand up the primary engine, wire identity (SSO/SCIM) and the governance catalog, build ingestion into open tables (batch and CDC/streaming), and migrate a high-value but bounded set of workloads. Establish automated table maintenance — compaction, clustering, snapshot expiry — from day one.

Phase 3
Scale & Multi-Engine (Months 6–9)

Expand to full production, attach additional engines per workload (Trino/Spark for engineering, the warehouse for BI, ML for data science) against the same governed tables, onboard users, and operationalize cost controls — warehouse sizing, auto-suspend, and workload isolation.

Phase 4
Optimize & Govern (Months 10–14)

Tune query and storage cost, refine clustering and materializations/reflections, harden FinOps attribution and budgets, extend governance to ML models and unstructured/AI data, and review the open-format exit path is still real — that you could move engines without moving data.


Section 8

What should you ask vendors about Data Lakehouse Platforms?

Use this checklist during evaluation to confirm each shortlisted platform keeps your data open, governed, and portable — not just fast in a demo.


Questions buyers ask

Frequently asked questions about Data Lakehouse Platforms

When would a 'Managed SQL platform' like Snowflake or BigQuery be a better fit than a 'Spark-native lakehouse' like Databricks, given the latter’s unified capabilities?

A Managed SQL platform is a better fit when broad SQL analytics and data sharing are the priority, and the platform team is lean. Snowflake or BigQuery offer near-zero infrastructure and fast time-to-value, which matters more than engine flexibility when most workloads are BI and SQL, even though Databricks excels at heavy data engineering and AI/ML.

For an organization already deep in AWS, what are the specific trade-offs of choosing the native cloud lakehouse (Athena/EMR + S3 Tables) versus a managed open Iceberg stack like Dremio or Starburst?

Choosing the native AWS lakehouse means composing Glue, Lake Formation, S3 Tables, Athena, and EMR, requiring more architectural ownership. While Dremio or Starburst offer a more turnkey managed Iceberg stack, the AWS native option provides integration and procurement savings within the existing identity, billing, and security, which can outweigh a marginally stronger standalone engine.

What are the hidden cost drivers to watch out for with Databricks, beyond the listed DBU consumption and cloud compute?

Beyond DBU consumption and underlying cloud compute, watch for costs from idle clusters, serverless vs. classic compute choices, and the specific edition tier. Databricks' premium pricing can escalate quickly without disciplined cluster governance, making these factors significant in the overall budget.

Our organization has a strict open-format mandate to avoid engine lock-in. What are the specific challenges or limitations we might encounter with an 'Open Iceberg stack' like Dremio, Starburst, or Onehouse compared to a hyperscaler’s offering?

An Open Iceberg stack standardizes on Iceberg in your own object storage, allowing genuinely swappable compute. However, these platforms might be narrower than hyperscalers; Dremio, for example, is an analytics-and-query lakehouse, not a full data-engineering-plus-ML suite, potentially requiring pairing with other tools for heavy pipelines and ML.

If we’re considering Microsoft Fabric, what specific aspects of its 'capacity-unit subscription' pricing model should we be most concerned about during rollout to avoid unexpected costs?

With Microsoft Fabric’s capacity-unit subscription, the primary concern is managing workload concurrency against shared capacity. Capacity-unit pricing pools compute across all workloads, meaning 'noisy neighbors' can impact performance and cost. Active management of capacity SKU size and pause/scale discipline are crucial to avoid unexpected expenses.

Section 9

Related Resources

Spotlight
Available placement · independent of CIOPages editorial
From the directory

Vendors in this category

Directory listings for the Data Lakehouse Platforms space— independent of this guide’s evaluation. Compare profiles in the CIOPages directory, or claim yours.

Apache Drill Claim
Apache Hive Claim
Apache Hudi Claim
Apache Pig Claim
Cloudera Claim
Delta Lake Claim
DuckDB Claim
Exasol Claim
Firebolt Claim
Browse all in the directory Represent one of these? Claim or spotlight your company
Tags:LakehouseDatabricksSnowflakeMicrosoft FabricBigQueryApache IcebergDelta LakeApache HudiUnity CatalogPolarisDremioStarburstOnehouseCloudera