Executive summary

The old warehouse-versus-lakehouse argument is no longer the most useful way to read the market. The important question is where control sits after storage, table format, catalog, governance, and compute can be pulled apart.

Open table formats reduce one kind of lock-in. They make it easier for data to live in cloud object storage and be read by more than one engine. Snowflake makes that pitch directly with its lakehouse analytics and open-format work: Lakehouse Analytics with Open Table Formats and Make Your Lakehouse AI-Ready. Databricks makes the same broad argument from the Delta Lake side: Lakehouse Storage. ClickHouse explains the architecture from a more specialized OLAP angle in its pages on data lakehouses, open table formats, and data lake support.

But open storage does not make the whole platform portable. The real control point moves up into the catalog, governance model, workload engine, and operating habits around the data estate. Snowflake and Databricks are fighting to own that layer. ClickHouse sits in a different position: it is not trying to be the full enterprise lakehouse in the same way, but it can win high-volume analytical workloads when query speed or cost-performance matters more than a single governed operating model.

Why now

The lakehouse stack has become more modular. A modern architecture can be split into cloud object storage, an open table format such as Iceberg or Delta Lake, a catalog and governance layer, and one or more compute engines. That separation creates a new strategic question for buyers: does one vendor need to own the whole stack, or can the enterprise mix catalogs and engines around shared data?

The answer is not simple. Open table formats help with storage interoperability, but the day-to-day operating model still depends on governance, lineage, identity, SQL dialects, orchestration, and workload tuning. Snowflake's open-format positioning is partly about reducing fear of storage lock-in while keeping the buyer inside Snowflake's governed experience. Databricks' Unity Catalog and Tabular/Iceberg moves point in the opposite direction: make the open lakehouse feel unified around Databricks' control plane. Starburst's discussion of Snowflake, Databricks, Tabular, and Iceberg is useful because it frames the fight as catalog and interoperability rather than file format alone.

The second reason this matters now is budget pressure. Consumption pricing is attractive at adoption time and painful when workloads sprawl. Procurement guides and practitioner comparisons keep circling the same issue: warehouse and lakehouse platforms are easy to start using, but harder to forecast and optimize at scale: Monetizely procurement guide, Databricks vs Snowflake technical comparison, and Snowflake revenue model. That does not make every cost comparison neutral. It does explain why buyers care about bringing specialized engines to specific workloads.

Value chain and buyer logic

The upstream layer is cloud object storage. This is where the raw data lives, usually inside AWS, Azure, or Google Cloud. Open table formats sit above that storage layer and define how tables, partitions, schema evolution, and metadata are represented. Iceberg and Delta Lake carry most of the Snowflake-Databricks strategic weight; Hudi remains relevant for streaming and upsert-heavy use cases. ClickHouse's open table formats overview and independent explainers such as Open Table Formats and the Open Data Lakehouse capture that layer.

The next layer is the catalog. This is where openness starts to get complicated. A catalog does more than point to files. It can manage access policy, governance context, schema definitions, lineage, and what different teams believe to be the authoritative data estate. Snowflake's Polaris work and Databricks' Unity Catalog strategy both matter because catalog control can preserve platform power even when the files underneath are more open. Databricks' Catalog Commits announcement is a good example of how the table-format fight becomes a catalog fight.

The compute layer is where workloads diverge. Snowflake wants to keep SQL, Python, AI, data engineering, and governance close to its managed platform. Databricks wants SQL warehousing, machine learning, notebooks, and data engineering to live around the lakehouse. ClickHouse wants to win workloads where real-time analytical speed and high concurrency matter. Its own comparisons should be treated as vendor evidence rather than neutral proof, but they still show where ClickHouse wants the fight to happen: cloud data warehouse comparison, real-time analytics platforms, and ClickHouse vs Databricks vs Snowflake join performance.

The buyer logic follows the same map. Central data teams buy governance, reliability, security review, and a managed operating model. Engineering teams care about workload fit, latency, and cost-performance. Finance teams care about consumption surprise. Executives care about consolidation. That is why the same enterprise can rationally use Snowflake for governed BI, Databricks for data science work, and ClickHouse for product analytics or event-heavy workloads.

Incumbents and challengers

Snowflake and Databricks are the two central incumbents because each is expanding into the other's historical zone. Snowflake is no longer just the managed cloud data warehouse. Databricks is no longer just the Spark and ML lakehouse platform. Public analysis from Public Comps and PitchGrade describes the same convergence: each vendor wants more of the enterprise data operating system.

Snowflake has scale and public-company visibility. Its FY2026 filings and related financial summaries show a large, still-growing business, but also the cost of defending and expanding the platform: Snowflake FY2026 10-K via Last10K, Snowflake quarterly results, and the SEC filing archive. The useful point is not that GAAP losses prove weakness. Enterprise software accounting, stock-based compensation, and growth investment can all blur that read. The useful point is that Snowflake has enough scale to keep pushing into adjacent workloads.

Databricks is harder to compare because it is private, but its strategic direction is clear. SQL warehousing, Unity Catalog, Delta Lake, and AI/ML all pull toward the same end state: a unified governed platform for data and AI workloads. The Tabular acquisition and Iceberg work matter because they keep Databricks close to the open-format conversation even when buyers are worried about lock-in.

ClickHouse is the challenger with the clearest specialized wedge. It is not a full substitute for Snowflake or Databricks across the whole data estate. It is better understood as a high-performance analytical engine that can peel away workloads where Snowflake or Databricks feel too broad, too expensive, or too slow for the use case. Coverage of ClickHouse's position against Snowflake and Databricks, including the Runtime piece and Finance Yahoo coverage, supports the market narrative, but the stronger evidence is architectural: ClickHouse has a plausible reason to win real-time OLAP and embedded analytics workloads.

Hyperscalers are the other force in the market. Microsoft Fabric, BigQuery, Redshift, and native cloud analytics services can enter the buying discussion through existing cloud commitments. That does not mean they replace Snowflake or Databricks one-for-one. It means pure-play vendors have to defend technical differentiation and procurement independence.

Where control and profit accrue

Control accrues in three places.

First, the catalog. If the catalog becomes the policy, lineage, and governance center, it can be more important than the open table format underneath. This is why Polaris, Unity Catalog, and open catalog interoperability matter.

Second, workload gravity. Once dashboards, pipelines, notebooks, access rules, UDFs, and orchestration jobs accumulate around a platform, migration becomes more than moving files. Tools such as XTable can help with table-format interoperability, but they do not erase all of the operating work around a mature data estate: Apache XTable explainer.

Third, performance-sensitive compute. If the workload is real-time analytics, user-facing analytics, logs, observability, or event-heavy AI application telemetry, specialized engines can win even when the wider enterprise platform stays with Snowflake or Databricks: ClickHouse real-time analytics platforms.

Profit follows those control points. Snowflake and Databricks make money when customers prefer an integrated managed platform over assembling the stack themselves. ClickHouse makes money when a specific workload is painful enough that buyers want a specialized engine. Hyperscalers make money when analytics is pulled back into the wider cloud relationship.

Bear case

The bear case for Snowflake and Databricks is that open storage plus better catalog interoperability weakens the premium attached to the full platform. If buyers can keep data in open formats, govern it through portable catalogs, and route more workloads to specialized engines, then the platform bundle gets less powerful.

That bear case should not be overstated. Fragmentation has a cost. Enterprises still need access control, auditability, support, integration, and reliability. Many buyers may still pay for the managed platform even when a cheaper engine exists for some workloads.

The bear case for ClickHouse is the mirror image. Its wedge is real, but bounded. If the buyer's main problem is governed enterprise data across many teams, Snowflake and Databricks remain difficult to displace. ClickHouse still has to show that workload-level excellence can expand into a durable platform position without inheriting all the complexity that made the incumbents useful in the first place: https://clickhouse.com/resources/engineering/real-time-analytics-platforms-a-practical-comparison.

What would change the thesis

The thesis weakens if catalog portability becomes routine. If buyers can move between Unity Catalog, Snowflake Polaris, and open-source alternatives without long dual-running projects, the catalog-control argument loses force.

The thesis strengthens if Snowflake and Databricks keep proving that adjacent workloads stick. Snowflake needs evidence that Python, AI, and data engineering work can expand the account without damaging operating leverage. Databricks needs evidence that SQL warehousing can become a durable core workload.

ClickHouse becomes more strategically important if public customer evidence shows it taking large production workloads from the lakehouse incumbents, beyond vendor benchmarks or specialized engineering teams: https://runtime.news/clickhouse-takes-aim-at-databricks-and-snowflake.

Watch next

Watch the Iceberg and Delta Lake ecosystem for signs that table-format interoperability is turning into real workload portability. Watch Polaris and Unity Catalog for evidence that catalog control is becoming more open or more consolidated. Watch Snowflake's filings for whether AI and data engineering expansion improves the operating story. Watch Databricks' SQL warehouse adoption for signs that it is becoming a mainstream warehouse replacement. Watch ClickHouse for proof that event-heavy analytics and AI telemetry can become a broad enterprise wedge rather than a specialized engine market.

Sources