Lakehouse platforms & table formats · Cloudera, Inc.
Cloudera Data Platform
Enterprise data platform unifying data engineering, warehousing, ML, and streaming across hybrid and multi-cloud environments.
Cloudera Data Platform (CDP) is an enterprise lakehouse suite descended from the Hadoop-era Cloudera and Hortonworks stacks, now rebuilt around open table formats and containerized services. It packages Cloudera Data Warehouse, Data Engineering (Spark), Machine Learning, and Data Flow (based on Apache NiFi and Kafka) under shared security, governance, and metadata services (SDX), and can run on-premises, on any major public cloud, or in a hybrid mix of both. CDP is aimed at large, often regulated enterprises with existing Hadoop-ecosystem investments (HDFS, Hive, HBase, Kafka, Spark) who need a path to a governed lakehouse without a full platform replacement. It differs from cloud-native lakehouses like Databricks in emphasizing hybrid and on-premises deployment alongside cloud.
At a glance
| Vendor | Cloudera, Inc. |
|---|---|
| Pricing model | Quote only |
| Free tier | No |
| Deployment | Cloud, Self-hosted |
| Open source | No |
| Best for | Large regulated enterprises with existing Hadoop-ecosystem investments needing hybrid or on-premises lakehouse deployment. |
Pricing
Consumption-based licensing (Cloudera Consumption Units) across compute and services; public pricing figures are not published and require contacting Cloudera sales.
Pricing has not been verified yet — see the vendor's site.
Features
- Unified data warehouse, data engineering, ML, and streaming services
- Shared governance and security via SDX (Shared Data Experience)
- Hybrid and multi-cloud deployment, including on-premises
- Built on open-source Hadoop-ecosystem components (Spark, Hive, Kafka, NiFi)
- Cloudera Machine Learning workbenches for data science teams
- Fine-grained data lineage and cataloging across workloads
- Support for open table formats including Apache Iceberg
Integrations
Profile last reviewed September 21, 2026