Storage & Compute tools

70 tools filed under Storage & Compute, in 8 categories. Open a profile for verified pricing, features and alternatives.

Cloud data warehouses

Tool Pricing model Free tier Open source
Amazon Redshift AWS's managed cloud data warehouse, available as provisioned clusters or an auto-scaling serverless option. Usage-based Yes No
Azure Synapse Analytics Microsoft's unified analytics service combining a dedicated SQL data warehouse, serverless SQL, and Apache Spark pools. Usage-based No No
Exasol In-memory, massively parallel analytics database available as a managed cloud service or self-hosted software. Quote only Yes No
Firebolt Cloud analytics database built for low-latency, high-concurrency queries with per-second usage billing. Usage-based Yes Yes
Google BigQuery Serverless cloud data warehouse billed by data scanned or by slot capacity, with no infrastructure to provision. Usage-based Yes No
Greenplum Open-source, massively parallel PostgreSQL-based data warehouse for large-scale analytical SQL. Open source + paid Yes Yes
Microsoft Fabric Microsoft's SaaS analytics platform unifying warehousing, lakehouse, pipelines, Power BI, and real-time analytics on OneLake. Usage-based Yes No
OpenText Vertica Columnar MPP analytics database, now owned by OpenText, deployable on-premises, in Kubernetes, or across major clouds. Quote only Yes No
Oracle Autonomous Data Warehouse Oracle's self-tuning, self-patching cloud data warehouse built on Oracle Database with elastic OCPU and storage scaling. Usage-based Yes No
SAP Datasphere SAP's cloud data warehousing and data-fabric service that virtualizes and harmonizes data across SAP and non-SAP sources. Quote only No No
SingleStore Distributed SQL database combining row and columnar storage for hybrid transactional/analytical (HTAP) workloads. Usage-based Yes No
Snowflake Cloud-native data warehouse that separates storage and compute and bills usage in per-second credits. Usage-based No No
Teradata VantageCloud Teradata's cloud-delivered evolution of its enterprise MPP data warehouse, unified under a single Teradata Unit pricing currency. Quote only No No

Graph databases & knowledge graphs

Tool Pricing model Free tier Open source
Amazon Neptune Fully managed AWS graph database supporting both property-graph (Gremlin, openCypher) and RDF (SPARQL) query models. Usage-based No No
ArangoDB Multi-model database combining graph, document and key/value storage under one engine and one query language, AQL. Quote only Yes No
Dgraph Open-source, natively distributed property-graph database with a GraphQL-derived query language, now stewarded by Hypermode. Open source + paid Yes Yes
JanusGraph Open-source distributed property-graph database queried with Gremlin, with pluggable storage and indexing backends. Open source + paid Yes Yes
Memgraph In-memory, Cypher-compatible property-graph database built in C++ for low-latency graph analytics and GraphRAG. Free tier + paid Yes No
NebulaGraph Open-source distributed property-graph database using nGQL, built for billion-edge-scale graphs at low latency. Free tier + paid Yes Yes
Neo4j Native property-graph database using Cypher, the language behind the new ISO GQL standard, for relationship-heavy analytics. Free tier + paid Yes No
Ontotext GraphDB RDF4J-compliant semantic graph database using SPARQL with OWL/RDFS reasoning, built for enterprise knowledge graphs. Free tier + paid Yes No
Stardog Enterprise RDF knowledge-graph platform using SPARQL, with OWL/RDFS reasoning and federation over relational sources. Free tier + paid Yes No
TigerGraph Massively parallel native graph database with its own GSQL language, built for deep multi-hop analytics at scale. Usage-based Yes No

In-process & embedded engines

Tool Pricing model Free tier Open source
Apache Arrow Language-independent columnar in-memory format that lets analytical engines exchange data without serialization overhead. Open source + paid Yes Yes
Apache DataFusion Open-source, embeddable Rust query engine and toolkit for building custom SQL and DataFrame engines on Apache Arrow. Open source + paid Yes Yes
chDB In-process OLAP SQL engine powered by ClickHouse, embedded directly in Python with no server to install. Open source + paid Yes Yes
clickhouse-local Single-binary command-line tool that runs ClickHouse's query engine locally over files, without installing a server. Open source + paid Yes Yes
DuckDB Open-source embedded analytical database that runs SQL queries in-process, without a server, on local or remote files. Open source + paid Yes Yes
MotherDuck Managed cloud service built on DuckDB, adding hybrid local/cloud execution, sharing, and scale-out storage. Free tier + paid Yes No
Polars Open-source, Rust-built in-process DataFrame library offering a fast, multi-threaded alternative to pandas for local analytics. Open source + paid Yes Yes
SQLite Serverless, in-process SQL database engine that reads and writes a single file, embedded inside applications worldwide. Open source + paid Yes Yes

Lakehouse platforms & table formats

Tool Pricing model Free tier Open source
Apache Hudi Open-source table format optimized for high-volume upserts, deletes, and incremental data-lake pipelines. Open source + paid Yes Yes
Apache Iceberg Open-source table format that adds warehouse-style transactions, schema evolution, and time travel to data-lake files. Open source + paid Yes Yes
Apache Paimon Open-source lake table format built for unified streaming and batch updates, originating from the Apache Flink community. Open source + paid Yes Yes
Cloudera Data Platform Enterprise data platform unifying data engineering, warehousing, ML, and streaming across hybrid and multi-cloud environments. Quote only No No
Databricks Managed lakehouse platform, built by Apache Spark's creators, combining Spark-based compute with governance and ML tooling. Usage-based Yes No
Delta Lake Open-source storage format that brings ACID transactions and versioning to Parquet files on a data lake. Open source + paid Yes Yes
Dremio SQL query and semantic layer engine for querying data lakes directly, built around Apache Iceberg and Apache Arrow. Usage-based Yes No
Onehouse Managed lakehouse service built by Apache Hudi's original creators, unifying ingestion and open table formats. Quote only No No

Query engines & federation

Tool Pricing model Free tier Open source
Apache Flink Open-source stream-processing engine for stateful, low-latency event processing at scale, with batch support. Open source + paid Yes Yes
Apache Spark Open-source distributed processing engine for large-scale batch, SQL, streaming and machine-learning workloads. Open source + paid Yes Yes
Dask Open-source Python library that parallelizes NumPy, pandas and scikit-learn workflows across cores or a cluster. Open source + paid Yes Yes
Presto Open-source distributed SQL engine originally built at Facebook for interactive queries across large datasets. Open source + paid Yes Yes
Ray Open-source Python framework for distributing machine-learning training, tuning and general compute across a cluster. Open source + paid Yes Yes
Starburst Commercial data platform built on Trino, offering managed and self-hosted federated SQL query with enterprise governance. Usage-based Yes No
Trino Open-source distributed SQL engine that queries data across warehouses, lakes and databases without moving it. Open source + paid Yes Yes

Real-time OLAP databases

Tool Pricing model Free tier Open source
Apache Doris Open-source MPP database combining real-time analytics with sub-second query response at large scale. Open source + paid Yes Yes
Apache Druid Open-source real-time analytics database that ingests streaming and batch data for interactive dashboards. Open source + paid Yes Yes
Apache Pinot Open-source distributed OLAP database built by LinkedIn for low-latency analytics serving at high query volume. Open source + paid Yes Yes
ClickHouse Open-source columnar database built for sub-second aggregate queries over billions of rows. Open source + paid Yes Yes
ClickHouse Cloud Fully managed cloud version of ClickHouse, billed by compute and storage rather than server count. Usage-based No No
Imply Managed and self-hosted Apache Druid platform built by Druid's original creators. Usage-based No No
StarRocks Open-source MPP database for sub-second analytics across real-time data and data-lake tables. Open source + paid Yes Yes
StarTree Fully managed cloud service for Apache Pinot, built by Pinot's original creators. Usage-based No No
Tinybird Managed real-time analytics platform, built on ClickHouse, for turning event streams into low-latency API endpoints. Free tier + paid Yes No

Time-series databases

Tool Pricing model Free tier Open source
Amazon Timestream AWS's managed time-series database, offered as a native serverless service or as managed InfluxDB. Usage-based No No
InfluxDB Purpose-built time-series database for metrics, sensor and IoT data, available self-hosted or as a managed cloud service. Free tier + paid Yes Yes
Prometheus Open-source monitoring system with a built-in time-series database, the de facto standard for Kubernetes and infrastructure metrics. Open source + paid Yes Yes
QuestDB High-throughput open-source time-series database with SQL, built for fast ingestion of market, IoT and metrics data. Free tier + paid Yes Yes
TimescaleDB Open-source time-series extension for PostgreSQL, giving full SQL and joins alongside time-series performance. Free tier + paid Yes Yes
VictoriaMetrics Open-source, Prometheus-compatible time-series database and monitoring backend built for high cardinality at low cost. Free tier + paid Yes Yes

Vector databases

Tool Pricing model Free tier Open source
Chroma Open-source embedding database designed for simple local-first development, with a managed serverless cloud tier. Free tier + paid Yes Yes
LanceDB Open-source, embedded vector database built on the Lance columnar format for multimodal AI data. Free tier + paid Yes Yes
Milvus Open-source vector database built for billion-scale similarity search, governed under the Linux Foundation. Open source + paid Yes Yes
pgvector Open-source PostgreSQL extension that adds vector similarity search directly inside a Postgres database. Open source + paid Yes Yes
Pinecone Fully managed, cloud-only vector database for similarity search and retrieval-augmented generation at scale. Free tier + paid Yes No
Qdrant Open-source vector database written in Rust, focused on filtered similarity search with a free managed cloud tier. Free tier + paid Yes Yes
Vespa Open-source big-data serving engine combining vector search, keyword search and ranking at large scale. Free tier + paid Yes Yes
Weaviate Open-source vector database with built-in hybrid (vector + keyword) search, available self-hosted or as managed cloud. Free tier + paid Yes Yes
Zilliz Cloud Fully managed cloud version of Milvus, built and operated by Milvus's original creators. Usage-based Yes No