Tools

Query engines & federation

7 tools compared: how each is priced, where it runs, and what to consider instead.

Tool Pricing model Free tier Open source
Apache Flink Open-source stream-processing engine for stateful, low-latency event processing at scale, with batch support. Open source + paid Yes Yes
Apache Spark Open-source distributed processing engine for large-scale batch, SQL, streaming and machine-learning workloads. Open source + paid Yes Yes
Dask Open-source Python library that parallelizes NumPy, pandas and scikit-learn workflows across cores or a cluster. Open source + paid Yes Yes
Presto Open-source distributed SQL engine originally built at Facebook for interactive queries across large datasets. Open source + paid Yes Yes
Ray Open-source Python framework for distributing machine-learning training, tuning and general compute across a cluster. Open source + paid Yes Yes
Starburst Commercial data platform built on Trino, offering managed and self-hosted federated SQL query with enterprise governance. Usage-based Yes No
Trino Open-source distributed SQL engine that queries data across warehouses, lakes and databases without moving it. Open source + paid Yes Yes

In the index now