Amundsen
A data discovery and metadata catalog platform, originally built at Lyft, for finding and understanding datasets across an organization.
Batch and streaming engines, orchestration, query federation and the BI layer that sits in front of them. Browse all 51 products in this category from WIEWAVE's published marketplace catalog — each one a hardened, maintained image you can deploy from your own cloud account.
Search the full catalogSelect any image for its supported distributions, marketplace availability and common questions.
A data discovery and metadata catalog platform, originally built at Lyft, for finding and understanding datasets across an organization.
Airflow with the scheduler, webserver and a Celery or Local executor already chosen, backed by PostgreSQL.
A cross-language, in-memory columnar data format and library for fast analytics and data interchange between systems.
Druid with its service roles laid out for a real cluster rather than the single-machine quickstart.
Flink with JobManager and TaskManager roles split, checkpointing configured and state backend pointed at durable storage.
Hadoop with HDFS and YARN configured for single-node or cluster start-up, and the classic port map documented on the box.
NiFi with TLS and a single-user authenticator configured, so the canvas isn't sitting open on first boot.
Spark with a matched JDK, executor memory sized to the instance and standalone or YARN submission both ready.
Superset with metadata in PostgreSQL, Celery workers for async queries and caching configured for dashboard load.
A web-based notebook for interactive data analytics and visualisation, with built-in support for Spark, SQL, and Scala.
Streaming SQL engine for Apache Kafka that runs continuous queries, joins and materialized views over topics without writing Java consumer code.
Analytics package for assembling a unified customer view on Google BigQuery; included data models and scripts depend on the publisher's build.
Parallel computing library that scales NumPy, pandas and scikit-learn workloads across cores or a cluster, with a scheduler and diagnostic dashboard.
Bundled platform image covering ingestion, storage and query for analytics workloads; the engines actually included depend on the publisher's configuration.
Client environment for the Databricks lakehouse, bundling the Databricks CLI, SDK and Apache Spark libraries for submitting jobs and querying tables.
Open-source business intelligence tool for building dashboards and charts over SQL databases and uploaded files, accessed through a browser interface.
Data quality library for Spark that declares constraints and computes metrics over large datasets, flagging anomalies inside ETL pipelines.
Dremio serves SQL queries directly over lake storage such as S3 and ADLS using Apache Arrow, with a web UI for datasets and reflections.
Node image intended for Elastic MapReduce style big data clusters; the exact Hadoop and Spark components included are not specified.
Fathom is a self-hosted, privacy-focused website analytics tool that tracks visitor traffic without cookies or personal data collection.
Feldera is an incremental computation engine that runs SQL queries continuously over streaming and batch data to produce always-up-to-date results.
Folium is a Python library that generates interactive Leaflet.js maps for visualising geospatial data in web pages and notebooks.
Tooling image for working with the AWS Glue Data Catalog metadata store, where table schemas are registered for engines such as Athena and Spark.
Real-time web log analyser that parses Apache, Nginx and S3 access logs and renders traffic reports in the terminal or as HTML dashboards.
Analytics image for exploring graph-structured data and producing insights; consult the publisher listing for the exact components and interfaces that are included.
Declarative orchestration platform where YAML-defined flows schedule and coordinate data tasks, with a web UI, executor and Postgres metadata store.
Business intelligence suite for reporting, dashboards, OLAP and data mining, deployed on its Java application server with a metadata database.
Privacy-focused web analytics server that records page views without cookies, storing data locally and exposing dashboards from a single lightweight binary.
Python framework from Spotify for building batch pipelines, resolving task dependencies, handling scheduling and exposing a central view of workflow state.
Matomo as self-hosted web analytics, with archiving moved to cron so reporting stays fast as data accumulates.
Privacy-focused web analytics server that records page views without cookies, storing data in SQLite and exposing a self-hosted reporting dashboard.
Metabase running against an external application database, so upgrades don't put your saved questions at risk.
Drop-in pandas replacement that distributes dataframe operations across available cores or a cluster using Ray or Dask, installed with its Python dependencies.
Transactional catalog for data lakes that gives Apache Iceberg tables Git-like branches, tags and commits, exposed over a REST API.
Deploys the Optable software stack; the upstream project behind this name is not clearly identified, so verify capabilities against vendor documentation first.
Kettle-based ETL suite with the Spoon designer and command-line runners for building, scheduling and executing data transformation and job workflows.
Privacy-focused, cookie-free web analytics server built in Elixir, storing pageview events in ClickHouse with PostgreSQL holding accounts and site settings.
Self-hosted product analytics suite with event capture, funnels, session replay and feature flags, backed by ClickHouse, PostgreSQL, Kafka and Redis.
Workflow orchestration platform for Python data pipelines providing scheduling, retries and observability, with a server UI and PostgreSQL metadata store.
Distributed SQL engine for interactive analytics that federates queries across Hive, S3, MySQL and other sources using coordinator and worker nodes.
The Python API for Apache Spark, preinstalled with a local Spark runtime for writing DataFrame and SQL jobs against distributed datasets.
Partner-packaged deployment of Prophecy, a low-code data engineering interface that generates Spark pipelines from a visual editor for ETL workloads.
Query and dashboard tool that connects to SQL and NoSQL sources, letting analysts share visualisations and scheduled reports from a browser.
Python module that connects notebooks and scripts to a SAS session, letting SAS procedures and datasets be driven directly from Python code.
Open-source control panel that tracks keyword rankings, backlinks, sitemaps and site audits across multiple websites, running on PHP with MySQL.
MPP analytical database delivering sub-second queries over real-time and batch data, speaking the MySQL protocol and querying lakehouse table formats directly.
Graphical ETL and data integration studio built on Eclipse, used to design jobs that extract, transform and load data across files, databases and APIs.
Streaming SQL engine built on ClickHouse that runs continuous queries over Kafka and other event sources alongside historical analytical tables.
Trino set up to federate queries across object storage, relational databases and Hive-compatible catalogs.
Python library for lazy, out-of-core DataFrames that processes billion-row tabular datasets through memory mapping, shipped with Jupyter and scientific packages.
Declarative grammar for interactive visualizations, providing the Vega and Vega-Lite runtimes plus a CLI that renders specifications to SVG, PNG or PDF.
Name the product and the distribution — we'll build, harden, certify and publish it on Azure, AWS or Google Cloud Marketplace.