Deequ with Apache Spark
Data quality library for Spark that declares constraints and computes metrics over large datasets, flagging anomalies inside ETL pipelines.
Open-source business intelligence tool for building dashboards and charts over SQL databases and uploaded files, accessed through a browser interface. Built and maintained by WIEWAVE for Azure Marketplace, on Ubuntu and Debian.
| Built & maintained by | WIEWAVE |
|---|---|
| Category | Analytics & Big Data |
| Operating systems | Ubuntu, Debian |
| Marketplaces | Azure Marketplace |
| Offer types | Public listing, with private offers on request |
| Marketplace | Status |
|---|---|
| Azure Marketplace | Published |
| AWS Marketplace | Available on request |
| Google Cloud Marketplace | Available on request |
Need it on another marketplace, or as a private offer for your organisation? cloud@wiewave.com
Each image follows its distribution's own provisioning model, package manager and security tooling — not one build relabelled several times.
LTS and interim releases, Minimal and Pro variants, built to Canonical's cloud-image conventions.
Stable and oldstable, with backports where a workload needs a newer runtime than the release ships.
The same four steps behind every offer we've published, including the hardening and CIS Benchmark checks every build goes through.
The distribution, licensing model and target marketplaces are agreed before anything is built.
Packer templates, Ansible provisioning and a pinned package set — then hardened, scanned and checked against the CIS Benchmark for its distribution.
Taken through each cloud's own certification pipeline before it goes live on the marketplace.
Rebuilt on the upstream security cadence and re-published, with old versions retired without breaking deployments.
If yours isn't here, ask our marketplace team directly.
Batch and streaming engines, orchestration, query federation and the BI layer that sits in front of them.
Data quality library for Spark that declares constraints and computes metrics over large datasets, flagging anomalies inside ETL pipelines.
Dremio serves SQL queries directly over lake storage such as S3 and ADLS using Apache Arrow, with a web UI for datasets and reflections.
Node image intended for Elastic MapReduce style big data clusters; the exact Hadoop and Spark components included are not specified.
Fathom is a self-hosted, privacy-focused website analytics tool that tracks visitor traffic without cookies or personal data collection.
Feldera is an incremental computation engine that runs SQL queries continuously over streaming and batch data to produce always-up-to-date results.
Folium is a Python library that generates interactive Leaflet.js maps for visualising geospatial data in web pages and notebooks.
Tell us the distribution, the marketplace and the commercial model — we'll build, certify and publish it as a public listing or a private offer.