KoboldCpp
Single-binary inference server for GGUF language models built on llama.cpp, exposing a chat web UI and an OpenAI-compatible API.
Deep learning stack pairing the Keras high-level API with a TensorFlow backend, preinstalled with Python bindings for training and inference. Built and maintained by WIEWAVE for Azure Marketplace, AWS Marketplace and Google Cloud Marketplace, on Ubuntu and Debian.
| Built & maintained by | WIEWAVE |
|---|---|
| Category | AI & Machine Learning |
| Operating systems | Ubuntu, Debian |
| Marketplaces | Azure Marketplace, AWS Marketplace and Google Cloud Marketplace |
| Offer types | Public listing, with private offers on request |
| Marketplace | Status |
|---|---|
| Azure Marketplace | Published |
| AWS Marketplace | Published |
| Google Cloud Marketplace | Published |
Need it on another marketplace, or as a private offer for your organisation? cloud@wiewave.com
Each image follows its distribution's own provisioning model, package manager and security tooling — not one build relabelled several times.
LTS and interim releases, Minimal and Pro variants, built to Canonical's cloud-image conventions.
Stable and oldstable, with backports where a workload needs a newer runtime than the release ships.
The same four steps behind every offer we've published, including the hardening and CIS Benchmark checks every build goes through.
The distribution, licensing model and target marketplaces are agreed before anything is built.
Packer templates, Ansible provisioning and a pinned package set — then hardened, scanned and checked against the CIS Benchmark for its distribution.
Taken through each cloud's own certification pipeline before it goes live on the marketplace.
Rebuilt on the upstream security cadence and re-published, with old versions retired without breaking deployments.
If yours isn't here, ask our marketplace team directly.
GPU-ready training and inference images with drivers, CUDA and frameworks already matched to each other.
Single-binary inference server for GGUF language models built on llama.cpp, exposing a chat web UI and an OpenAI-compatible API.
Differentiable computer vision library for PyTorch, supplying image transforms, geometry, filtering and feature operators as GPU-ready tensor operations.
Kubernetes custom resources for serving machine learning models, handling autoscaling, canary rollouts and standardised inference endpoints across frameworks.
Preinstalled LangChain libraries alongside Flowise, a drag-and-drop builder for wiring LLM chains, agents and retrieval flows through a browser canvas.
Combines the LangChain framework with Langflow's visual editor so you can prototype retrieval-augmented generation pipelines and export them as runnable Python code.
Visual low-code environment for composing LLM applications, running the Langflow server with its Python backend, component library and web canvas.
Tell us the distribution, the marketplace and the commercial model — we'll build, certify and publish it as a public listing or a private offer.