Llama 3 70B
Deployment of Meta's 70-billion-parameter Llama 3 model, sized for multi-GPU inference with transformer libraries and CUDA dependencies preinstalled.
Meta's seven-billion-parameter Llama 2 model with tokenizer and PyTorch inference scripts prepared for GPU-based text generation and fine-tuning. Built and maintained by WIEWAVE for Azure Marketplace, AWS Marketplace and Google Cloud Marketplace, on Ubuntu and Debian.
| Built & maintained by | WIEWAVE |
|---|---|
| Category | AI & Machine Learning |
| Operating systems | Ubuntu, Debian |
| Marketplaces | Azure Marketplace, AWS Marketplace and Google Cloud Marketplace |
| Offer types | Public listing, with private offers on request |
| Marketplace | Status |
|---|---|
| Azure Marketplace | Published |
| AWS Marketplace | Published |
| Google Cloud Marketplace | Published |
Need it on another marketplace, or as a private offer for your organisation? cloud@wiewave.com
Each image follows its distribution's own provisioning model, package manager and security tooling — not one build relabelled several times.
LTS and interim releases, Minimal and Pro variants, built to Canonical's cloud-image conventions.
Stable and oldstable, with backports where a workload needs a newer runtime than the release ships.
The same four steps behind every offer we've published, including the hardening and CIS Benchmark checks every build goes through.
The distribution, licensing model and target marketplaces are agreed before anything is built.
Packer templates, Ansible provisioning and a pinned package set — then hardened, scanned and checked against the CIS Benchmark for its distribution.
Taken through each cloud's own certification pipeline before it goes live on the marketplace.
Rebuilt on the upstream security cadence and re-published, with old versions retired without breaking deployments.
If yours isn't here, ask our marketplace team directly.
GPU-ready training and inference images with drivers, CUDA and frameworks already matched to each other.
Deployment of Meta's 70-billion-parameter Llama 3 model, sized for multi-GPU inference with transformer libraries and CUDA dependencies preinstalled.
Eight-billion-parameter Llama 3 chat model served through a Python inference stack, suited to single-GPU generation and lightweight fine-tuning experiments.
Compiled llama.cpp HTTP server exposing an OpenAI-compatible API for running quantized GGUF models on CPU or GPU hardware.
Data framework connecting LLMs to private documents, providing indexing, retrieval and query engines in Python with common loader and vector store integrations.
Tooling image aimed at inspecting and evaluating large language model behaviour; the packaged application's exact feature set should be confirmed before use.
Drop-in OpenAI-compatible API server that runs local models for text, embeddings, images and audio without calling external inference services.
Tell us the distribution, the marketplace and the commercial model — we'll build, certify and publish it as a public listing or a private offer.