GPU instances on Akamai Cloud, purpose-built for distributed AI inference. Up to 8 RTX PRO 6000 Blackwell GPUs per Linode, 96 GB of GDDR7 VRAM per card, and the throughput to serve large models at scale — delivered as an on-demand OpEx, not a CapEx.
Akamai Cloud computing
Akamai Cloud lets you choose the right GPU for your workload. And for the scale and complexity of data-center workloads — especially AI inference — NVIDIA RTX PRO™ 6000 Blackwell Server Edition is engineered for the job. Deploy cloud GPUs on demand, pay per hour, and scale ML, AI, and data-processing workloads with confidence.
A GPU Linode’s vCPU cores are dedicated to you — never shared between customers. Your software runs at peak speed and efficiency, even at 100% CPU all day, every day.
Turn GPU CapEx into OpEx. Predictable hourly pricing and low-cost egress (US$0.005/GB in most regions) let you test and scale without draining your infrastructure budget.
Up to 60% lower latency and 3× higher throughput — for up to 86% less cost with image generation and AI workloads compared to equivalent hyperscaler GPUs.
Configure compute, memory, and storage to optimize for your workload, then launch GPU instances across Akamai Cloud locations — meeting your users and data wherever they are.
Manage infrastructure flexibly with our UI, API, CLI, Terraform provider, and developer tool integrations. Custom images and CI/CD pipelines supported.
Every GPU Linode ships with email and phone support for all customers. Backups keep your data safe with automated daily, weekly, and biweekly snapshots.
NVIDIA RTX PRO™ 6000 Blackwell Server Edition
Each card pairs 5th-generation Tensor Cores and 4th-generation RT Cores with 96 GB of GDDR7 ECC memory — 24,064 CUDA cores pushing 120 TFLOPS of FP32 performance. That architecture is ideally suited for AI inferencing, providing the throughput needed for large-scale model deployment.
Fleet includes NVIDIA RTX PRO™ 6000 Blackwell Server Edition, NVIDIA RTX™ 4000 Ada, and NVIDIA Quadro RTX™ 6000. Plans are equally at home with graphics, visualization, and video workflows.
Plans & pricing
The same RTX PRO 6000 Blackwell performance, in sizes that fit your workload — with dedicated vCPU cores and generous RAM. Pricing is per GPU per hour (US$3.00/GPU-hr on-demand, as published; confirm live pricing at linode.com/pricing).
| Configuration | 1 GPU | 2 GPU | 4 GPU | 8 GPU |
|---|---|---|---|---|
| GPU cards | 1 | 2 | 4 | 8 |
| GPU memory (VRAM) | 96 GB | 192 GB | 384 GB | 768 GB |
| vCPU cores (dedicated) | 16 | 32 | 64 | 128 |
| Memory (RAM) | 176 GB | 352 GB | 704 GB | 1,408 GB |
| Storage | 1,024 GB | – | – | 6,597 GB |
| Network bandwidth (outbound) | 16 Gbps | |||
8-card plans add vNUMA: Virtual Non-Uniform Memory Access exposes the native PCIe device topology to the VM, so job schedulers and GPU communication libraries map hardware-localized resources efficiently. Some new accounts may require a $100 deposit to deploy GPU Linodes.
Recommended workloads
GPU Linodes are optimized for high-throughput, low-latency inference at production scale — with large GPU memory, next-generation Tensor Cores, and architectural efficiency that sustains token throughput and fast first-response latency.
96 GB of VRAM per GPU in a high-throughput architecture mitigates the “bottlenecking” found in shared cloud resources — powering agents that process text, visuals, and audio in real time.
Process, reason, and respond within a line of thought. Akamai Cloud has the GPUs to take conversations global — scaling out while keeping latency direct and predictable.
RT, Tensor, and CUDA cores accelerate perception pipelines — from robotics and autonomous systems to real-time analytics on streaming video.
GPU-accelerated encoding converts massive streams in real time. Pair with the NVIDIA RTX™ 4000 Ada plan for live 8K transcoding and AI upscaling at the edge.
Real-time ray tracing and advanced shading — mesh shading, variable rate shading, and multi-view rendering — in a single GPU.
Give Hadoop, Spark, and Storm the parallel compute they need to process terabyte-scale datasets — the volume, velocity, and variety of modern data.
How it works
Choose NVIDIA RTX PRO™ 6000 Blackwell or Ada architectures, matching performance requirements to your budget.
Customize compute, memory, and storage to optimize for your unique workload.
Launch GPU instances in minutes across Akamai Cloud locations, meeting your users and data wherever they are.
Achieve ambitious AI goals with high-performance NVIDIA compute designed for immersion, autonomy, and scale.
Global availability
NVIDIA RTX PRO 6000 Blackwell Server Edition is rolling out in 20 regions worldwide (limited deployment availability) — from Amsterdam to Tokyo, Singapore to Toronto.
RTX 4000 Ada is available in Chicago 2, Frankfurt 2, Osaka, Paris, Seattle, and Singapore. Quadro RTX 6000 is available in Atlanta, Newark, Frankfurt, Mumbai, and Singapore.
Start with up to US$100 in Akamai Cloud credits. Launch NVIDIA RTX PRO™ 6000 Blackwell GPU Linodes in minutes — managed Kubernetes (LKE), backups, and 24/7 support included.