AMD Instinct MI350X
288GB HBM3e flagship with native FP4/FP6 — the TCO leader for frontier inference
Reserved and on-demand, self-service access
Reserved and committed-use discounts up to 60%. All prices GBP, ex-VAT.
MI350X is AMD's flagship accelerator built on the CDNA 4 architecture, delivering unprecedented memory capacity and the new FP4/FP6 datapaths required by frontier model serving.
Highlights
- Industry-leading 288GB HBM3e per card
- Native FP4/FP6 support for next-gen inference
- Up to 40% lower TCO vs H200 on LLM inference
Software stack
Best-fit scenarios
- Frontier LLM training (100B+ parameters)
- Long-context inference (1M+ tokens)
- Mixture-of-Experts hosting
Pricing by configuration
All available SKUs for this GPU across bare-metal, VM and container form factors. Reserved and committed-use discounts available.
On-demand starting at £7.40/h
| Configuration | GPUs | VRAM total | CPU | Memory | Storage | Network | £/hr | £/mo |
|---|---|---|---|---|---|---|---|---|
Bare Metal — 8× MI350X Popular mi350x-bm-8 | 8 | 2.3 TB | 192 vCPU (dual EPYC 9554) | 2 TB DDR5 | 8× 3.84 TB NVMe | RoCE, full bisection | £54.40 | £39,600 |
Virtual Machine — 1× MI350X mi350x-vm-1 | 1 | 288GB (PCIe passthrough) | 32 vCPU | 256 GB DDR5 | 1× 1.92 TB NVMe | 100G | £8.00 | £5,850 |
Container — 1× MI350X mi350x-ctr-1 | 1 | 288GB | 16 vCPU | 128 GB | 500 GB ephemeral | 100G | £7.40 | £5,400 |
Prices shown are list prices in GBP (ex-VAT). Reserved capacity and committed-use discounts up to 60% available — contact sales for a custom quote.
Available form factors
Same silicon, three consumption models — pick the deployment model that matches your workload, then see every GPU that fits.
Whole server, no virtualization — raw performance and full root.
KVM-based with GPU passthrough — multi-tenant isolation.
Pay per pod-hour with elastic Kubernetes scaling.
Technova at a glance
Full-stack AI cloud
GPUs, networking, storage, managed Slurm & Kubernetes on one platform.
Manual UK deployment
Every resource deployed and configured by our UK engineering team.
UK & US data residency
UK workloads stay in Glasgow & London. US workloads in Texas. You choose the region.
Reliable
Historical uptime over 99.9% with sensible SLAs and fair compensation.
Dual-stack CUDA + ROCm
Professional full-stack expertise across both NVIDIA and AMD.
World-class support
Proactive support from ML engineers and infrastructure specialists.
Ready to deploy AMD Instinct MI350X?
Spin up a private cluster in days, or talk to our sales engineers about custom configurations.