RDMA over Converged Ethernet (RoCE v2)
The fabric connecting our AMD Instinct and NVIDIA GPU clusters
Overview
RoCE v2 is the default interconnect for our AMD Instinct MI350X, MI300X and MI325X instances, as well as for VM and Container form factors where customers need RDMA performance without a separate InfiniBand fabric. All GPU models on our platform — across Bare Metal, Virtual Machines and Containers — can be connected via RoCE, with Managed Slurm and Kubernetes scheduling jobs directly onto RoCE-connected GPU pools.
Key specifications
- Bandwidth
- 800 Gbps per port
- Latency
- <2 µs port-to-port
- Topology
- Spine-leaf with full bisection
- Congestion Control
- PFC + ECN + DCQCN
- Collectives
- NCCL (NVIDIA) / RCCL (AMD)
- Available form factors
- Bare Metal, VM, Container
Core advantages
Connects every GPU on our platform
AMD MI350X, MI300X, MI325X, NVIDIA H200 and RTX PRO 6000 all ship with RoCE-capable NICs. Bare Metal, VM and Container instances can be mixed on the same RoCE fabric without re-cabling.
Integrated with Managed Slurm
Our Managed Slurm service places training jobs with RoCE topology awareness — NCCL and RCCL collectives run natively over RoCE, with per-job bandwidth monitoring built into the scheduler dashboard.
Integrated with Kubernetes Service
Kubernetes pods on TKS get RoCE-aware pod placement via the GPU operator. Volcano gang scheduling ensures distributed training pods land on the same RoCE fabric segment for low-latency AllReduce.
800G per port across all data centres
Glasgow, London and Texas data centres all run 800G RoCE spine-leaf fabrics with full bisection bandwidth — the same fabric that serves our AI Token Foundry, Intelligent Automation and Bioinformatics solutions.
Ideal for
Ready to put Technova to work?
Talk to our team about a custom GPU cluster, managed Slurm or one of our vertical AI solutions.