
Platform Services
Managed Slurm and Kubernetes for AI — GPU-aware scheduling, multi-tenant isolation, integrated observability and automated lifecycle management.
Managed Slurm
A fully managed Slurm batch scheduler on Kubernetes — purpose-built for large-scale GPU training and HPC workloads.
- Priority queues & fair-share
- Mixed workload support
- HPC-compatible interface
Kubernetes Service
A managed Kubernetes control plane engineered for GPU workloads — from single-node tests to multi-region training clusters.
- Rapid cluster provisioning
- Tenant isolation
- Federated multi-region scaling
One platform across the AI lifecycle
From initial prototyping through foundation-scale training to production inference — scheduling, isolation and observability are built in.
Faster iteration cycles
Spin up test environments in minutes, then promote them to production with GPU-aware scheduling and automatic autoscaling.
Lower operational overhead
Reduce outages and manual toil with smart GPU placement, automated lifecycle management and integrated monitoring.
Clear cost predictability
Lifecycle-managed bare-metal and VM instances with scale-to-zero billing make capacity planning straightforward and transparent.
What you get with every managed service
24/7 SRE
PagerDuty, escalation trees, RCA on every Sev-1.
Upstream-aligned
We track upstream releases and roll out tested upgrades.
Multi-tenant ready
Namespaces, quotas, RBAC — production-grade by default.
GPU-aware
Schedulers that understand GPUs, fabric and storage.
Observability
Prometheus, Grafana, OpenCost and audit logs out of the box.
Day-2 automation
Upgrades, scaling, patching — fully automated and observable.
Frequently asked questions
Common questions about our Platform Services.
Can I run mixed workloads across the platform?
Yes. Slurm and Kubernetes are integrated on the same fabric, so batch training runs, containerised inference services and interactive notebooks can coexist under consistent, GPU-aware scheduling policies.
When should I use Slurm vs Kubernetes?
Slurm suits large-scale distributed training and traditional HPC batch jobs that need scheduled queues and MPI/NCCL collectives. Kubernetes (TKS) is better for containerised services, model serving endpoints and MLOps pipelines that benefit from rapid provisioning, autoscaling and namespace isolation. The two services interoperate for hybrid workflows.
How do teams, projects and permissions work — can I enforce team-level isolation?
Identity is organised into organisations and projects, each with RBAC roles governing exactly who can see and do what. Resources are logically segregated per project, all actions are audit-logged, and enterprise SSO is supported via SAML or OIDC. Network and orchestration-level isolation ensure teams only access their own workloads and data.
How is performance tuned for NVIDIA and AMD hardware?
Our reference platform configurations align with NVIDIA and AMD guidance: curated GPU node profiles, validated CUDA/ROCm and driver releases, and tuned network/storage topologies. This delivers consistent, repeatable performance across both silicon families for training and inference.
Start building on managed Slurm and Kubernetes
Begin with a managed scheduler today, or talk to our engineers about a custom configuration tailored to your workloads.