Managed Slurm

Managed Slurm

Distributed model training and HPC job scheduling — fully managed

Overview

A fully managed Slurm batch scheduler on Kubernetes — purpose-built for large-scale GPU training and HPC workloads.

Technova Managed Slurm deploys and operates a production-grade Slurm workload manager on top of Kubernetes, tuned for distributed AI training and traditional HPC. We handle the scheduler controller, job accounting, autoscaling, and MPI/NCCL cluster bring-up so your team can submit jobs instead of managing infrastructure. Priority queues and fair-share policies keep R&D timelines on track, while familiar sbatch/srun tooling means HPC teams can transition to AI without rewriting scripts.

Capabilities

Priority queues & fair-share

Scheduled queues with priority, fair-share and backfill policies — designed for multi-hundred-GPU training runs that need predictable start times and resource allocation.

Mixed workload support

Batch training, containerised services and interactive development sessions run side by side under a single scheduler, with consistent GPU-aware placement across all job types.

HPC-compatible interface

Standard Slurm commands — sbatch, srun, sacct, squeue — work out of the box. Existing MPI workflows, job scripts and partition configurations carry over without modification.

Scheduler on Kubernetes

The Slurm controller, workers and accounting daemons run as Kubernetes workloads, inheriting the platform's reliability, rolling upgrades and integrated observability stack.

Topology-aware autoscaling

Node autoscaler places jobs with NVLink/RoCE affinity and MIG-aware partitioning for RTX PRO 6000 and MI300X. Worker pools expand from zero based on queue depth and contract back to zero when idle.

Job-level telemetry

Each job gets a live dashboard covering GPU utilisation, collective bandwidth (NCCL/RCCL), storage I/O and accrued cost — with configurable alerts for under-utilisation and slow nodes.

Benefits

Predictable R&D schedules via priority queues and fair-share allocation
Run batch training, services and interactive jobs under one scheduler
Keep existing Slurm scripts, MPI workflows and partition configs unchanged
No need for in-house HPC operations staff — we manage the controller and accounting
Natively connected to Technova shared filesystem and object storage
Pay only for active GPUs — idle clusters scale to zero automatically

Integrations

PyTorch DDP
DeepSpeed
Megatron-LM
JAX
MPI
NCCL / RCCL
sbatch / srun / sacct

How it works

Step 01

Submit

Submit your sbatch script with priority and partition hints. Standard Slurm job descriptors are accepted without modification.

Step 02

Schedule

The scheduler evaluates queue priority and topology, then places the job on the optimal GPU pool with fair-share enforcement.

Step 03

Scale

Kubernetes-autoscaled worker nodes start in seconds. The GPU operator handles NCCL/RCCL initialisation and device plugin registration.

Step 04

Observe

Track utilisation, network throughput and cost on a per-job dashboard. Alerts fire on performance drift; a full audit log is retained after completion.

Pricing

Controller fee£0.10 / node-hour
Managed service fee15% of compute spend
Idle clustersFree (scale to zero)
Multi-tenant accountsIncluded

Ready to put Technova to work?

Talk to our team about a custom GPU cluster, managed Slurm or one of our vertical AI solutions.