Managed Slurm
Distributed model training and HPC job scheduling — fully managed
Overview
A fully managed Slurm batch scheduler on Kubernetes — purpose-built for large-scale GPU training and HPC workloads.
Technova Managed Slurm deploys and operates a production-grade Slurm workload manager on top of Kubernetes, tuned for distributed AI training and traditional HPC. We handle the scheduler controller, job accounting, autoscaling, and MPI/NCCL cluster bring-up so your team can submit jobs instead of managing infrastructure. Priority queues and fair-share policies keep R&D timelines on track, while familiar sbatch/srun tooling means HPC teams can transition to AI without rewriting scripts.
Capabilities
Priority queues & fair-share
Scheduled queues with priority, fair-share and backfill policies — designed for multi-hundred-GPU training runs that need predictable start times and resource allocation.
Mixed workload support
Batch training, containerised services and interactive development sessions run side by side under a single scheduler, with consistent GPU-aware placement across all job types.
HPC-compatible interface
Standard Slurm commands — sbatch, srun, sacct, squeue — work out of the box. Existing MPI workflows, job scripts and partition configurations carry over without modification.
Scheduler on Kubernetes
The Slurm controller, workers and accounting daemons run as Kubernetes workloads, inheriting the platform's reliability, rolling upgrades and integrated observability stack.
Topology-aware autoscaling
Node autoscaler places jobs with NVLink/RoCE affinity and MIG-aware partitioning for RTX PRO 6000 and MI300X. Worker pools expand from zero based on queue depth and contract back to zero when idle.
Job-level telemetry
Each job gets a live dashboard covering GPU utilisation, collective bandwidth (NCCL/RCCL), storage I/O and accrued cost — with configurable alerts for under-utilisation and slow nodes.
Benefits
Integrations
How it works
Submit
Submit your sbatch script with priority and partition hints. Standard Slurm job descriptors are accepted without modification.
Schedule
The scheduler evaluates queue priority and topology, then places the job on the optimal GPU pool with fair-share enforcement.
Scale
Kubernetes-autoscaled worker nodes start in seconds. The GPU operator handles NCCL/RCCL initialisation and device plugin registration.
Observe
Track utilisation, network throughput and cost on a per-job dashboard. Alerts fire on performance drift; a full audit log is retained after completion.
Pricing
Ready to put Technova to work?
Talk to our team about a custom GPU cluster, managed Slurm or one of our vertical AI solutions.