Skip to content

SLURM

Prepare the Database

Clone mknoxnv/ubuntu-slurm — we'll take config and service file templates from it throughout controller and worker setup:

git clone https://github.com/mknoxnv/ubuntu-slurm.git

Tip

Preferably clone it into the shared NFS path so each node can copy from it, with one central point to modify files.

Install MariaDB

sudo apt-get install mariadb-server
sudo systemctl enable mysql
sudo systemctl start mysql

Enter MySQL to create and initialize the database:

sudo mysql -u root
create database slurm_acct_db;
create user 'slurm'@'localhost';
set password for 'slurm'@'localhost' = password('slurmdbpass');
grant usage on *.* to 'slurm'@'localhost';
grant all privileges on slurm_acct_db.* to 'slurm'@'localhost';
flush privileges;
exit

slurmdbpass is a live credential, not just a placeholder

This is still the actual password in production. It should be rotated — when you do, change it in two places together: the set password statement above, and PASSWORD in ubuntu-slurm/slurmdbd.conf below. They must match or slurmdbd won't be able to authenticate to the database.

Build SLURM from Source

Tip

Preferably download it into the shared NFS path so each node can use it.

wget https://download.schedmd.com/slurm/slurm-23.11.0.tar.bz2
tar xvjf slurm-23.11.0.tar.bz2
cd slurm-23.11.0

Configuration

During configuration, try to resolve most warnings and errors — they'll cause problems in later steps, forcing you to reconfigure and rebuild. Installing the needed libraries first helps:

sudo apt-get update
sudo apt-get install -y libnuma-dev
sudo apt-get install -y libpam0g-dev
sudo apt-get install -y libhdf5-dev
sudo apt-get install -y liblz4-dev libhwloc-dev
sudo apt-get install -y libgtk2.0-dev libglib2.0-dev
sudo apt-get install -y librdkafka-dev
sudo apt-get install -y libdbus-1-dev
sudo apt-get install -y check
sudo apt-get install -y liblua5.3-dev
sudo apt-get install -y libreadline-dev
sudo apt-get install -y linux-headers-$(uname -r)
sudo apt-get install -y freeipmi-tools libfreeipmi-dev
sudo apt-get install -y rrdtool librrd-dev
sudo apt-get install -y libjson-c-dev libjansson-dev
sudo apt-get install -y libjwt-dev
sudo apt-get install -y libhttp-parser-dev
sudo apt-get install -y libyaml-dev
sudo apt-get install -y man2html
sudo apt-get install -y mailutils
sudo apt-get install -y libhdf5-dev
sudo apt-get install -y libmysqlclient-dev

Note

Deactivate conda before doing the next step — it conflicts with the packages being used here.

./configure --prefix=/usr --sysconfdir=/etc/slurm --enable-pam --with-pam_dir=/lib/x86_64-linux-gnu/security/ --without-shared-libslurm

Some warnings at this stage are safe to ignore (e.g. around optional plugins you don't need).

Make

sudo make
sudo make contrib
sudo make install DESTDIR=/tmp/slurm-package

Exit the SLURM folder:

cd ..

Install fpm:

sudo apt-get install ruby ruby-dev rubygems build-essential
sudo gem install --no-document fpm
fpm -s dir -t deb -v 1.0 -n slurm-23.11.0 -C /tmp/slurm-package .

Install

The SLURM package is now ready and can be used on all nodes:

sudo dpkg -i slurm-23.11.0_1.0_amd64.deb

Or, if you reconfigured and are reinstalling:

sudo dpkg --install --force-overwrite slurm-23.11.0_1.0_amd64.deb

Install this same .deb on the Login Node and Worker Nodes too.

Configure & Start Services

Directories

sudo mkdir -p /etc/slurm /etc/slurm/prolog.d /etc/slurm/epilog.d /var/spool/slurm/ctld /var/spool/slurm/d /var/log/slurm
sudo chown slurm /var/spool/slurm/ctld /var/spool/slurm/d /var/log/slurm

This directory should also be created:

sudo mkdir -p /var/spool/slurm/d
sudo chown slurm /var/spool/slurm/d

And this one:

sudo mkdir -p /run/slurmd
sudo chown slurm:slurm /run/slurmd

slurmdbd.conf

Copy it before starting the SLURM services. The PASSWORD in this file must match the one set above in Prepare the Database:

sudo cp ubuntu-slurm/slurmdbd.conf /etc/slurm/
sudo chmod 600 /etc/slurm/slurmdbd.conf
sudo chown slurm:slurm /etc/slurm/slurmdbd.conf

Copy the SLURM control and db services:

sudo cp ubuntu-slurm/slurmdbd.service /etc/systemd/system/
sudo cp ubuntu-slurm/slurmctld.service /etc/systemd/system/

slurm.conf

Start from ubuntu-slurm/slurm.conf. You can check detected system info (note: GPUs aren't included here) with:

sudo slurmd -C

Key values to set: ControlMachine to the actual controller hostname, SelectType=select/cons_tres, and your compute nodes at the end of the file. Below is the current canonical config for this cluster — use it as the reference copy rather than reconstructing it from scratch:

/etc/slurm/slurm.conf
#
# slurm.conf file generated by configurator.html.
#
# See the slurm.conf man page for more information.
#
ClusterName=compute-cluster
ControlMachine=server02
#ControlAddr=
#BackupController=
#BackupAddr=
#
SlurmUser=slurm
#SlurmdUser=root
SlurmctldPort=6817
SlurmdPort=6818
AuthType=auth/munge
#JobCredentialPrivateKey=
#JobCredentialPublicCertificate=
StateSaveLocation=/var/spool/slurm/ctld
SlurmdSpoolDir=/var/spool/slurm/d
SwitchType=switch/none
MpiDefault=none
SlurmctldPidFile=/var/run/slurmctld.pid
SlurmdPidFile=/var/run/slurmd.pid
ProctrackType=proctrack/cgroup
PluginDir=/usr/lib/slurm
#FirstJobId=
ReturnToService=1
#MaxJobCount=
#PlugStackConfig=
#PropagatePrioProcess=
#PropagateResourceLimits=
#PropagateResourceLimitsExcept=
#Prolog=/etc/slurm/prolog.d/*
#Epilog=/etc/slurm/epilog.d/*
#SrunProlog=
#SrunEpilog=
#TaskProlog=
#TaskEpilog=
TaskPlugin=task/cgroup
#TrackWCKey=no
#TreeWidth=50
#TmpFS=
#UsePAM=
#
# TIMERS
SlurmctldTimeout=300
SlurmdTimeout=300
InactiveLimit=0
MinJobAge=300
KillWait=30
Waittime=0
#
#
LaunchParameters=use_interactive_step
#
# SCHEDULING
SchedulerType=sched/backfill
#SchedulerAuth=
SelectType=select/cons_tres
SelectTypeParameters=CR_Core_Memory,CR_CORE_DEFAULT_DIST_BLOCK,CR_ONE_TASK_PER_CORE
#FastSchedule=1
#PriorityType=priority/multifactor
#PriorityDecayHalfLife=14-0
#PriorityUsageResetPeriod=14-0
#PriorityWeightFairshare=100000
#PriorityWeightAge=1000
#PriorityWeightPartition=10000
#PriorityWeightJobSize=1000
#PriorityMaxAge=1-0
#
# LOGGING
SlurmctldDebug=3
SlurmctldLogFile=/var/log/slurmctld.log
SlurmdDebug=3
SlurmdLogFile=/var/log/slurmd.log
JobCompType=jobcomp/none
#JobCompLoc=
#
# ACCOUNTING
JobAcctGatherType=jobacct_gather/cgroup
#JobAcctGatherFrequency=30
#
AccountingStorageTRES=gres/gpu
DebugFlags=CPU_Bind,gres
AccountingStorageType=accounting_storage/slurmdbd
AccountingStorageHost=localhost
#AccountingStorageLoc=
AccountingStoragePass=/var/run/munge/munge.socket.2
AccountingStorageUser=slurm
AccountingStorageEnforce=qos,limits
#
# COMPUTE NODES
GresTypes=gpu
DefMemPerNode=64000
NodeName=server02 Gres=gpu:6 CPUs=255 RealMemory=2051942 State=UNKNOWN
NodeName=jrcai01 Gres=gpu:2 CPUs=48  RealMemory=64132 State=UNKNOWN
NodeName=jrcai02 Gres=gpu:2 CPUs=48 RealMemory=64124 State=UNKNOWN
NodeName=jrcai18 Gres=gpu:2 CPUs=32 RealMemory=257545 State=UNKNOWN
NodeName=jrcai08 Gres=gpu:3 CPUs=64 RealMemory=257541 State=UNKNOWN
NodeName=jrcai23 State=UNKNOWN
PartitionName=Normal Default=YES Nodes=server02 QoS=normal_partition MaxTime=24:00:00  State=UP
PartitionName=LoginNode  Nodes=jrcai23  State=DOWN Hidden=YES
PartitionName=RTX3090 Nodes=jrcai01,jrcai02,jrcai08 QoS=rtx3090_partition MaxTime=24:00:00  State=UP
PartitionName=A6000 Nodes=jrcai18  MaxTime=24:00:00 QoS=a6000_partition State=UP
PartitionName=interactive Nodes=server02,jrcai[01,02,08,18] QoS=interactive  MaxTime=1:00:00  State=UP

Two settings worth knowing about

  • AccountingStorageEnforce=qos,limits — actually enforces the QoS limits you set up (see Resource Limits and Deadline Budgets), rather than just tracking them.
  • LaunchParameters=use_interactive_step — makes salloc run a shell on the allocated node directly, instead of staying on the login node.

When adding a new compute node, append a NodeName= line like the ones above (use sudo slurmd -C on that node to get its detected specs), then copy the updated file to /etc/slurm/ everywhere:

sudo cp ubuntu-slurm/slurm.conf /etc/slurm/
sudo chown slurm:slurm /etc/slurm/slurm.conf

gres.conf

Edit the default gres.conf and add the GPUs that SLURM will manage:

sudo nano ubuntu-slurm/gres.conf

It should be something like this — one line per actual GPU device on the controller:

NodeName=server02 Name=gpu File=/dev/nvidia0
NodeName=server02 Name=gpu File=/dev/nvidia1
NodeName=server02 Name=gpu File=/dev/nvidia2
NodeName=server02 Name=gpu File=/dev/nvidia3
NodeName=server02 Name=gpu File=/dev/nvidia4
NodeName=server02 Name=gpu File=/dev/nvidia5

Then copy it to the system:

sudo cp ubuntu-slurm/gres.conf /etc/slurm/gres.conf

cgroup.conf

Edit the cgroup file:

sudo nano ubuntu-slurm/cgroup.conf

Comment out some lines, as follows:

#CgroupAutomount=yes
#CgroupReleaseAgentDir="/etc/slurm/cgroup"

ConstrainCores=yes
ConstrainDevices=yes
ConstrainRAMSpace=yes
#TaskAffinity=ye

Then copy it to the system:

sudo cp ubuntu-slurm/cgroup.conf /etc/slurm/cgroup.conf

cgroup_allowed_devices_file.conf

Copy it as-is:

sudo cp ubuntu-slurm/cgroup_allowed_devices_file.conf /etc/slurm/cgroup_allowed_devices_file.conf

Configure cgroups (GRUB)

sudo nano /etc/default/grub

Add cgroup and swap to GRUB_CMDLINE_LINUX:

GRUB_CMDLINE_LINUX="cgroup_enable=memory swapaccount=1"
sudo update-grub

Start the SLURM Services

sudo systemctl daemon-reload
sudo systemctl enable slurmdbd
sudo systemctl start slurmdbd
sudo systemctl enable slurmctld
sudo systemctl start slurmctld

If the controller is also going to be a worker/compute node:

sudo cp ubuntu-slurm/slurmd.service /etc/systemd/system/
sudo systemctl enable slurmd
sudo systemctl start slurmd
sudo systemctl status slurmd

Debugging

Check the status of the SLURM database service:

sudo systemctl status slurmdbd

Check the status of the SLURM controller:

sudo systemctl status slurmctld

If the status output isn't detailed enough, start with higher verbosity:

sudo /usr/sbin/slurmctld -Dvvv
sudo /usr/sbin/slurmdbd -Dvvv