SLURM¶
Set Up SLURM User and Group¶
Every machine needs SLURM installed, with the slurm user/group sharing the same UID and GID across all nodes — in practice, these are the only UID/GID that need to stay synchronized for the cluster to work. On all machines:
Install Built SLURM¶
Install dependencies:
sudo apt-get update
sudo apt-get install -y libnuma-dev
sudo apt-get install -y libpam0g-dev
sudo apt-get install -y libhdf5-dev
sudo apt-get install -y liblz4-dev libhwloc-dev
sudo apt-get install -y libgtk2.0-dev libglib2.0-dev
sudo apt-get install -y librdkafka-dev
sudo apt-get install -y libdbus-1-dev
sudo apt-get install -y check
sudo apt-get install -y liblua5.3-dev
sudo apt-get install -y libreadline-dev
sudo apt-get install -y linux-headers-$(uname -r)
sudo apt-get install -y freeipmi-tools libfreeipmi-dev
sudo apt-get install -y rrdtool librrd-dev
sudo apt-get install -y libjson-c-dev libjansson-dev
sudo apt-get install -y libjwt-dev
sudo apt-get install -y libhttp-parser-dev
sudo apt-get install -y libyaml-dev
sudo apt-get install -y man2html
sudo apt-get install -y mailutils
sudo apt-get install -y libhdf5-dev
sudo apt-get install -y libmysqlclient-dev
sudo apt-get install -y lua5.3 liblua5.3-0 lua-posix lua-filesystem
Install SLURM that was already built during controller setup — it should exist in the shared NFS dir:
SLURM Configuration¶
Make the directories needed:
sudo mkdir -p /etc/slurm /etc/slurm/prolog.d /etc/slurm/epilog.d /var/spool/slurm/ctld /var/spool/slurm/d /var/log/slurm
sudo chown slurm /var/spool/slurm/ctld /var/spool/slurm/d /var/log/slurm
This directory should also be created:
And this one:
slurm.conf¶
slurm.conf is edited once on the controller (see Controller: slurm.conf for the canonical copy) and copied to every node — you just need to make sure this worker's NodeName= entry is in the # COMPUTE NODES section there.
Check this node's detected system info (GPUs aren't included) with:
Once the controller's copy includes this node, copy the file here too:
gres.conf¶
Edit the default gres.conf and add the GPUs that SLURM will manage on this node:
Use this node's own hostname and list one line per actual GPU device it has — e.g. for a node with 2 GPUs:
Then copy it to the system:
cgroup.conf¶
Edit the cgroup file:
Comment out some lines, as follows:
#CgroupAutomount=yes
#CgroupReleaseAgentDir="/etc/slurm/cgroup"
ConstrainCores=yes
ConstrainDevices=yes
ConstrainRAMSpace=yes
#TaskAffinity=ye
Then copy it to the system:
cgroup_allowed_devices_file.conf¶
Copy it as-is:
Configure cgroups (GRUB)¶
Add cgroup and swap to GRUB_CMDLINE_LINUX:
Start SLURM¶
Check status:
Important
Make sure to update the SLURM files on all nodes in your cluster.
Important
Ensure the SLURM server/controller has all nodes defined in /etc/hosts.