Skip to content

Deadline Priority QoS — Design & Implementation (v2, tested)

Goal

Give any advisor group a priority boost when they have an upcoming deadline, without preemption. A group submits with --qos=deadline_<name> to jump ahead of normal jobs. Each group has its own GPU-hour budget; once spent, jobs under that QoS are rejected and the group falls back to normal scheduling. Normal jobs (no --qos) are never affected.

Preemption was ruled out: with 17 groups (many on formerly personally-owned workstations), forcibly killing running jobs would erode trust in the shared cluster and cause people to avoid submitting altogether.

Key Correction from v1 — Read This First

A single shared deadline QoS does not give per-group budgets — tested directly on the cluster: two different accounts assigned the same QoS drained one shared counter. There's no sacctmgr syntax that scopes GrpTRESMins to a specific (account, QoS) pair; set GrpTRESMins=... qos=deadline on an account applies the limit to the whole account, blocking normal jobs too.

The only isolation that works: one QoS per group, each with its own GrpTRESMins.

Naming Convention

Accounts are named grp_<name>. The matching deadline QoS is named deadline_<name> (drop grp_, add deadline_).

Example: account grp_mohammedsinan → QoS deadline_mohammedsinan.

See Account Structure for how accounts/groups are organized more generally.

Budget Math

GrpTRESMins=gres/gpu=X is in GPU-minutes, calculated as:

GPU-minutes = (number of GPUs) x (minutes)

Conversions used so far:

Duration Minutes
2 days 2,880
6 days 8,640
30 days (1 month) 43,200
35 days 50,400

Examples:

  • 2 GPUs × 2 days → 2 × 2,880 = 5,760
  • 2 GPUs × 6 days → 2 × 8,640 = 17,280
  • 3 GPUs × 1 month → 3 × 43,200 = 129,600
  • 4 GPUs × 35 days → 4 × 50,400 = 201,600

Pick whatever combination of (GPUs × days) represents a group's realistic deadline crunch, multiply, and that's their GrpTRESMins value. Budget resets yearly (see below), so this is effectively "total allowance per year," not per-deadline — if a group needs multiple deadline pushes a year, size the number accordingly (e.g. three separate 2-GPU/2-day pushes → 5,760 × 3).

1. Create the Group's QoS

sudo sacctmgr -i add qos deadline_mohammedsinan set Priority=1000000 Flags=DenyOnLimit,NoDecay GrpTRESMins=gres/gpu=5760
  • Priority=1000000 — only needs to be the highest Priority value among all QoS on the cluster; the actual number is normalized, not absolute.
  • Flags=DenyOnLimit — reject the job at submission once budget is exhausted, with a clear error, instead of leaving it pending silently.
  • Flags=NoDecay — this QoS's own usage/GrpTRESMins will NOT be auto-decayed or reset by PriorityDecayHalfLife / PriorityUsageResetPeriod in slurm.conf. Confirmed from SchedMD docs and a slurm-users mailing list thread: fairshare for the account is untouched by this flag — only this QoS's own budget tracking is exempted from the global decay/reset schedule. This means yearly reset must be done manually (see below) — nothing resets it automatically once NoDecay is set.

To modify an existing QoS instead of creating (e.g. adjusting budget later):

sudo sacctmgr -i modify qos deadline_mohammedsinan set GrpTRESMins=gres/gpu=201600

Note

sacctmgr add qos silently does nothing — "Nothing new added" — if the QoS already exists; use modify to change settings on one that's already there.

2. Confirm PriorityWeightQOS Is Actually Nonzero

The QoS Priority= value only matters if slurm.conf weights it into the job priority formula:

scontrol show config | grep PriorityWeight

On this cluster: PriorityWeightAge=4000, PriorityWeightFairShare=10000, PriorityWeightQOS=0 — meaning QoS currently contributes nothing to priority regardless of the QoS's own Priority= value. To make deadline jobs reliably outrank normal jobs, PriorityWeightQOS needs to exceed the sum of the other active weights (here, 4000 + 10000 = 14000):

PriorityWeightQOS=20000

Set in /etc/slurm/slurm.conf, then:

sudo scontrol reconfigure

Verify with sprio -l — compare a deadline-QoS job's QOS column against a normal job's.

3. Grant the QoS to the Account

sudo sacctmgr -i modify account grp_mohammedsinan set qos+=deadline_mohammedsinan

Always use qos+=, never plain qos=

qos+= adds to the existing list — never use plain qos= to add a QoS, since that replaces the whole list. This was hit during testing: an account with no explicit QoS list (just inheriting normal implicitly) had normal silently wiped out and replaced with only deadline, which caused a bug where even plain jobs with no --qos flag got billed against the deadline budget and got rejected. If this ever happens, recover with:

sudo sacctmgr -i modify account grp_mohammedsinan set qos+=normal
sudo sacctmgr -i modify account grp_mohammedsinan set DefaultQOS=normal

4. Confirm Both QoS Are Present, and the Default Is Correct

sacctmgr show assoc account=grp_mohammedsinan format=Account,User,QOS,DefaultQOS tree

Expect to see normal,deadline_mohammedsinan and DefaultQOS=normal.

5. Confirm the Budget Landed Correctly

scontrol show assoc_mgr | grep -A16 "QOS=deadline_mohammedsinan"

Look for GrpTRESMins=...gres/gpu=5760(0) — format is limit(used).

Submitting Jobs

If grp_mohammedsinan is the user's DefaultAccount, --account can be omitted; --qos must always be explicit (default QoS is normal, so nothing draws from the deadline budget unless requested):

sbatch --qos=deadline_mohammedsinan --gres=gpu:2 my_job.sh

Confirm a user's actual default account before assuming it can be omitted:

sacctmgr show user withassoc format=User,Account,DefaultAccount where account=grp_mohammedsinan

Checking Remaining Budget

scontrol show assoc_mgr | grep -A16 "QOS=deadline_mohammedsinan"

GrpTRESMins=...gres/gpu=5760(N) → remaining = 5760 - N.

Note

This counter is cumulative total consumed, not concurrent usage — it does not go back down when a job finishes. It only decreases via an explicit reset (see below).

A Subtlety Observed During Testing: Admission Uses Worst-Case Time, Not Actual Runtime

This cluster's job_submit.lua force-sets interactive (salloc/srun) jobs' time limit to 120 minutes regardless of requested --time. During testing, this caused a submission-time rejection against a small test budget (60 GPU-minutes) even for a job that only ran 90 seconds — because DenyOnLimit projects worst-case usage (GPUs × time-limit) before allowing submission. Once admitted, however, the usage actually recorded reflected real elapsed time (~3 GPU-minutes for a 90-second job), not the 120-minute reservation. Net effect: real per-group budgets (thousands of GPU-minutes) won't be affected by this, but it's worth knowing a job can be denied at submission for looking expensive on paper even if it would only run briefly.

Yearly / Manual Budget Reset

Because NoDecay is set, nothing resets automatically — this must be done explicitly, per QoS:

sudo sacctmgr -i modify qos deadline_mohammedsinan set RawUsage=0

Reset all 17 groups in one pass:

for acct in $(sacctmgr show account format=Account%30 -n | tr -d ' ' | grep 'grp_' | sort -u); do
  name=${acct#grp_}
  qos="deadline_${name}"
  echo "Resetting $qos"
  sudo sacctmgr -i modify qos "$qos" set RawUsage=0
done

Verify a reset actually took effect:

scontrol show assoc_mgr | grep -A16 "QOS=deadline_mohammedsinan"

UsageRaw should read 0.000000 and GrpTRESMins=...gres/gpu=129600(0).

Automate with cron, running Jan 1 each year. Save the loop above as a script, e.g. /SLURM/home/slurm-admin/scripts/deadline_reset.sh, then:

sudo crontab -e

Add:

0 0 1 1 * /bin/bash /SLURM/home/slurm-admin/scripts/deadline_reset.sh >> /var/log/deadline_reset.log 2>&1

Not yet verified end-to-end

Test the single-account reset above, confirm via scontrol show assoc_mgr that UsageRaw and GrpTRESMins usage actually zero out, before relying on the cron job for all 17 groups untested.

Preventing Budget Sharing / Leakage

Per-group QoS already prevents one group from touching another's pool directly. Remaining risk vectors:

  • Users in multiple accounts can submit under a different account's --account= and --qos=deadline_<other> to draw on that group's budget. Audit: sacctmgr show user withassoc format=User,Account,DefaultAccount — flag anyone under more than one account.
  • Account "lending" — a PI adds an outside user to their account. This is an admin-only action (sacctmgr add user ... account=...), so it's auditable; restrict who can do this.
  • Monthly usage report as a detection net:

    sreport cluster AccountUtilizationByUser qos=deadline_mohammedsinan start=2026-01-01 -T gres/gpu
    

    Repeat per group, or script across all 17.