Deadline Priority QoS — Design & Implementation (v2, tested)¶
Goal¶
Give any advisor group a priority boost when they have an upcoming deadline, without preemption. A group submits with --qos=deadline_<name> to jump ahead of normal jobs. Each group has its own GPU-hour budget; once spent, jobs under that QoS are rejected and the group falls back to normal scheduling. Normal jobs (no --qos) are never affected.
Preemption was ruled out: with 17 groups (many on formerly personally-owned workstations), forcibly killing running jobs would erode trust in the shared cluster and cause people to avoid submitting altogether.
Key Correction from v1 — Read This First¶
A single shared deadline QoS does not give per-group budgets — tested directly on the cluster: two different accounts assigned the same QoS drained one shared counter. There's no sacctmgr syntax that scopes GrpTRESMins to a specific (account, QoS) pair; set GrpTRESMins=... qos=deadline on an account applies the limit to the whole account, blocking normal jobs too.
The only isolation that works: one QoS per group, each with its own GrpTRESMins.
Naming Convention¶
Accounts are named grp_<name>. The matching deadline QoS is named deadline_<name> (drop grp_, add deadline_).
Example: account grp_mohammedsinan → QoS deadline_mohammedsinan.
See Account Structure for how accounts/groups are organized more generally.
Budget Math¶
GrpTRESMins=gres/gpu=X is in GPU-minutes, calculated as:
Conversions used so far:
| Duration | Minutes |
|---|---|
| 2 days | 2,880 |
| 6 days | 8,640 |
| 30 days (1 month) | 43,200 |
| 35 days | 50,400 |
Examples:
- 2 GPUs × 2 days → 2 × 2,880 = 5,760
- 2 GPUs × 6 days → 2 × 8,640 = 17,280
- 3 GPUs × 1 month → 3 × 43,200 = 129,600
- 4 GPUs × 35 days → 4 × 50,400 = 201,600
Pick whatever combination of (GPUs × days) represents a group's realistic deadline crunch, multiply, and that's their GrpTRESMins value. Budget resets yearly (see below), so this is effectively "total allowance per year," not per-deadline — if a group needs multiple deadline pushes a year, size the number accordingly (e.g. three separate 2-GPU/2-day pushes → 5,760 × 3).
1. Create the Group's QoS¶
sudo sacctmgr -i add qos deadline_mohammedsinan set Priority=1000000 Flags=DenyOnLimit,NoDecay GrpTRESMins=gres/gpu=5760
Priority=1000000— only needs to be the highestPriorityvalue among all QoS on the cluster; the actual number is normalized, not absolute.Flags=DenyOnLimit— reject the job at submission once budget is exhausted, with a clear error, instead of leaving it pending silently.Flags=NoDecay— this QoS's own usage/GrpTRESMinswill NOT be auto-decayed or reset byPriorityDecayHalfLife/PriorityUsageResetPeriodin slurm.conf. Confirmed from SchedMD docs and a slurm-users mailing list thread: fairshare for the account is untouched by this flag — only this QoS's own budget tracking is exempted from the global decay/reset schedule. This means yearly reset must be done manually (see below) — nothing resets it automatically onceNoDecayis set.
To modify an existing QoS instead of creating (e.g. adjusting budget later):
Note
sacctmgr add qos silently does nothing — "Nothing new added" — if the QoS already exists; use modify to change settings on one that's already there.
2. Confirm PriorityWeightQOS Is Actually Nonzero¶
The QoS Priority= value only matters if slurm.conf weights it into the job priority formula:
On this cluster: PriorityWeightAge=4000, PriorityWeightFairShare=10000, PriorityWeightQOS=0 — meaning QoS currently contributes nothing to priority regardless of the QoS's own Priority= value. To make deadline jobs reliably outrank normal jobs, PriorityWeightQOS needs to exceed the sum of the other active weights (here, 4000 + 10000 = 14000):
Set in /etc/slurm/slurm.conf, then:
Verify with sprio -l — compare a deadline-QoS job's QOS column against a normal job's.
3. Grant the QoS to the Account¶
Always use qos+=, never plain qos=
qos+= adds to the existing list — never use plain qos= to add a QoS, since that replaces the whole list. This was hit during testing: an account with no explicit QoS list (just inheriting normal implicitly) had normal silently wiped out and replaced with only deadline, which caused a bug where even plain jobs with no --qos flag got billed against the deadline budget and got rejected. If this ever happens, recover with:
4. Confirm Both QoS Are Present, and the Default Is Correct¶
Expect to see normal,deadline_mohammedsinan and DefaultQOS=normal.
5. Confirm the Budget Landed Correctly¶
Look for GrpTRESMins=...gres/gpu=5760(0) — format is limit(used).
Submitting Jobs¶
If grp_mohammedsinan is the user's DefaultAccount, --account can be omitted; --qos must always be explicit (default QoS is normal, so nothing draws from the deadline budget unless requested):
Confirm a user's actual default account before assuming it can be omitted:
Checking Remaining Budget¶
GrpTRESMins=...gres/gpu=5760(N) → remaining = 5760 - N.
Note
This counter is cumulative total consumed, not concurrent usage — it does not go back down when a job finishes. It only decreases via an explicit reset (see below).
A Subtlety Observed During Testing: Admission Uses Worst-Case Time, Not Actual Runtime¶
This cluster's job_submit.lua force-sets interactive (salloc/srun) jobs' time limit to 120 minutes regardless of requested --time. During testing, this caused a submission-time rejection against a small test budget (60 GPU-minutes) even for a job that only ran 90 seconds — because DenyOnLimit projects worst-case usage (GPUs × time-limit) before allowing submission. Once admitted, however, the usage actually recorded reflected real elapsed time (~3 GPU-minutes for a 90-second job), not the 120-minute reservation. Net effect: real per-group budgets (thousands of GPU-minutes) won't be affected by this, but it's worth knowing a job can be denied at submission for looking expensive on paper even if it would only run briefly.
Yearly / Manual Budget Reset¶
Because NoDecay is set, nothing resets automatically — this must be done explicitly, per QoS:
Reset all 17 groups in one pass:
for acct in $(sacctmgr show account format=Account%30 -n | tr -d ' ' | grep 'grp_' | sort -u); do
name=${acct#grp_}
qos="deadline_${name}"
echo "Resetting $qos"
sudo sacctmgr -i modify qos "$qos" set RawUsage=0
done
Verify a reset actually took effect:
UsageRaw should read 0.000000 and GrpTRESMins=...gres/gpu=129600(0).
Automate with cron, running Jan 1 each year. Save the loop above as a script, e.g. /SLURM/home/slurm-admin/scripts/deadline_reset.sh, then:
Add:
0 0 1 1 * /bin/bash /SLURM/home/slurm-admin/scripts/deadline_reset.sh >> /var/log/deadline_reset.log 2>&1
Not yet verified end-to-end
Test the single-account reset above, confirm via scontrol show assoc_mgr that UsageRaw and GrpTRESMins usage actually zero out, before relying on the cron job for all 17 groups untested.
Preventing Budget Sharing / Leakage¶
Per-group QoS already prevents one group from touching another's pool directly. Remaining risk vectors:
- Users in multiple accounts can submit under a different account's
--account=and--qos=deadline_<other>to draw on that group's budget. Audit:sacctmgr show user withassoc format=User,Account,DefaultAccount— flag anyone under more than one account. - Account "lending" — a PI adds an outside user to their account. This is an admin-only action (
sacctmgr add user ... account=...), so it's auditable; restrict who can do this. -
Monthly usage report as a detection net:
Repeat per group, or script across all 17.