Teaching GPU cluster Slurm partitions

In Slurm, a partition can be considered a resource abstraction that groups a set of compute nodes into a single resource pool. A partition's configuration defines resource limits, scheduling policies, and access controls for the nodes within that pool.

When you submit a job, Slurm allocates resources from the selected partition based on your job's resource requirements, the resources currently available, and any restrictions configured for that partition.

MFCF may adjust partition configurations over time based on observed usage patterns and the availability of computing resources.

The teaching GPU cluster currently has one partition, gpu-gen, which provides access to all GPU compute nodes in the cluster.

  • gpu-gen total available resources:
    Partition name gpu-gen
    Total available memory 2.3 TB
    Max Cores 288 cores
    Threads per core 2 Threads
    GPU devices 8 GTX1080ti devices, 5 RTX 6000 Ada, 1 RTX PRO 6000 and 3 L40S Ada
    GPU memory per device 12 GB to 48 GB depending on GPU
    Compute Nodes

    gpu-pt1-03,
    gpu-pt1-04,
    gpu-pt1-05,
    gpu-pt1-06

    gpu-gen partition specifications
  • gpu-gen per user resource limits:
    Max runtime (h) 12 hour
    Max Nodes 1 Node
    gpu-gen partition per user limits