ECE Cluster Job Submission Guide
CPU Queues Information
| Settings | CPU Queue Type | |||
|---|---|---|---|---|
| Short | Medium | Long | ||
| Default Time Limit | 6 hr | 12 hr | 1 day | |
| Maximum Time Limit | 12 hr | 36 hr | 3 days | |
| Max Job running per Group | 3 | 2 | 1 | |
| Storage quota | 4 TB per Group | |||
| Default CPU Cores | 20/30 | 40/60 | 80/128 | |
| No. of available GPUs per job | 1 | 1 | 1 | |
| Token Cost | 0.5 per job | 1 per job | 2 per job | |
| Settings | GPU Queue | |||
|---|---|---|---|---|
| gpu-short | gpu-long | |||
| Default Time Limit | 6 hr | 2 day | ||
| Max Job running per Group | 2 | 1 | ||
| Storage quota | 4 TB per Group | |||
| No. of available GPUs per job | 1 | 1 | ||
| Token Cost | 0.5 per job | 1 per job | ||
| Queue | Description |
|---|---|
| Short | MIG-enabled (2x 4= 8 GPUs @ 35 GB each) GPUs (DGX H200) |
| Long | Dedicated GPUs (2x 2H200=4 GPUs @141 GB) |
1. Short Queue
srun --account=it --partition=gpu-short --qos=gpu-short --ntask=1 --pty hostname
srun --account=it --partition=short --qos=short --ntasks=1 --cpus-per-task=1 --nodelist=compute2 hostname
For GPU job
srun --account=it --partition=gpu-short --gres=gpu:1g.35gb:1 --time=00:05:00 --pty nvidia-smi
2. Medium Queue
srun --account=it --partition=medium --qos=medium --ntasks=1 --cpus-per-task=1 hostname
srun --account=it --partition=medium --qos=medium --ntasks=1 --cpus-per-task=1 --nodelist=compute2 hostname
3. Long Queue
srun --account=it --partition=long --qos=long --ntasks=1 --cpus-per-task=1 hostname
srun --account=it --partition=long --qos=long --ntasks=1 --cpus-per-task=1 --nodelist=compute3 hostname
For GPU job
srun --account=it --partition=gpu-long --qos=gpu-long --ntasks=1 --cpus-per-task=1 --gres=gpu:1 nvidia-smi
4. Job script
#!/bin/bash #SBATCH --time=06:00:00 #SBATCH --job-name=Any_name #SBATCH --account=GroupName #SBATCH --partition=gpu-short #SBATCH --qos=gpu-short #SBATCH --ntasks=1 #SBATCH --gres=gpu:gpu:1g.35gb:1 #SBATCH --mem=64G nvidia-smi > nvidialog.txt
Save the script as jobscript.slurm and submit using:
sbatch jobscript.slurm
5. User Access Policy
6. Storage Policy
(b)Maintaining project directories.
(c)Taking backups of important data.
(d)Avoiding unnecessary duplication of datasets.
7. Token Policy
To ensure the equitable utilisation of GPU resources, a token-based scheduling policy shall be implemented.
8. Job Scheduling Policy
Prohibited Activities (Policy Violations)
- Any job found running directly on the Master Node will be terminated immediately. An automated notification will be sent to the user informing them of the policy violation and advising them to submit jobs through the Slurm job scheduler.
- Bypassing the GPU quota system.
- Obtaining unauthorized administrative privileges.
- Cryptocurrency mining.
- Commercial or non-academic workloads.
- Keeping GPUs idle intentionally.
- Sharing login credentials.
- Illegal or prohibited activities.
Monitoring and Enforcement
GPU jobs are logged, monitored and audited. Misuse may result in reduced quota, temporary suspension or permanent loss of access.
Support and Exceptions
- Additional GPU hours may be requested with Faculty approval.
- User tiers are managed by administrators.
- Contact: helpdesk@iiitd.ac.in
Agreement
By using the ECE Cluster you agree to comply with this policy.
By using the ECE Cluster you agree to comply with this policy.