Submitting Jobs on PERUN – Complete Guide¶
This guide explains practical Slurm batch scripts for the PERUN supercomputer.
Table of Contents¶
- Storage Overview
- Basic Job Templates
- Advanced Examples
- Troubleshooting
- Slurm Basics
- Environment Variables
- Best Practices
- FAQ
- Complete Working Example
1. Storage Overview¶
PERUN provides three storage tiers:
- HOME –
/mnt/home– personal space, quota 500 GB, use for code/configs - PROJECT –
/mnt/project– shared team data, long-term storage - SCRATCH –
/mnt/scratch– fast Lustre storage, for running computations only
Manual Data Management Required
There is no automatic staging or syncing of data to/from SCRATCH. You must copy data to SCRATCH yourself before running a job, and copy results back manually once the job finishes.
Important
Data on SCRATCH is automatically deleted 60 days after last access. Always move results you want to keep back to PROJECT or HOME.
Practical rule:
- Prepare input data and code in HOME or PROJECT
- Copy data to SCRATCH before running a computation, e.g.
/mnt/scratch/$USER/myjob/ - Run computations on SCRATCH (fast I/O)
- Copy results you want to keep back to PROJECT
- Don't leave anything important on SCRATCH
2. Basic Job Templates¶
2.1 Single GPU Training¶
#!/bin/bash
#SBATCH --job-name=train_model
#SBATCH --output=%x_%j.out
#SBATCH --error=%x_%j.err
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=48G
#SBATCH --time=24:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
# Copy input data to scratch
rsync -a "$SLURM_SUBMIT_DIR"/data "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
# Your training code
python3 "$SLURM_SUBMIT_DIR"/train.py \
--data data/dataset.csv \
--checkpoint checkpoints/ \
--output results/
# Copy results back to project storage
rsync -a checkpoints/ results/ "$SLURM_SUBMIT_DIR"/output/
2.2 Multi-GPU Training (DDP)¶
#!/bin/bash
#SBATCH --job-name=train_ddp
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_long
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:4
#SBATCH --ntasks=4
#SBATCH --cpus-per-task=8
#SBATCH --mem=128G
#SBATCH --time=48:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
srun python3 -m torch.distributed.run \
--nproc_per_node=4 \
train_ddp.py
rsync -a "$SCRATCH_DIR"/results/ "$SLURM_SUBMIT_DIR"/results/
2.3 CPU-Only Job¶
#!/bin/bash
#SBATCH --job-name=preprocess
#SBATCH --output=%x_%j.out
#SBATCH --partition=cpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/data/raw/ "$SCRATCH_DIR"/data/raw/
cd "$SCRATCH_DIR"
python3 "$SLURM_SUBMIT_DIR"/preprocess_data.py \
--input data/raw/ \
--output data/processed/
rsync -a data/processed/ "$SLURM_SUBMIT_DIR"/data/processed/
3. Advanced Examples¶
3.1 Multi-Node DDP Training¶
#!/bin/bash
#SBATCH --job-name=ddp_multinode
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_long
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --nodes=4
#SBATCH --gres=gpu:4
#SBATCH --ntasks-per-node=4
#SBATCH --cpus-per-task=8
#SBATCH --time=96:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
MASTER_ADDR=$(scontrol show hostnames "$SLURM_NODELIST" | head -n 1)
MASTER_PORT=29500
export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
echo "Training on $SLURM_NNODES nodes, $SLURM_NTASKS GPUs total"
echo "Master: $MASTER_ADDR:$MASTER_PORT"
srun python3 -m torch.distributed.run \
--nnodes=$SLURM_NNODES \
--nproc_per_node=4 \
--rdzv_backend=c10d \
--rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT \
train_ddp.py
rsync -a "$SCRATCH_DIR"/results/ "$SLURM_SUBMIT_DIR"/results/
3.2 Hyperparameter Search (Job Array)¶
#!/bin/bash
#SBATCH --job-name=hparam_search
#SBATCH --output=logs/search_%A_%a.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --array=0-99%10
#SBATCH --time=12:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_${SLURM_ARRAY_JOB_ID}_${SLURM_ARRAY_TASK_ID}"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
SEED=$SLURM_ARRAY_TASK_ID
python3 train.py \
--seed $SEED \
--lr $(python3 -c "print(0.0001 * (1.5 ** $SEED))") \
--output results/seed_${SEED}/
rsync -a results/seed_${SEED}/ "$SLURM_SUBMIT_DIR"/results/seed_${SEED}/
3.3 Checkpoint Resume¶
#!/bin/bash
#SBATCH --job-name=resume_training
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"/checkpoints
cd "$SCRATCH_DIR"
# Copy previous checkpoint to scratch
if [ -f "$SLURM_SUBMIT_DIR/results_job_3500/checkpoints/best_model.pt" ]; then
cp "$SLURM_SUBMIT_DIR/results_job_3500/checkpoints/best_model.pt" checkpoints/
echo "Resumed from previous checkpoint"
fi
python3 "$SLURM_SUBMIT_DIR"/train.py --resume checkpoints/best_model.pt
rsync -a checkpoints/ "$SLURM_SUBMIT_DIR"/checkpoints/
3.4 Periodic Checkpoint Backup During Long Runs¶
#!/bin/bash
#SBATCH --job-name=custom_sync
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --time=24:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
echo "Running in: $SCRATCH_DIR"
python3 train.py &
TRAIN_PID=$!
# Periodically back up critical checkpoints while training runs
while kill -0 $TRAIN_PID 2>/dev/null; do
sleep 3600
if [ -f checkpoints/latest.pt ]; then
rsync -a checkpoints/latest.pt "$SLURM_SUBMIT_DIR"/backup_checkpoints/
echo "$(date): Synced intermediate checkpoint"
fi
done
wait $TRAIN_PID
# Final sync
rsync -a checkpoints/ results/ "$SLURM_SUBMIT_DIR"/output/
4. Troubleshooting¶
4.1 Job Failed Immediately¶
Symptom
Job exits with "Permission denied" or "No such file"
Solution
Make sure /mnt/scratch/$USER exists and is writable. Create the job's scratch subdirectory yourself with mkdir -p at the start of your script before writing to it.
4.2 Output Files Not Found After Job¶
Symptom
Can't find results after job completes
Solution
Results are only where you explicitly copied them with rsync/cp. Check your script's final sync step, and check /mnt/scratch/$USER/job_XXXXX/ if the job failed before that step ran.
4.3 Large Checkpoints Missing¶
Symptom
Some checkpoints are missing from your output directory
Possible causes:
- Job hit the time limit before your final
rsynccommand ran - Disk quota exceeded on PROJECT/HOME
- Copy was too large and did not finish within the remaining walltime
Solution
# Manually recover from scratch (if the data still exists)
rsync -avP /mnt/scratch/$USER/job_XXXXX/ ~/recovered_results/
Consider adding periodic intermediate syncs for long-running jobs (see 3.4).
4.4 Job Slower Than Expected¶
Symptom
Training is slow despite using scratch
Diagnostics:
# Check if actually running in scratch
squeue -j $JOBID -o "%i %Z" # WorkDir should be /mnt/scratch/...
# Check if data is actually in scratch
ls -lh /mnt/scratch/$USER/job_$JOBID/data/
4.5 Monitoring Live Progress¶
# Tail output file
tail -f %x_%j.out
# Check job's current working directory contents
ssh <node> 'ls -lh /mnt/scratch/$USER/job_$JOBID/'
5. Slurm Basics (Cheat Sheet)¶
Essential Commands¶
# Submit job
sbatch job.sh
# Check queue
squeue -u $USER
# Job details
scontrol show job XXXXX
# Cancel job
scancel XXXXX
# Job history
sacct -j XXXXX --format=JobID,State,Elapsed,MaxRSS,ReqMem
Common SBATCH Directives¶
#SBATCH --job-name=my_job # Job name
#SBATCH --output=%x_%j.out # Output file (%x=name, %j=jobid)
#SBATCH --error=%x_%j.err # Error file
#SBATCH --partition=gpu_short # Queue: cpu_short, cpu_long, gpu_short, gpu_long
#SBATCH --account=perun2501234 # Project account
#SBATCH --qos=perun2501234 # Quality of Service
#SBATCH --gres=gpu:2 # Request 2 GPUs (gpu_short/gpu_long only)
#SBATCH --nodes=1 # Number of nodes
#SBATCH --ntasks=1 # Number of processes
#SBATCH --cpus-per-task=8 # CPUs per process
#SBATCH --mem=64G # Memory
#SBATCH --time=24:00:00 # Time limit (HH:MM:SS)
Choosing a Partition
| Partition | Nodes | Time Limit | Use For |
|---|---|---|---|
cpu_short |
cn01–cn32 | 2 days | Short CPU jobs |
cpu_long |
cn01–cn32 | 4 days | Long CPU jobs |
gpu_short |
gpu01–gpu26 | 2 days | Short GPU/AI jobs |
gpu_long |
gpu01–gpu26 | 4 days | Long GPU/AI jobs |
All CPU nodes (cn01–cn32) and all GPU nodes (gpu01–gpu26) are available in both the short and long partition — choose based on required walltime, not node availability.
Account and QoS
Replace perun2501234 with your actual project number. Check your available accounts and QoS limits with:
Output Filename Patterns¶
| Pattern | Meaning | Example |
|---|---|---|
%x |
Job name | train_model |
%j |
Job ID | 3928 |
%A |
Array job ID | 4000 |
%a |
Array task ID | 5 |
%N |
Node name | gpu01 |
6. Environment Variables¶
Available in Jobs¶
$SLURM_JOB_ID # Job ID
$SLURM_JOB_NAME # Job name
$SLURM_SUBMIT_DIR # Directory where sbatch was called
$SLURM_CPUS_PER_TASK # CPUs requested
$SLURM_NTASKS # Total tasks
$SLURM_NNODES # Number of nodes
$SLURM_NODELIST # List of nodes
$CUDA_VISIBLE_DEVICES # Visible GPUs (set by Slurm)
Define your own scratch path variables at the top of your script as needed, e.g.:
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
DATA_DIR="$SCRATCH_DIR/data"
RESULTS_DIR="$SCRATCH_DIR/results"
7. Best Practices¶
DO¶
- Always use scratch for I/O-heavy computation – much faster than PROJECT/HOME
- Copy input data to scratch before your job starts working with it
- Copy results back to PROJECT/HOME before the job ends
- Use
%x_%j.outfor output files – easier to track - Always set
--accountand--qos - Request appropriate resources – don't over-request
- Set realistic time limits
- Choose the right partition –
_shortfor jobs under 2 days,_longfor up to 4 days - Test with short jobs first
- Keep code in Git
DON'T¶
- Don't write large files directly to HOME – use scratch
- Don't assume scratch is populated automatically – copy data yourself
- Don't leave results only on scratch – copy them back, or they will be deleted after 60 days of inactivity
- Don't forget your final
rsync/cpstep – if the job is killed mid-run, unsaved results are lost - Don't request 8 GPUs if you use 1
- Don't use
--nodelistin production - Don't run interactive jobs 24/7 – use batch jobs
- Don't use
gpu_longfor short jobs – usegpu_shortto get scheduled faster
8. FAQ¶
Do I need to change my Python code?
No, but make sure your script's paths point to the scratch directory you created, not the submit directory, when reading/writing large data.
What if my dataset is 5TB?
Don't copy it if avoidable. Keep large datasets in /datasets/ and reference them directly, or copy only the subset you need.
Can I submit from any directory?
Yes, but you are responsible for copying any needed input data to scratch yourself.
How long does copying take?
Roughly 1–2 seconds per GB with rsync, depending on load.
What if my job is killed mid-training?
Any results not yet copied back to PROJECT/HOME are lost. Use periodic intermediate syncs for long jobs.
How do I find my account and QoS?
Run: sacctmgr show user $USER withassoc format=account,qos
Which partition should I use?
Use gpu_short / cpu_short for jobs under 2 days. Use gpu_long / cpu_long for jobs up to 4 days.
9. Complete Working Example¶
#!/bin/bash
################################################################################
# BERT-Large Fine-tuning on SKQuAD Dataset
# Expected runtime: ~2 hours on 2x H200 GPUs
################################################################################
#SBATCH --job-name=skquad_bert
#SBATCH --output=%x_%j.out
#SBATCH --error=%x_%j.err
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:2
#SBATCH --cpus-per-task=16
#SBATCH --mem=96G
#SBATCH --time=04:00:00
SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
# Copy input data to scratch
rsync -a "$SLURM_SUBMIT_DIR"/data "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"
export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
export CUDA_LAUNCH_BLOCKING=0
echo "═══════════════════════════════════════════════════════════"
echo "Job ID: $SLURM_JOB_ID"
echo "Job Name: $SLURM_JOB_NAME"
echo "Node: $(hostname)"
echo "GPUs: $CUDA_VISIBLE_DEVICES"
echo "Scratch: $SCRATCH_DIR"
echo "═══════════════════════════════════════════════════════════"
echo
nvidia-smi --query-gpu=name,memory.total --format=csv
echo
python3 "$SLURM_SUBMIT_DIR"/train_bert.py \
--model_name bert-large-uncased \
--dataset skquad \
--output_dir checkpoints/ \
--num_train_epochs 3 \
--per_device_train_batch_size 16 \
--learning_rate 3e-5 \
--warmup_steps 500 \
--save_steps 1000 \
--logging_steps 100 \
--fp16
echo
echo "═══════════════════════════════════════════════════════════"
echo "Training complete! Copying results back to project storage..."
echo "═══════════════════════════════════════════════════════════"
# Copy results back
rsync -a checkpoints/ "$SLURM_SUBMIT_DIR"/results_job_$SLURM_JOB_ID/checkpoints/
Submit the Job
Monitor Progress