Skip to content

Submitting Jobs on PERUN – Complete Guide

This guide explains practical Slurm batch scripts for the PERUN supercomputer.


Table of Contents

  1. Storage Overview
  2. Basic Job Templates
  3. Advanced Examples
  4. Troubleshooting
  5. Slurm Basics
  6. Environment Variables
  7. Best Practices
  8. FAQ
  9. Complete Working Example

1. Storage Overview

PERUN provides three storage tiers:

  • HOME – /mnt/home – personal space, quota 500 GB, use for code/configs
  • PROJECT – /mnt/project – shared team data, long-term storage
  • SCRATCH – /mnt/scratch – fast Lustre storage, for running computations only

Manual Data Management Required

There is no automatic staging or syncing of data to/from SCRATCH. You must copy data to SCRATCH yourself before running a job, and copy results back manually once the job finishes.

Important

Data on SCRATCH is automatically deleted 60 days after last access. Always move results you want to keep back to PROJECT or HOME.

Practical rule:

  1. Prepare input data and code in HOME or PROJECT
  2. Copy data to SCRATCH before running a computation, e.g. /mnt/scratch/$USER/myjob/
  3. Run computations on SCRATCH (fast I/O)
  4. Copy results you want to keep back to PROJECT
  5. Don't leave anything important on SCRATCH

2. Basic Job Templates

2.1 Single GPU Training

#!/bin/bash
#SBATCH --job-name=train_model
#SBATCH --output=%x_%j.out
#SBATCH --error=%x_%j.err
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=48G
#SBATCH --time=24:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"

# Copy input data to scratch
rsync -a "$SLURM_SUBMIT_DIR"/data "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

# Your training code
python3 "$SLURM_SUBMIT_DIR"/train.py \
    --data data/dataset.csv \
    --checkpoint checkpoints/ \
    --output results/

# Copy results back to project storage
rsync -a checkpoints/ results/ "$SLURM_SUBMIT_DIR"/output/

2.2 Multi-GPU Training (DDP)

#!/bin/bash
#SBATCH --job-name=train_ddp
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_long
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:4
#SBATCH --ntasks=4
#SBATCH --cpus-per-task=8
#SBATCH --mem=128G
#SBATCH --time=48:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

srun python3 -m torch.distributed.run \
    --nproc_per_node=4 \
    train_ddp.py

rsync -a "$SCRATCH_DIR"/results/ "$SLURM_SUBMIT_DIR"/results/

2.3 CPU-Only Job

#!/bin/bash
#SBATCH --job-name=preprocess
#SBATCH --output=%x_%j.out
#SBATCH --partition=cpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/data/raw/ "$SCRATCH_DIR"/data/raw/
cd "$SCRATCH_DIR"

python3 "$SLURM_SUBMIT_DIR"/preprocess_data.py \
    --input data/raw/ \
    --output data/processed/

rsync -a data/processed/ "$SLURM_SUBMIT_DIR"/data/processed/

3. Advanced Examples

3.1 Multi-Node DDP Training

#!/bin/bash
#SBATCH --job-name=ddp_multinode
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_long
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --nodes=4
#SBATCH --gres=gpu:4
#SBATCH --ntasks-per-node=4
#SBATCH --cpus-per-task=8
#SBATCH --time=96:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

MASTER_ADDR=$(scontrol show hostnames "$SLURM_NODELIST" | head -n 1)
MASTER_PORT=29500

export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK

echo "Training on $SLURM_NNODES nodes, $SLURM_NTASKS GPUs total"
echo "Master: $MASTER_ADDR:$MASTER_PORT"

srun python3 -m torch.distributed.run \
    --nnodes=$SLURM_NNODES \
    --nproc_per_node=4 \
    --rdzv_backend=c10d \
    --rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT \
    train_ddp.py

rsync -a "$SCRATCH_DIR"/results/ "$SLURM_SUBMIT_DIR"/results/

3.2 Hyperparameter Search (Job Array)

#!/bin/bash
#SBATCH --job-name=hparam_search
#SBATCH --output=logs/search_%A_%a.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --array=0-99%10
#SBATCH --time=12:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_${SLURM_ARRAY_JOB_ID}_${SLURM_ARRAY_TASK_ID}"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

SEED=$SLURM_ARRAY_TASK_ID

python3 train.py \
    --seed $SEED \
    --lr $(python3 -c "print(0.0001 * (1.5 ** $SEED))") \
    --output results/seed_${SEED}/

rsync -a results/seed_${SEED}/ "$SLURM_SUBMIT_DIR"/results/seed_${SEED}/

3.3 Checkpoint Resume

#!/bin/bash
#SBATCH --job-name=resume_training
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"/checkpoints
cd "$SCRATCH_DIR"

# Copy previous checkpoint to scratch
if [ -f "$SLURM_SUBMIT_DIR/results_job_3500/checkpoints/best_model.pt" ]; then
    cp "$SLURM_SUBMIT_DIR/results_job_3500/checkpoints/best_model.pt" checkpoints/
    echo "Resumed from previous checkpoint"
fi

python3 "$SLURM_SUBMIT_DIR"/train.py --resume checkpoints/best_model.pt

rsync -a checkpoints/ "$SLURM_SUBMIT_DIR"/checkpoints/

3.4 Periodic Checkpoint Backup During Long Runs

#!/bin/bash
#SBATCH --job-name=custom_sync
#SBATCH --output=%x_%j.out
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:1
#SBATCH --time=24:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"
rsync -a "$SLURM_SUBMIT_DIR"/ "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

echo "Running in: $SCRATCH_DIR"

python3 train.py &
TRAIN_PID=$!

# Periodically back up critical checkpoints while training runs
while kill -0 $TRAIN_PID 2>/dev/null; do
    sleep 3600
    if [ -f checkpoints/latest.pt ]; then
        rsync -a checkpoints/latest.pt "$SLURM_SUBMIT_DIR"/backup_checkpoints/
        echo "$(date): Synced intermediate checkpoint"
    fi
done

wait $TRAIN_PID

# Final sync
rsync -a checkpoints/ results/ "$SLURM_SUBMIT_DIR"/output/

4. Troubleshooting

4.1 Job Failed Immediately

Symptom

Job exits with "Permission denied" or "No such file"

Solution

Make sure /mnt/scratch/$USER exists and is writable. Create the job's scratch subdirectory yourself with mkdir -p at the start of your script before writing to it.

4.2 Output Files Not Found After Job

Symptom

Can't find results after job completes

Solution

Results are only where you explicitly copied them with rsync/cp. Check your script's final sync step, and check /mnt/scratch/$USER/job_XXXXX/ if the job failed before that step ran.

4.3 Large Checkpoints Missing

Symptom

Some checkpoints are missing from your output directory

Possible causes:

  • Job hit the time limit before your final rsync command ran
  • Disk quota exceeded on PROJECT/HOME
  • Copy was too large and did not finish within the remaining walltime

Solution

# Manually recover from scratch (if the data still exists)
rsync -avP /mnt/scratch/$USER/job_XXXXX/ ~/recovered_results/

Consider adding periodic intermediate syncs for long-running jobs (see 3.4).

4.4 Job Slower Than Expected

Symptom

Training is slow despite using scratch

Diagnostics:

# Check if actually running in scratch
squeue -j $JOBID -o "%i %Z"  # WorkDir should be /mnt/scratch/...

# Check if data is actually in scratch
ls -lh /mnt/scratch/$USER/job_$JOBID/data/

4.5 Monitoring Live Progress

# Tail output file
tail -f %x_%j.out

# Check job's current working directory contents
ssh <node> 'ls -lh /mnt/scratch/$USER/job_$JOBID/'

5. Slurm Basics (Cheat Sheet)

Essential Commands

# Submit job
sbatch job.sh

# Check queue
squeue -u $USER

# Job details
scontrol show job XXXXX

# Cancel job
scancel XXXXX

# Job history
sacct -j XXXXX --format=JobID,State,Elapsed,MaxRSS,ReqMem

Common SBATCH Directives

#SBATCH --job-name=my_job         # Job name
#SBATCH --output=%x_%j.out        # Output file (%x=name, %j=jobid)
#SBATCH --error=%x_%j.err         # Error file
#SBATCH --partition=gpu_short     # Queue: cpu_short, cpu_long, gpu_short, gpu_long
#SBATCH --account=perun2501234    # Project account
#SBATCH --qos=perun2501234        # Quality of Service
#SBATCH --gres=gpu:2              # Request 2 GPUs (gpu_short/gpu_long only)
#SBATCH --nodes=1                 # Number of nodes
#SBATCH --ntasks=1                # Number of processes
#SBATCH --cpus-per-task=8         # CPUs per process
#SBATCH --mem=64G                 # Memory
#SBATCH --time=24:00:00           # Time limit (HH:MM:SS)

Choosing a Partition

Partition Nodes Time Limit Use For
cpu_short cn01–cn32 2 days Short CPU jobs
cpu_long cn01–cn32 4 days Long CPU jobs
gpu_short gpu01–gpu26 2 days Short GPU/AI jobs
gpu_long gpu01–gpu26 4 days Long GPU/AI jobs

All CPU nodes (cn01–cn32) and all GPU nodes (gpu01–gpu26) are available in both the short and long partition — choose based on required walltime, not node availability.

Account and QoS

Replace perun2501234 with your actual project number. Check your available accounts and QoS limits with:

sacctmgr show user $USER withassoc format=account,qos

Output Filename Patterns

Pattern Meaning Example
%x Job name train_model
%j Job ID 3928
%A Array job ID 4000
%a Array task ID 5
%N Node name gpu01

6. Environment Variables

Available in Jobs

$SLURM_JOB_ID              # Job ID
$SLURM_JOB_NAME            # Job name
$SLURM_SUBMIT_DIR          # Directory where sbatch was called
$SLURM_CPUS_PER_TASK       # CPUs requested
$SLURM_NTASKS              # Total tasks
$SLURM_NNODES              # Number of nodes
$SLURM_NODELIST            # List of nodes
$CUDA_VISIBLE_DEVICES      # Visible GPUs (set by Slurm)

Define your own scratch path variables at the top of your script as needed, e.g.:

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
DATA_DIR="$SCRATCH_DIR/data"
RESULTS_DIR="$SCRATCH_DIR/results"

7. Best Practices

DO

  • Always use scratch for I/O-heavy computation – much faster than PROJECT/HOME
  • Copy input data to scratch before your job starts working with it
  • Copy results back to PROJECT/HOME before the job ends
  • Use %x_%j.out for output files – easier to track
  • Always set --account and --qos
  • Request appropriate resources – don't over-request
  • Set realistic time limits
  • Choose the right partition – _short for jobs under 2 days, _long for up to 4 days
  • Test with short jobs first
  • Keep code in Git

DON'T

  • Don't write large files directly to HOME – use scratch
  • Don't assume scratch is populated automatically – copy data yourself
  • Don't leave results only on scratch – copy them back, or they will be deleted after 60 days of inactivity
  • Don't forget your final rsync/cp step – if the job is killed mid-run, unsaved results are lost
  • Don't request 8 GPUs if you use 1
  • Don't use --nodelist in production
  • Don't run interactive jobs 24/7 – use batch jobs
  • Don't use gpu_long for short jobs – use gpu_short to get scheduled faster

8. FAQ

Do I need to change my Python code?

No, but make sure your script's paths point to the scratch directory you created, not the submit directory, when reading/writing large data.

What if my dataset is 5TB?

Don't copy it if avoidable. Keep large datasets in /datasets/ and reference them directly, or copy only the subset you need.

Can I submit from any directory?

Yes, but you are responsible for copying any needed input data to scratch yourself.

How long does copying take?

Roughly 1–2 seconds per GB with rsync, depending on load.

What if my job is killed mid-training?

Any results not yet copied back to PROJECT/HOME are lost. Use periodic intermediate syncs for long jobs.

How do I find my account and QoS?

Run: sacctmgr show user $USER withassoc format=account,qos

Which partition should I use?

Use gpu_short / cpu_short for jobs under 2 days. Use gpu_long / cpu_long for jobs up to 4 days.


9. Complete Working Example

#!/bin/bash
################################################################################
# BERT-Large Fine-tuning on SKQuAD Dataset
# Expected runtime: ~2 hours on 2x H200 GPUs
################################################################################

#SBATCH --job-name=skquad_bert
#SBATCH --output=%x_%j.out
#SBATCH --error=%x_%j.err
#SBATCH --partition=gpu_short
#SBATCH --account=perun2501234
#SBATCH --qos=perun2501234
#SBATCH --gres=gpu:2
#SBATCH --cpus-per-task=16
#SBATCH --mem=96G
#SBATCH --time=04:00:00

SCRATCH_DIR="/mnt/scratch/$USER/job_$SLURM_JOB_ID"
mkdir -p "$SCRATCH_DIR"

# Copy input data to scratch
rsync -a "$SLURM_SUBMIT_DIR"/data "$SCRATCH_DIR"/
cd "$SCRATCH_DIR"

export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
export CUDA_LAUNCH_BLOCKING=0

echo "═══════════════════════════════════════════════════════════"
echo "Job ID:   $SLURM_JOB_ID"
echo "Job Name: $SLURM_JOB_NAME"
echo "Node:     $(hostname)"
echo "GPUs:     $CUDA_VISIBLE_DEVICES"
echo "Scratch:  $SCRATCH_DIR"
echo "═══════════════════════════════════════════════════════════"
echo

nvidia-smi --query-gpu=name,memory.total --format=csv
echo

python3 "$SLURM_SUBMIT_DIR"/train_bert.py \
    --model_name bert-large-uncased \
    --dataset skquad \
    --output_dir checkpoints/ \
    --num_train_epochs 3 \
    --per_device_train_batch_size 16 \
    --learning_rate 3e-5 \
    --warmup_steps 500 \
    --save_steps 1000 \
    --logging_steps 100 \
    --fp16

echo
echo "═══════════════════════════════════════════════════════════"
echo "Training complete! Copying results back to project storage..."
echo "═══════════════════════════════════════════════════════════"

# Copy results back
rsync -a checkpoints/ "$SLURM_SUBMIT_DIR"/results_job_$SLURM_JOB_ID/checkpoints/

Submit the Job

sbatch train_skquad.sh

Monitor Progress

tail -f skquad_bert_XXXXX.out