Skip to content

8.4 Restarts and HPC: preserve correctness while scaling

These are unexecuted teaching inputs and starting models. Original diagrams are schematics, not calculated results. Validate version-specific syntax, licensed or authorized data, numerical convergence and the scientific model before using this workflow.

8.4.1 Model, units and provenance

PW cutoffs and energies use Ry, common force output uses Ry/bohr, and pressure uses kbar. Geometry cards state their coordinate units. Different executables have distinct grammars and time-unit conventions.

Shared inputs, conventions and evidence

Original schematic: Restarts and HPC: preserve correctness while scaling. No numerical results are claimed.
Original schematic: Restarts and HPC: preserve correctness while scaling. No numerical results are claimed.

8.4.2 Unexecuted inputs and explicit deltas

Use the accompanying instructions to identify the parent calculation and placement of every delta; a snippet is not automatically a standalone input. Preserve all blank-line and file-provenance requirements.

8.4.2.1 Input block 1

! Initial run, with enough margin before the scheduler kills it:
&CONTROL
 calculation='scf', prefix='si_restart', outdir='./scratch/si_restart',
 pseudo_dir='./pseudo', restart_mode='from_scratch', max_seconds=120.0
/
! Use the remaining validated Si-base input unchanged.
! Continuation: same physical input, prefix/outdir and compatible layout;
! change restart_mode='restart', with a suitable larger time budget.

8.4.2.2 Input block 2

# Illustrative launch patterns; use the cluster's supported launcher.
export OMP_NUM_THREADS=1
mpirun -np 8 pw.x -nk 2 -in inputs/si.restart.in > outputs/si.restart.out
# Validate the rank/pool choice against the actual k-point count and build.

8.4.3 Worked investigation

8.4.3.1 Intuition and prerequisites

Production work fails most expensively when a job appears restartable but its required files were never saved or were overwritten. Parallel speedup is useful only if the same scientific problem reaches the same converged state. Use a modest converged Si or Al workload first; do not benchmark a different mesh on each machine. Know your scheduler's wall-time limits and the filesystem accessible from every rank.

8.4.3.2 Original controlled restart protocol

See input block 1 above.

8.4.3.3 Scaling, checks and exercise

Benchmark fixed work at a few sensible MPI/thread/pool combinations. Record elapsed time, peak memory where available, SCF iteration count, communication/I/O timing and final energy. More pools may reduce memory or expose k parallelism but can become ineffective when there are too few k points. Do not assume a Γ-only molecular calculation scales like a dense metallic mesh. Check filesystem quota before large wavefunction writes and archive a minimal reproducibility bundle before cleaning scratch.

Exercise: compare uninterrupted and restarted runs, then intentionally omit a copied scratch component in a separate test and learn to identify the failure without corrupting the original. Build a submission checklist with required files, expected storage, stop margin and success checks. A timing result with different iteration counts needs interpretation; speedup from looser convergence is not parallel speedup.

8.4.5 Sources and further reading