Practical Guide · E

Inspect QE HPC Calculations from the Terminal

Inspect Quantum ESPRESSO files and Slurm records from a terminal while keeping scheduler state, program completion, numerical evidence, and scientific acceptance separate.

Purpose

Terminal inspection answers bounded operational questions: Where is the calculation? Which executable produced a file? Is a recorded job pending, running, or finished? Does a particular output contain termination and solver markers? Which artifacts exist, and are their bytes the ones that were reviewed?

Those observations are not interchangeable. A Slurm COMPLETED state does not prove that pw.x reached self-consistency. JOB DONE. does not prove convergence of the target observable. An SCF marker does not establish a ground state, a converged phonon dispersion, or a scientific conclusion. Record each observation with its command, path, job ID, and observation time.

The commands below are read-only unless a side effect is explicitly identified. They use the committed bcc-fe-spin-qe case when a concrete file is needed, so they can be run from the repository root without inventing a calculation result.

This is the terminal part of a larger human workflow. Before accepting the calculation, also open the structure in an appropriate viewer, inspect convergence tables and plots, and compare the physical result with the intended model and relevant literature. Terminal evidence can locate and classify a problem; it cannot replace those visual and scientific checks.

Prepare

Anchor the inspection before reading a log:

case_root=examples/cases/bcc-fe-spin-qe
(
  cd -- "$case_root"
  pwd -P
)
test -d "$case_root"

Changing directory inside the subshell proves that the recorded case path is reachable without silently changing the caller’s shell. It does not prove that the path belongs to the intended run or that any output is valid.

pwd -P proves the shell’s resolved current directory at that moment. It does not prove that the directory is the intended project authority, that no process is writing there, or that a similarly named scratch directory is equivalent.

Inventory the relevant executables without starting a calculation:

for exe in pw.x ph.x bands.x dos.x projwfc.x q2r.x matdyn.x; do
  if path=$(command -v "$exe" 2>/dev/null); then
    printf '%-10s %s\n' "$exe" "$path"
  else
    printf '%-10s %s\n' "$exe" 'NOT FOUND'
  fi
done

if mpi_path=$(command -v mpirun 2>/dev/null); then
  printf '%-10s %s\n' 'mpirun' "$mpi_path"
  mpirun --version | head -n 5
else
  printf '%-10s %s\n' 'mpirun' 'NOT FOUND'
fi

if pw_path=$(command -v pw.x 2>/dev/null); then
  "$pw_path" -h </dev/null 2>&1 | head -n 5
else
  printf '%-10s %s\n' 'pw.x' 'NOT FOUND' >&2
fi

command -v proves only which command the current shell would resolve. mpirun --version identifies the currently resolved launcher, not the launcher used by a previous job and not its compatibility with the installed QE build or the scheduler. The help banner can identify a local executable build when that build supports -h. None of these commands proves that a calculation succeeded. For a completed run, compare this shell evidence with the version banner inside the exact saved output:

find "$case_root/output" -type f -name '*.out' \
  -exec grep -Hn -m 1 -E 'Program (PWSCF|PHONON|BANDS|DOS|PROJWFC|Q2R|MATDYN) v\.' {} +

grep is used because it is normally present on HPC systems. If rg is installed, it can provide a faster recursive search, but a search tool does not strengthen the evidence in the matching line.

Check capacity in both blocks and inodes:

df -h -- . "${TMPDIR:-.}"
df -ih -- . "${TMPDIR:-.}"

This can reveal that the inspected filesystem or declared temporary directory is full. It cannot prove that a compute node sees the same mount, that quota remains available, or that a future write will succeed.

Build a bounded file inventory:

ls -lah -- "$case_root"
find "$case_root" -maxdepth 4 -type f \
  -printf '%TY-%Tm-%TdT%TH:%TM:%TS %12s %p\n' | sort
find "$case_root" -maxdepth 4 -type f -size 0 -print

ls and find establish names, sizes, and filesystem timestamps visible to this process. They do not establish file completeness, content identity, provenance, or whether a zero-byte file is an expected stderr capture or a failed output. GNU find -printf is implementation-specific; on another system, use that site’s supported formatting while preserving the same fields.

Run

Invoke only the branch the question requires

Select one executable stage after checking its parent evidence:

  • pw.x reads the reviewed input for the declared SCF, relaxation, NSCF, or bands-parent task.
  • bands.x requires a compatible completed pw.x bands calculation.
  • dos.x requires a compatible dense-grid pw.x NSCF parent.
  • projwfc.x requires the compatible pw.x prefix, outdir, and wavefunctions.
  • ph.x starts the requested response branch from a compatible accepted pw.x parent.
  • q2r.x requires the complete compatible dynamical-matrix set for the declared q mesh.
  • matdyn.x requires the corresponding q2r.x force constants.

Set the three variables to one reviewed stage. The block refuses existing stdout or stderr and invokes only that one program:

: "${QE_PROGRAM:?Set one executable: pw.x, bands.x, dos.x, projwfc.x, ph.x, q2r.x, or matdyn.x}"
: "${QE_INPUT:?Set QE_INPUT to the reviewed input for that stage}"
: "${QE_STAGE:?Set a filesystem-safe stage name}"
case "$QE_PROGRAM" in
  pw.x|bands.x|dos.x|projwfc.x|ph.x|q2r.x|matdyn.x) ;;
  *) printf 'Unsupported QE_PROGRAM: %s\n' "$QE_PROGRAM" >&2; exit 2 ;;
esac
test -f "$QE_INPUT"
test ! -e "$QE_STAGE.out"
test ! -e "$QE_STAGE.err"

if "$QE_PROGRAM" -in "$QE_INPUT" > "$QE_STAGE.out" 2> "$QE_STAGE.err"; then
  qe_status=0
else
  qe_status=$?
fi
printf '%s\n' "$qe_status" > "$QE_STAGE.exit-status"
tail -n 40 -- "$QE_STAGE.out" "$QE_STAGE.err"
test "$qe_status" -eq 0

This proves only that the shell attempted the one selected executable and retained its exit status, stdout, and stderr. The parent list does not define one mandatory pw.x → bands.x → dos.x → projwfc.x → ph.x → q2r.x → matdyn.x sequence. Keep every branch in its own prefix/outdir and record the actual launcher, working directory, input hash, and parent artifact identity.

Before a submission, a Slurm installation may support a non-submitting syntax and feasibility check:

sbatch --test-only job.slurm

This checks how the local scheduler parses the request and may estimate feasibility. It does not submit a job, test the QE executable, confirm file paths on a compute node, or guarantee a start time.

The next command has a side effect: it submits a job. Use it only after checking the script, account, partition, working directory, output paths, and ownership:

submission=$(sbatch --parsable job.slurm)
job_id=${submission%%;*}
printf 'submitted job_id=%s\n' "$job_id"

The returned ID proves that the scheduler accepted a request. It does not prove that the job started or that QE completed.

For a recorded job ID, inspect current and accounting views separately:

job_id=${job_id:?Set job_id to the exact recorded Slurm job ID}
squeue -j "$job_id" -o '%.18i %.9T %.10M %.10l %.6D %R'
sacct -j "$job_id" --format=JobIDRaw,JobName%24,State,ExitCode,Elapsed,Timelimit,AllocCPUS,MaxRSS
scontrol show job -dd "$job_id"

squeue is a current queue observation. An empty result can mean completion, purge, the wrong cluster, or the wrong ID. sacct is scheduler accounting and can preserve step exit codes after the queue entry disappears, but availability and fields depend on site configuration. scontrol show job -dd exposes the scheduler’s detailed record, including work directory, command, resources, and reason fields when retained. None reads the scientific meaning of QE output.

Inspect processes owned by the current user on the current host without signalling them:

ps -u "$USER" -o pid,ppid,stat,etime,cmd --forest

This proves only that matching local processes were visible at that instant. It does not inspect another compute node, establish scheduler ownership, identify the correct calculation without its recorded path and job ID, or authorize cancellation.

Follow a known live log by name:

live_log=${live_log:?Set live_log to the exact stdout path for the recorded stage}

# Follow the currently open file descriptor; stop with Ctrl-C.
tail -f -- "$live_log"

# Alternative when the named file may be rotated or recreated; stop with Ctrl-C.
tail -F -- "$live_log"

tail -f follows the currently open file descriptor. tail -F follows the filename and retries if it disappears, which is useful for rotation but can attach to replacement bytes. Neither proves scheduler state, normal program termination, SCF convergence, or scientific validity.

tail -F proves that new bytes become visible at that path and continues across file replacement. Silence does not prove a hung job, and new text does not prove progress or convergence. Exit with Ctrl-C; this stops only the local viewer.

Cancellation changes scheduler state and can terminate a running calculation. Use it only for a job that you own and have positively identified. Re-read the job record, require an exact confirmation, then request cancellation:

job_id=${job_id:?Set job_id to the exact recorded Slurm job ID}
job_owner=$(squeue -h -j "$job_id" -o '%u' | head -n 1)
test -n "$job_owner"
test "$job_owner" = "$USER"
scontrol show job -dd "$job_id"
read -r -p "Type CANCEL $job_id to confirm: " confirmation
test "$confirmation" = "CANCEL $job_id"
scancel "$job_id"

This sends a cancellation request for the confirmed job ID; it does not prove that every job step stopped, that output buffers were flushed, or that restart data are usable. Recheck squeue, sacct, scontrol, active processes, and output timestamps. A slow, quiet, or scientifically blocked job is not by itself a reason or authorization to cancel it.

Check

Choose an exact output rather than searching whichever file happens to be newest:

out="$case_root/output/attempt-02-pmix/fm-k12.out"
test -f "$out"
head -n 40 -- "$out"
tail -n 80 -- "$out"
less -S +G -- "$out"

head is useful for the program banner, parallel layout, and input echo. tail is useful for the most recent termination context. less permits non-destructive review of the complete file. A plausible header and tail do not prove that the middle is intact; inspect the complete stage and bind its hash.

Count and search without silently treating a marker as acceptance:

wc -l -- "$out"
if grep -n -E \
  'Program (PWSCF|PHONON|BANDS|DOS|PROJWFC|Q2R|MATDYN)|JOB DONE|convergence has been achieved|No convergence has been achieved|^!.*total energy|Fermi energy|Forces acting on atoms|Total force|total[[:space:]]+stress|ATOMIC_POSITIONS' \
  -- "$out"; then
  :
else
  grep_status=$?
  case "$grep_status" in
    1) printf '%s\n' 'No requested marker was found.' >&2 ;;
    *) exit "$grep_status" ;;
  esac
fi

awk '
  /^!.*total energy/ { energy=$0 }
  /the Fermi energy is/ { fermi=$0 }
  END {
    if (energy == "") exit 1
    print energy
    if (fermi != "") print fermi
  }
' "$out"

grep -n -i -E \
  'warning|error in routine|stopping|not converged|no convergence|segmentation fault|out of memory|oom-kill' \
  -- "$out" || true

wc -l counts newline-terminated records only; it does not prove completeness. The guarded grep returns 0 when at least one requested marker is found, 1 when none is found, and a value greater than 1 for an access or execution error; only status 1 is translated into “not found.” The awk command prints the last matching total-energy line and, when present, the last matching Fermi-energy line. It is a bounded text extraction, not a unit-aware parser and not evidence of convergence.

The first search locates evidence candidates; it does not decide which gate they satisfy. The second search locates adverse text; an empty result means only that these patterns were absent. Read surrounding lines and stderr because software- and site-specific failures use other wording.

For a relaxation, inspect one coherent output together with the exact input. Reject a concatenated stdout, show the active-coordinate mask, extract the last complete force and final-coordinate blocks, and display the last complete stress tensor when the active cell requires one:

relax_in=${relax_in:?Set relax_in to the exact relaxation input path}
relax_out=${relax_out:?Set relax_out to the exact relaxation stdout path}
test -f "$relax_in"
test -f "$relax_out"
test "$(grep -cF 'Program PWSCF v.' -- "$relax_out")" -eq 1
sed -n '/^ATOMIC_POSITIONS/,/^K_POINTS/p' "$relax_in"

awk '
  /Forces acting on atoms/ {block=$0 ORS; inside=1; next}
  inside {block=block $0 ORS}
  inside && /Total force =/ {last=block; inside=0}
  END {if (last == "") exit 1; printf "%s", last}
' "$relax_out"

awk '
  /Begin final coordinates/ {block=$0 ORS; inside=1; next}
  inside {block=block $0 ORS}
  inside && /End final coordinates/ {last=block; inside=0}
  END {if (last == "") exit 1; printf "%s", last}
' "$relax_out"

awk '
  /total[[:space:]]+stress/ {block=$0 ORS; rows=3; next}
  rows > 0 {block=block $0 ORS; rows--; if (rows == 0) last=block}
  END {
    if (last == "") print "No complete stress block found; stress is not assessed."
    else printf "%s", last
  }
' "$relax_out"

The input rows expose explicit if_pos flags when present; otherwise confirm the version-matching documented default before classifying free components. Apply the force gate to every free Cartesian component, not aggregate Total force. A stress block is required only for the active cell components declared by the protocol, but its absence is then unresolved rather than a pass. The final coordinates, force components, stress components, electronic markers, ionic stop record, and program completion remain separate evidence. None proves a minimum or physical validity.

For a declared ph.x or matdyn.x output, locate printed mode frequencies:

ph_out=${ph_out:?Set ph_out to the exact ph.x or matdyn.x stdout path}
if grep -n -E 'freq[[:space:]]*\([[:space:]]*[0-9]+\)[[:space:]]*=' "$ph_out"; then
  :
else
  grep_status=$?
  case "$grep_status" in
    1) printf '%s\n' 'No printed mode-frequency line was found.' >&2 ;;
    *) exit "$grep_status" ;;
  esac
fi

A frequency match proves only that a line with the requested QE text shape was printed. It does not establish complete q-point coverage, acoustic-sum-rule treatment, phonon convergence, or dynamical stability.

This reports matching frequencies and units at the q points present in that output. It does not establish q-space coverage, acoustic-sum-rule quality, displacement or q-mesh convergence, dynamical stability, anharmonic stability, or experimental agreement. Preserve negative or imaginary modes instead of filtering them out.

Different QE executables create different evidence:

ExecutableWhat terminal evidence can establishWhat it cannot establish alone
pw.xProgram identity, recorded exit, SCF marker, energies, occupations, forces, stress, and printed structures when requestedBasis, k-point, smearing, geometry, magnetic-state, or target-observable convergence
ph.xThe requested DFPT stage, q-point text, response convergence, dynamical-matrix output, and terminationFull q-space coverage, interpolation convergence, dynamical stability, EPC convergence, or temperature validity
bands.xPost-processing completion and existence of the declared band-data artifactAdequacy of the parent SCF/path, full-zone extrema, or a physical band gap
dos.xTotal-DOS post-processing and the declared energy-grid artifactBrillouin-zone convergence, correct normalization, or projection closure
projwfc.xProjection output and orbital/channel labels for one parent stateBasis completeness, unique chemical bonding, or equality between summed projections and total DOS without an explicit check
q2r.xConversion from a declared dynamical-matrix set to a force-constant artifactCompleteness or correctness of that set, real-space convergence, or stability
matdyn.xFrequencies/eigenvectors at requested q points from a declared force-constant artifactCorrect parent lineage, q-grid convergence, non-analytic-correction validity, anharmonic stability, or experimental agreement

Inventory downstream files and zero-byte captures explicitly:

find "$case_root" -type f \
  \( -name '*.dyn*' -o -name '*.fc' -o -name '*.freq' -o -name '*.modes' \
     -o -name '*.dos' -o -name '*.pdos*' -o -name '*.dat*' -o -name '*.json' \) \
  -printf '%12s %p\n' | sort
find "$case_root" -type f -size 0 -print

Existence and nonzero size are artifact checks, not semantic checks. A zero-byte stderr can be expected; a zero-byte dynamical matrix can be terminal adverse evidence. Interpret each file against the stage that was supposed to write it.

Compare inputs and bind bytes:

if diff -u -- \
  "$case_root/input/fm-k10.scf.in" \
  "$case_root/input/fm-k12.scf.in"; then
  printf '%s\n' 'Inputs are byte-for-byte identical.'
else
  diff_status=$?
  case "$diff_status" in
    1) printf '%s\n' 'Inputs differ; inspect the unified diff above.' >&2 ;;
    *) exit "$diff_status" ;;
  esac
fi

sha256sum -- \
  "$case_root/manifest.json" \
  "$case_root/input/fm-k12.scf.in" \
  "$out"

For diff, exit 0 means identical inputs, exit 1 means a real difference, and an exit value greater than 1 is an operational error that must not be reported as a scientific difference. sha256sum binds exact bytes; neither command establishes which input is scientifically appropriate.

diff makes textual changes inspectable; it does not decide whether two calculations are scientifically comparable or reveal defaults that were not printed. sha256sum binds exact bytes; matching hashes do not prove correctness, completeness, convergence, or provenance beyond the declared binding.

Read

Termination. Combine scheduler state, step exit code, program banner, fatal text, termination marker, and expected artifacts. Any one item is incomplete. The committed bcc Fe case preserves a first Slurm/Open MPI launch failure before PWSCF and a second attempt whose four QE stages each exited zero and printed JOB DONE..

SCF. Read the final SCF marker, iteration history, residual threshold, occupations, spin state, and warnings. Internal self-consistency says the declared electronic iteration stopped under its criterion. It does not establish the lowest electronic state or convergence with respect to cutoff, k mesh, smearing, cell, or candidate set.

Energy and Fermi level. Use the final exclamation-mark total energy, its units and normalization, and the occupation model. A printed Fermi energy is code-defined output for that state; comparing it across different references, charges, cells, smearings, or Hamiltonians is not automatically meaningful. Small last-iteration energy changes are not a substitute for an external convergence series.

Forces, stress, and coordinates. Check that force/stress printing was requested, units are recorded, constrained components are understood, and the coordinates correspond to the exact intended frame. Small forces on a symmetry-fixed one-atom SCF model do not prove relaxation, dynamical stability, or a global minimum. Stress can diagnose a fixed cell but does not by itself authorize changing it.

Warnings. Preserve warnings with their surrounding context and software version. A warning-free text search cannot prove absence of silent numerical problems. A warning does not always invalidate a result, but ignoring it without a bounded assessment leaves the claim unsupported.

Phonons. Record every computed q point, dynamical-matrix identity, q2r input set, force-constant file, matdyn path, acoustic-sum-rule treatment, and imaginary-frequency convention. One Gamma calculation cannot support a dispersion or whole-Brillouin-zone stability claim. The committed bcc Fe case contains no phonon artifacts, so it supports no phonon conclusion.

Artifacts. Match inputs, stdout, stderr, derived tables, figures, restart objects, and manifests by path, size, hash, and lineage. Artifact integrity can prove that reviewed bytes were retained. It cannot upgrade a failed numerical check or validate a scientific interpretation.

Keep the complete output open in less while reading any extracted table or plot. Inspect the iteration sequence around an apparent oscillation or warning instead of relying on the last matching line. For a relaxation, compare the final coordinates with the starting structure in a viewer and look for unexpected reconstruction or short contacts. For bands, DOS, and phonons, open the actual figure and underlying table; visual plausibility complements, but never replaces, observable-specific convergence.

If it fails

First classify the layer that failed. A pending allocation is not a QE failure. A launcher/MPI error before the program banner is not SCF nonconvergence. A nonzero wrapper exit can coexist with successful child stages. JOB DONE. can coexist with an observable-specific convergence failure.

Preserve the failed attempt before changing anything. Retain the job ID, submission command, scheduler records, exact work directory, stdout, stderr, inputs, exit codes, and hashes. Put a repaired run in a new attempt directory with a causal link to the failure; do not overwrite the adverse record.

If the output stops moving, recheck squeue, sacct, scontrol, file timestamps, sizes, and the program’s known buffering behavior. Do not infer a hang from tail -F silence. Do not cancel by default. If a cancellation is actually authorized, the operator must re-establish job ownership, scope, downstream effects, and restart safety immediately before the side effect.

Next

Run the read-only companion from the repository root:

python3 examples/practical-guides/qe_hpc_terminal_inspection.py

It checks the sizes and SHA-256 values declared by the committed case manifest, the preserved launch failure, the four recorded Attempt 02 exit codes, QE/SCF markers, selected energies and Fermi levels, and the failed numerical-convergence record. It does not call Slurm, inspect a live process, submit or cancel a job, or modify a file.

After terminal triage, use the worked audit to decide which scientific check failed. Continue with symptom-first Troubleshooting or return to the relevant A–D topic; do not proceed to a stronger claim until that check has new, traceable evidence under a predeclared criterion.

Official sources

Ways to work: PythonQuantum ESPRESSO

Companion checked with: Python 3.12; Quantum ESPRESSO 7.5 committed-output format.

Reproducibility note

The companion material was checked with Python 3.12; Quantum ESPRESSO 7.5 committed-output format. It tests only the bounded software or analysis behaviour described here; it does not establish numerical convergence, model validity, or a material property.

The manifest records at least one failed check, so not every declared check passed. Inspect the failed evidence layer before using the case; an earlier program execution or analysis step may still have completed.

Open the exact case record or its manifest.