API reference

The whole public surface is re-exported at the top level, so import nb2slurm is the only import you need.

Workflow

class nb2slurm.Workflow(name, notebooks, kernel, varying, resources=<factory>, project_dir='.', conda_env=None, setup=<factory>, mounts=<factory>, runner_name='run_workflow.py', concurrency=0, output_dir='output', done_csv='done/done.csv', jobs_json='jobs.json', environment=None, kernels=<factory>, extra_environments=<factory>, submitted_jobs=<factory>)[source]

Bases: object

A notebook chain plus the resources and cluster details it needs.

The docstring comments below are the constructor arguments; only name, notebooks, kernel and varying are required.

Parameters:
name: str

project name, used for the SLURM job names

notebooks: list[str]

the notebook chain, in order; the first one writes settings.json

kernel: str

Jupyter kernel the notebooks are executed with (must exist on the cluster)

varying: list[str]

what changes per job, in the order the levels of jobs.json nest

resources: dict

#SBATCH resources, e.g. dict(nodes=1, cpus=2, time="04:00:00")

project_dir: str = '.'

project root that everything else is relative to

conda_env: str | None = None

conda env activated in the job (defaults to environment.name if given)

setup: list[str]

raw shell lines run before the job, e.g. module load Python/3.11

mounts: list[dict]

optional rclone mounts, [{"remote": ..., "mountpoint": ...}]

runner_name: str = 'run_workflow.py'

filename of the generated papermill driver

concurrency: int = 0

max jobs running at once per submit (0 = no limit)

output_dir: str = 'output'

root for per-subject outputs (relative to project root)

done_csv: str = 'done/done.csv'

idempotency ledger (relative to project root)

jobs_json: str = 'jobs.json'

nested JSON describing the jobs to run (one leaf path = one job)

environment: Environment | None = None

primary conda env + kernel to run in

kernels: dict

per-notebook kernel overrides, {notebook_path: kernel}

extra_environments: list[Environment]

extra envs to also create

submitted_jobs: list[str]

job ids we have submitted this session (used by status/cancel)

build()[source]

Render the generated scripts into <project>/scripts/.

Return type:

dict[str, Path]

write_jobs_txt(jobs_json=None)[source]

Flatten the jobs JSON into scripts/jobs.txt (one job per line).

This is what the generated bash submitters read, so they never have to parse JSON. The order matches varying.

Parameters:

jobs_json (str | None)

Return type:

Path

create_environment(ssh=None, overwrite=False)[source]

Create the conda env(s) + Jupyter kernel(s) on the HPC (or locally).

Creates the primary environment plus any extra_environments (each to its own environment_<name>.yml). Run once before the first submit; re-running updates the env in place. Pass overwrite=True to delete and rebuild from scratch. Returns a list of results.

Parameters:
remove_environment(ssh=None)[source]

Delete the conda env(s) + Jupyter kernel(s) on the HPC (or locally).

Removes the primary environment and any extra_environments. Safe when nothing is there (a missing env/kernel is ignored) — handy to recover from a half-built env. Returns a list of results.

Parameters:

ssh (SSHConfig | None)

jobs_from_json(jobs_json=None)[source]

Read the jobs JSON and return one tuple of varying values per job.

Parameters:

jobs_json (str | None)

Return type:

list[tuple[str, …]]

build_outputs(jobs_json=None)[source]

Pre-create the full output tree from the jobs JSON, under output_dir.

Parameters:

jobs_json (str | None)

Return type:

dict[str, Path]

submit(items=None, ssh=None, dry_run=False, jobs_json=None, concurrency=None)[source]

Submit one SLURM job per item. Returns the submitted job ids.

By default the jobs are read from the nested JSON (self.jobs_json): each root-to-leaf path becomes one job. Override by passing an explicit items list, or a different jobs_json path.

Concurrency is throttled with SLURM dependencies: job N waits for job N-concurrency to finish (afterany), so at most concurrency run at once without holding the queue open. Pass concurrency here to override self.concurrency for this call; set it to 0 (or None on the workflow) to submit every job at once with no dependencies — handy for a handful of quick, independent jobs.

Parameters:
Return type:

list[str]

check(ssh=None, raise_on_error=True)[source]

Verify the cluster is ready before submitting.

Confirms remote_dir, the notebooks, the built scripts, the conda env and the Jupyter kernel all exist — so a missing piece is a clear message here rather than a cryptic SLURM failure later. Returns a list of result dicts; raises RuntimeError on the first failure unless raise_on_error=False.

Parameters:
Return type:

list[dict]

status(ssh=None, user=None)[source]

Return current queue entries as a list of dicts (parsed squeue).

Parameters:
Return type:

list[dict]

cancel(ssh=None, job_ids=None)[source]

Cancel jobs. Defaults to the ones submitted this session.

Parameters:
Return type:

None

reset_done(ssh=None)[source]

Clear the done ledger so the next submit() reruns every job.

Deletes done_csv on the cluster when ssh is given, or locally otherwise. Safe to call even if the ledger doesn’t exist yet. To rerun only a subset, edit the CSV directly (it’s just key per row) or pass an explicit items= list to submit().

Parameters:

ssh (SSHConfig | None)

Return type:

None

push(ssh, delete=False, dry_run=False)[source]

Upload the project to the cluster (source only — never outputs).

Syncs notebooks/, scripts/, jobs.json, environment.yml, … up to remote_dir. The output dir (output/ by default) and done/ are excluded, so pushing your latest notebook edits can never wipe results already on the cluster.

Parameters:
pull(ssh, delete=False, dry_run=False)[source]

Download results from the cluster (outputs only).

Pulls only the output dir (output/ by default) and the done/ ledger back to the project. It never fetches notebooks/ or scripts/, so a results-sync can’t overwrite a notebook you changed locally while jobs were running.

Parameters:

Environment

class nb2slurm.Environment(name, kernel, python='3.11', channels=<factory>, conda_packages=<factory>, pip_packages=<factory>)[source]

Bases: object

A conda environment + Jupyter kernel to create for the workflow.

Parameters:
name: str

conda environment name (conda activate <name> in the job)

kernel: str

Jupyter kernel to register; must match Workflow(kernel=...)

python: str = '3.11'

Python version for the environment

channels: list[str]

conda channels, in priority order

conda_packages: list[str]

packages installed with conda/mamba

pip_packages: list[str]

packages installed with pip (nb2slurm itself is needed in the job)

to_yaml()[source]

Render an environment.yml for conda/mamba.

Return type:

str

write(project_dir='.', filename='environment.yml')[source]

Write the environment.yml into the project directory.

Parameters:
Return type:

Path

exists(ssh=None, project_dir='.')[source]

Return True if the conda env already exists — on the HPC (ssh) or locally.

Parameters:
Return type:

bool

remove(ssh=None, project_dir='.', stream=True)[source]

Delete the conda env and its Jupyter kernel — on the HPC (ssh) or locally.

Safe to call when nothing is there yet (a missing env/kernel is ignored). Use it to recover from a half-built env or to force a clean rebuild.

Parameters:
Return type:

CommandResult

create(ssh=None, project_dir='.', filename='environment.yml', stream=True, overwrite=False)[source]

Create the env and register the kernel — on the HPC (ssh) or locally.

Writes environment.yml first if it is missing. If the env already exists it is updated in place; pass overwrite=True to delete and rebuild it from scratch. stream=True (the default) echoes conda/pip output live, since a solve + downloads can take minutes and would otherwise look like a hang.

Parameters:
Return type:

CommandResult

SSH

class nb2slurm.SSHConfig(host, user, remote_dir, port=22, key_filename=None, password=None, passphrase=None, extra_connect_kwargs=<factory>)[source]

Bases: object

Connection details for the HPC login node.

Provide a key_filename (or rely on an agent/known config). remote_dir is the project directory on the cluster that the generated scripts live in; commands are run from there.

Auth notes:

  • passphrase decrypts a passphrase-protected private key (this is what Snellius and most clusters use — the “password” you type is your key’s passphrase, not a server account password).

  • password is for actual password authentication (rare on HPC).

  • Best of all is loading the key into ssh-agent (ssh-add): then you need neither here, and rsync (push/pull) also works without prompts.

Parameters:
  • host (str)

  • user (str)

  • remote_dir (str)

  • port (int)

  • key_filename (str | None)

  • password (str | None)

  • passphrase (str | None)

  • extra_connect_kwargs (dict)

host: str

login node hostname

user: str

your username on the cluster

remote_dir: str

the project directory on the cluster; commands run from there

port: int = 22

SSH port

key_filename: str | None = None

private key path, e.g. ~/.ssh/id_ed25519 (~ is expanded for you)

password: str | None = None

account password, if your cluster uses one (never written to disk by save_config)

passphrase: str | None = None

passphrase unlocking an encrypted private key (never written to disk either)

extra_connect_kwargs: dict

extra keyword arguments passed straight to paramiko.SSHClient.connect

key_path()[source]

The private key path with ~ expanded, or None if unset.

paramiko opens key_filename directly and does not expand ~, so we resolve it here (e.g. ~/.ssh/id_rsa -> the absolute path).

Return type:

str | None

rsync_ssh()[source]

The -e transport string rsync should use (ssh + port + key).

Return type:

str

rsync_target(subpath='')[source]

A user@host:remote_dir/<subpath> spec for rsync.

Parameters:

subpath (str)

Return type:

str

test_connection(command='hostname && whoami')[source]

Try to connect and run a trivial command; print a clear ok/fail.

A quick first check before push/submit. Returns True on success. On an encrypted key it prints the passphrase hint from _connect.

Parameters:

command (str)

Return type:

bool

run(command, cwd=None, stream=False)[source]

Run a single command on the cluster and return its result.

Output is drained continuously while the command runs, so a chatty command (conda env create, pip install) can’t fill paramiko’s channel window and deadlock against recv_exit_status. Pass stream=True to also echo output live — useful for long-running builds where you’d otherwise see nothing until they finish.

Parameters:
Return type:

CommandResult

nb2slurm.generate_key(path=None, key_type='rsa', bits=4096, comment=None, overwrite=False, show=True)[source]

Create an SSH keypair at path (+ <path>.pub).

key_type is "rsa" (default, bits wide) or "ed25519" — the modern, fixed-size type, recommended as it sidesteps the legacy RSA/DSA baggage some setups trip over. When path is omitted it defaults to ~/.ssh/id_rsa or ~/.ssh/id_ed25519 to match key_type.

Returns (private_path, public_path). The private key is written 0600 and the public key in authorized_keys format. An existing key is left alone unless overwrite=True, so this is safe to call repeatedly. Point your SSHConfig(key_filename=...) at path.

nb2slurm can’t install the key for you — many HPCs disable password login, so there’s no way in. With show=True (default) the public key is printed so you can copy it into your cluster’s key-upload page (or its ~/.ssh/authorized_keys); nb2slurm.public_key(path) reprints it later.

Parameters:
Return type:

Tuple[Path, Path]

nb2slurm.public_key(path='~/.ssh/id_rsa')[source]

Return the public key line for path (reads <path>.pub).

This is the text you paste into your HPC — print(nb2slurm.public_key()) then copy it into the cluster’s key-upload page (or ~/.ssh/authorized_keys on a login node, if your HPC lets you edit it directly).

Parameters:

path (str)

Return type:

str

Inside your notebooks

class nb2slurm.Settings[source]

Bases: object

Read/write the per-run settings.json.

static write(outdir, settings)[source]

Write settings to <outdir>/settings.json and return the path.

Call this from your first notebook. outdir is supplied to that notebook by the generated runner, so you never hardcode a path.

Parameters:
Return type:

Path

static load(settings_path)[source]

Load a settings file. Call this at the top of every later notebook.

Parameters:

settings_path (str | Path)

Return type:

dict[str, Any]

nb2slurm.on_hpc()[source]

Return True if running inside a nb2slurm/SLURM batch job.

Return type:

bool

Jobs, outputs and saved config

class nb2slurm.Structure(spec=None)[source]

Bases: object

A nested-dict description of the jobs (and their output folders) for a run.

Parameters:

spec (Mapping[str, Any] | None)

classmethod from_json(path)[source]

Load the hierarchy from a JSON file.

Parameters:

path (str | Path)

Return type:

Structure

jobs()[source]

Return one tuple of path components per job (root-to-leaf).

Return type:

list[tuple[str, …]]

paths(base='.')[source]

Return {relative/path: Path} for every leaf. No I/O.

Parameters:

base (str | Path)

Return type:

dict[str, Path]

build(base='.')[source]

Create every leaf folder under base and return {relative/path: Path}.

Parameters:

base (str | Path)

Return type:

dict[str, Path]

class nb2slurm.Done(csv_file)[source]

Bases: object

A CSV ledger of completed subjects, with concurrency-safe writes.

Parameters:

csv_file (str | Path)

is_done(key)[source]

Return True if key is already recorded.

Parameters:

key (str)

Return type:

bool

mark(key)[source]

Record key as done (no-op if already present).

Parameters:

key (str)

Return type:

None

clear()[source]

Delete the ledger so every key is treated as not-done again.

No-op if the file doesn’t exist. It’s just a CSV, so removing individual rows by hand also works if you only want to rerun a subset.

Return type:

None

nb2slurm.save_config(path, *, workflow, ssh=None)[source]

Write the workflow (and optional SSH) config to path as JSON.

Parameters:
Return type:

Path

nb2slurm.load_config(path)[source]

Read a config file back into (workflow, ssh).

ssh is None if none was saved. Passwords are never stored, so set ssh.password yourself afterwards if your cluster needs one.

Parameters:

path (str | Path)

Return type:

tuple[Workflow, SSHConfig | None]