API reference¶
The whole public surface is re-exported at the top level, so import nb2slurm is
the only import you need.
Workflow¶
- class nb2slurm.Workflow(name, notebooks, kernel, varying, resources=<factory>, project_dir='.', conda_env=None, setup=<factory>, mounts=<factory>, runner_name='run_workflow.py', concurrency=0, output_dir='output', done_csv='done/done.csv', jobs_json='jobs.json', environment=None, kernels=<factory>, extra_environments=<factory>, submitted_jobs=<factory>)[source]¶
Bases:
objectA notebook chain plus the resources and cluster details it needs.
The docstring comments below are the constructor arguments; only
name,notebooks,kernelandvaryingare required.- Parameters:
- conda_env: str | None = None¶
conda env activated in the job (defaults to
environment.nameif given)
- environment: Environment | None = None¶
primary conda env + kernel to run in
- extra_environments: list[Environment]¶
extra envs to also create
- write_jobs_txt(jobs_json=None)[source]¶
Flatten the jobs JSON into
scripts/jobs.txt(one job per line).This is what the generated bash submitters read, so they never have to parse JSON. The order matches
varying.
- create_environment(ssh=None, overwrite=False)[source]¶
Create the conda env(s) + Jupyter kernel(s) on the HPC (or locally).
Creates the primary
environmentplus anyextra_environments(each to its ownenvironment_<name>.yml). Run once before the first submit; re-running updates the env in place. Passoverwrite=Trueto delete and rebuild from scratch. Returns a list of results.
- remove_environment(ssh=None)[source]¶
Delete the conda env(s) + Jupyter kernel(s) on the HPC (or locally).
Removes the primary
environmentand anyextra_environments. Safe when nothing is there (a missing env/kernel is ignored) — handy to recover from a half-built env. Returns a list of results.- Parameters:
ssh (SSHConfig | None)
- jobs_from_json(jobs_json=None)[source]¶
Read the jobs JSON and return one tuple of varying values per job.
- build_outputs(jobs_json=None)[source]¶
Pre-create the full output tree from the jobs JSON, under
output_dir.
- submit(items=None, ssh=None, dry_run=False, jobs_json=None, concurrency=None)[source]¶
Submit one SLURM job per item. Returns the submitted job ids.
By default the jobs are read from the nested JSON (
self.jobs_json): each root-to-leaf path becomes one job. Override by passing an explicititemslist, or a differentjobs_jsonpath.Concurrency is throttled with SLURM dependencies: job N waits for job N-
concurrencyto finish (afterany), so at mostconcurrencyrun at once without holding the queue open. Passconcurrencyhere to overrideself.concurrencyfor this call; set it to0(orNoneon the workflow) to submit every job at once with no dependencies — handy for a handful of quick, independent jobs.
- check(ssh=None, raise_on_error=True)[source]¶
Verify the cluster is ready before submitting.
Confirms remote_dir, the notebooks, the built scripts, the conda env and the Jupyter kernel all exist — so a missing piece is a clear message here rather than a cryptic SLURM failure later. Returns a list of result dicts; raises RuntimeError on the first failure unless
raise_on_error=False.
- status(ssh=None, user=None)[source]¶
Return current queue entries as a list of dicts (parsed squeue).
- reset_done(ssh=None)[source]¶
Clear the done ledger so the next
submit()reruns every job.Deletes
done_csvon the cluster whensshis given, or locally otherwise. Safe to call even if the ledger doesn’t exist yet. To rerun only a subset, edit the CSV directly (it’s justkeyper row) or pass an explicititems=list tosubmit().- Parameters:
ssh (SSHConfig | None)
- Return type:
None
- push(ssh, delete=False, dry_run=False)[source]¶
Upload the project to the cluster (source only — never outputs).
Syncs notebooks/, scripts/, jobs.json, environment.yml, … up to
remote_dir. The output dir (output/by default) anddone/are excluded, so pushing your latest notebook edits can never wipe results already on the cluster.
- pull(ssh, delete=False, dry_run=False)[source]¶
Download results from the cluster (outputs only).
Pulls only the output dir (
output/by default) and thedone/ledger back to the project. It never fetches notebooks/ or scripts/, so a results-sync can’t overwrite a notebook you changed locally while jobs were running.
Environment¶
- class nb2slurm.Environment(name, kernel, python='3.11', channels=<factory>, conda_packages=<factory>, pip_packages=<factory>)[source]¶
Bases:
objectA conda environment + Jupyter kernel to create for the workflow.
- Parameters:
- write(project_dir='.', filename='environment.yml')[source]¶
Write the
environment.ymlinto the project directory.
- exists(ssh=None, project_dir='.')[source]¶
Return True if the conda env already exists — on the HPC (ssh) or locally.
- remove(ssh=None, project_dir='.', stream=True)[source]¶
Delete the conda env and its Jupyter kernel — on the HPC (ssh) or locally.
Safe to call when nothing is there yet (a missing env/kernel is ignored). Use it to recover from a half-built env or to force a clean rebuild.
- create(ssh=None, project_dir='.', filename='environment.yml', stream=True, overwrite=False)[source]¶
Create the env and register the kernel — on the HPC (ssh) or locally.
Writes
environment.ymlfirst if it is missing. If the env already exists it is updated in place; passoverwrite=Trueto delete and rebuild it from scratch.stream=True(the default) echoes conda/pip output live, since a solve + downloads can take minutes and would otherwise look like a hang.
SSH¶
- class nb2slurm.SSHConfig(host, user, remote_dir, port=22, key_filename=None, password=None, passphrase=None, extra_connect_kwargs=<factory>)[source]¶
Bases:
objectConnection details for the HPC login node.
Provide a
key_filename(or rely on an agent/known config).remote_diris the project directory on the cluster that the generated scripts live in; commands are run from there.Auth notes:
passphrasedecrypts a passphrase-protected private key (this is what Snellius and most clusters use — the “password” you type is your key’s passphrase, not a server account password).passwordis for actual password authentication (rare on HPC).Best of all is loading the key into
ssh-agent(ssh-add): then you need neither here, and rsync (push/pull) also works without prompts.
- Parameters:
- password: str | None = None¶
account password, if your cluster uses one (never written to disk by save_config)
- passphrase: str | None = None¶
passphrase unlocking an encrypted private key (never written to disk either)
- key_path()[source]¶
The private key path with
~expanded, orNoneif unset.paramiko opens
key_filenamedirectly and does not expand~, so we resolve it here (e.g.~/.ssh/id_rsa-> the absolute path).- Return type:
str | None
- test_connection(command='hostname && whoami')[source]¶
Try to connect and run a trivial command; print a clear ok/fail.
A quick first check before push/submit. Returns True on success. On an encrypted key it prints the passphrase hint from
_connect.
- run(command, cwd=None, stream=False)[source]¶
Run a single command on the cluster and return its result.
Output is drained continuously while the command runs, so a chatty command (
conda env create,pip install) can’t fill paramiko’s channel window and deadlock againstrecv_exit_status. Passstream=Trueto also echo output live — useful for long-running builds where you’d otherwise see nothing until they finish.
- nb2slurm.generate_key(path=None, key_type='rsa', bits=4096, comment=None, overwrite=False, show=True)[source]¶
Create an SSH keypair at
path(+<path>.pub).key_typeis"rsa"(default,bitswide) or"ed25519"— the modern, fixed-size type, recommended as it sidesteps the legacy RSA/DSA baggage some setups trip over. Whenpathis omitted it defaults to~/.ssh/id_rsaor~/.ssh/id_ed25519to matchkey_type.Returns
(private_path, public_path). The private key is written 0600 and the public key inauthorized_keysformat. An existing key is left alone unlessoverwrite=True, so this is safe to call repeatedly. Point yourSSHConfig(key_filename=...)atpath.nb2slurm can’t install the key for you — many HPCs disable password login, so there’s no way in. With
show=True(default) the public key is printed so you can copy it into your cluster’s key-upload page (or its~/.ssh/authorized_keys);nb2slurm.public_key(path)reprints it later.
- nb2slurm.public_key(path='~/.ssh/id_rsa')[source]¶
Return the public key line for
path(reads<path>.pub).This is the text you paste into your HPC —
print(nb2slurm.public_key())then copy it into the cluster’s key-upload page (or~/.ssh/authorized_keyson a login node, if your HPC lets you edit it directly).
Inside your notebooks¶
- class nb2slurm.Settings[source]¶
Bases:
objectRead/write the per-run
settings.json.
Jobs, outputs and saved config¶
- class nb2slurm.Structure(spec=None)[source]¶
Bases:
objectA nested-dict description of the jobs (and their output folders) for a run.
- Parameters:
spec (Mapping[str, Any] | None)
- class nb2slurm.Done(csv_file)[source]¶
Bases:
objectA CSV ledger of completed subjects, with concurrency-safe writes.
- Parameters:
csv_file (str | Path)