Docs menu
- Home
- Docs
- Requirements
- Automation
Automation: one command, locked and seeded
Automation removes manual work when you or others repeat the workflow: one command, locked versions, fetched inputs and recorded seeds. It is assessed beside the level and is not a step toward L3.
Last updated
AutomationSeparate from the level
1 confirmed gap · 1 needs evidence · 2/4 confirmed
- ?One command produces the resultsNeeds evidence
- The software environment is pinnedConfirmed gap
- Input retrieval is automatedConfirmed
- Random seeds and variability are controlledConfirmed
Level relationshipSeparate from the level
Automation makes a workflow cheap to repeat, for you and for anyone checking it. Evidence of automation in the files does not show that the workflow runs, so it never raises the level. Seeds are not always enough: the PyTorch notes warn that results can differ between CPU and GPU.
In this guide
One command produces the results
- What counts
- A
Makefile,Snakefile, Nextflow pipeline orrun.shregenerates every key result without manual or graphical steps. It may start from a deposited, documented intermediate dataset. - Common gaps
- Scripts in an order only you know, a GUI step, or a start from an undocumented intermediate. Code using an external service: a Warning.
- How to fix
- Put every step into one entry point, from fetching data to the last figure. Make tracking services optional and say how to run without them:
.PHONY: all data smoke
all: data results/table1.csv results/figure2.pdf
data:
bash scripts/get_data.sh
results/table1.csv: data
python -m analysis.table1 --out $@
results/figure2.pdf: results/table1.csv analysis/figure2.R
Rscript analysis/figure2.R
smoke:
python -m analysis.table1 --subsample 0.01 --out /tmp/table1_smoke.csvThe software environment is pinned
- What counts
- Exact versions of every package in a lock file or container, a base image pinned by digest or tag, and runtime version stated.
- Common gaps
- Version ranges, a
latestbase image, a stale lock, results from a runtime the pins exclude. Notebooks run in an undefined environment: a Warning. - How to fix
- Lock from a clean environment that runs your analysis and produce the results in it. Containers start from a versioned base image:
# Python: lock every direct and indirect version in uv.lock
uv lock
# R: record the R version and every package version in renv.lock
Rscript -e 'renv::snapshot()'Input retrieval is automated
- What counts
- The run command downloads each open input from a persistent identifier and checks its checksum; the docs say which token any protected download needs.
- Common gaps
- Links someone must click, unzip and move by hand, an S3 bucket that needs a key nobody explains, or data behind a registration.
- How to fix
- For restricted data, offer a small open subset. Call a fetch step from your one command, such as pooch, and name the token variable:
#!/usr/bin/env bash
# scripts/get_data.sh: fetch and verify every open input
set -euo pipefail
mkdir -p data/raw
curl -fL -o data/raw/counts.csv.gz \
"https://zenodo.org/records/0000000/files/counts.csv.gz"
sha256sum -c data/checksums.sha256Random seeds and variability are controlled
- What counts
- Every stochastic call in result code is seeded, GPU steps use deterministic settings, and any remaining run-to-run variation is documented with the results.
- Common gaps
- An unseeded split, a GPU step without deterministic flags, or a library that varies even when seeded. Parallel workers without reproducible streams: a Warning.
- How to fix
- Take the seed from the command line, set every generator in one place, use reproducible parallel streams, and report the spread over seeds, as the NeurIPS checklist asks:
import argparse, random
import numpy as np
import torch
parser = argparse.ArgumentParser()
parser.add_argument("--seed", type=int, default=0)
args = parser.parse_args()
random.seed(args.seed)
np.random.seed(args.seed)
torch.manual_seed(args.seed)
torch.use_deterministic_algorithms(True) # error on nondeterministic opsTests and notebooks
- What counts
- A small test runs the analysis on example data in continuous integration, and result notebooks are saved after one clean run from top to bottom.
- Common gaps
- No tests, cells run out of order, or code that silently changes results: a quiet fallback, a disabled save, an unreachable branch.
- How to fix
- These are Warnings and never change the level. Restart the kernel and run each notebook top to bottom before saving, and make errors fail loudly.
Related
See where your project stands
Run the analysis on your paper or repository. Every finding names its area and links back here.