Docs menu
- Home
- Docs
- Requirements
- Documentation
Documentation: a workflow others can follow
L2 Workflow documented adds a route others can follow: versions and identifiers, how inputs were prepared, which script makes which result, and instructions that match the repository.
Last updated
DocumentationRequired for L2
1 confirmed gap · 1 needs evidence · 6/8 confirmed
- Dependencies and versions are listedConfirmed
- The exact code version is identifiedConfirmed
- Data sources have stable identifiersConfirmed gap
- Data preparation steps are documentedConfirmed
- ?Code is linked to the paper’s resultsNeeds evidence
- Instructions explain how to run the analysisConfirmed
- Documented commands agree with the codeConfirmed
- Machine assumptions documentedConfirmed
Level relationshipRequired for L2
Documentation turns available materials into a workflow. Small omissions stop it: an unnamed version, a renamed script, a path that exists on one laptop. Of over 9,000 R files from Harvard Dataverse, 74% failed to complete without error (Trisovic et al., 2022). A Materials and workflow check assesses this area; Warnings here block no level.
In this guide
- Dependencies and versions are listed
- The exact code version is identified
- Versions of tools and models
- Data sources have stable identifiers
- Data preparation steps are documented
- Code is linked to the paper’s results
- Instructions explain how to run the analysis
- Documented commands agree with the code
- Machine assumptions documented
- A tidy repository
Dependencies and versions are listed
- What counts
- A standard file lists every package the code imports, each with a version or range:
requirements.txt,environment.yml,pyproject.toml,renv.lockorDESCRIPTION. - Common gaps
- No dependency file at all, imports that appear in no file, or a listed package without any version.
- How to fix
- Generate the list from a clean environment that runs your analysis, give every package a version, and update it in the same commit as the code.
The exact code version is identified
- What counts
- The paper cites an immutable version that contains the code behind every result: a release, tag, commit hash or version DOI.
- Common gaps
- Only a branch or a concept DOI, a cited tag that does not exist, or result code added after the cited version.
- How to fix
- Tag the version you submit, archive it so it gets its own version DOI, and cite that DOI in the paper:
git tag -a v1.0-paper -m "Code as submitted with the manuscript"
git push origin v1.0-paperVersions of tools and models
- What counts
- Forks, other repositories and models the code uses carry a fixed tag, commit or model revision, and every version the paper states matches the pins and logs.
- Common gaps
- A branch-only fork, a model without a revision, a Methods version the lock contradicts. Result files changed after the cited release: a Warning.
- How to fix
- Pin each fork and model to a commit or revision, copy versions into the Methods from the lock file, and cite a new release after changing result code.
Data sources have stable identifiers
- What counts
- Each dataset has an accession, DOI, versioned record or release URL; live services carry a release or query date; controlled data say where and how to apply.
- Common gaps
- A dataset named only in prose, a cloud-drive link, a placeholder DOI, a link to the wrong file, or controlled data with no procedure.
- How to fix
- List each dataset with source, identifier, version and access in the README, pin models to a revision, and replace personal links. The Materials guide table works.
Data preparation steps are documented
- What counts
- Every input that shared code does not produce, deposited processed data included, comes with its steps: tool, version, non-default parameters and any manual curation.
- Common gaps
- A table edited by hand, a file only commented-out code writes, or a model without its training code. Commented-out steps elsewhere: a Warning.
- How to fix
- Make each step a script or workflow rule, re-enable or remove commented-out steps, and describe any manual step exactly:
rule filter_cells:
input: "data/raw/counts.h5ad"
output: "data/processed/filtered.h5ad"
params: min_genes=200
shell: "python -m pipeline.filter {input} {output} --min-genes {params.min_genes}"
rule train_model:
input: "data/processed/filtered.h5ad"
output: "models/classifier.pt"
shell: "python -m pipeline.train {input} {output} --seed 0"Code is linked to the paper’s results
- What counts
- A table or entry links each key figure, panel, table and reported number to the one script, notebook section or command that produces it.
- Common gaps
- No map at all, a map that skips some figures, or
analysis_v2.pybesideanalysis_final.py. Scripts nothing names: a Warning. - How to fix
- Add a results table to the README, one row per paper item, keep one copy of each script, and say what any other script is for:
## Results
| Paper item | Command | Output |
|------------|------------------------------------|----------------------------------|
| Figure 2 | Rscript analysis/figure2.R | results/figure2/figure2.csv |
| Figure 3B | python -m analysis.fig3 --panel b | results/figure3/panel_b.csv |
| Table 2 | python -m analysis.evaluate_cohort | results/claims/table2_auroc.json |Instructions explain how to run the analysis
- What counts
- The README, paper or linked docs cover setup, where inputs go or how they are fetched, the run order and every manual step, or that there is none.
- Common gaps
- No README, data read from a folder outside the repository that no doc names, or docs promising an automatic download that never happens.
- How to fix
- Write a short “How to reproduce” section: setup, where each input goes, then each step with its command and the outputs it creates, in order.
Documented commands agree with the code
- What counts
- Every command, file, option, variable, folder and install name the docs give matches the repository exactly, and result notebooks are saved without errors.
- Common gaps
- A renamed script,
data.csvforData.csv.gz, an install name taken by another package, or leftover<your-path>text. A TODO link in prose: a Warning. - How to fix
- In a fresh clone, check every name the README uses, create the folders it writes to, replace placeholders and save notebooks without errors.
Machine assumptions documented
- What counts
- Result code builds paths from the project folder or a setting, has no unexplained host or system check, and the README states GPU, memory and run time.
- Common gaps
- Paths like
/Users/you/data, a script that sources your own shell profile or refuses other systems. Unstated GPU or memory needs: a Warning. - How to fix
- Build paths from the project root, let data locations be overridden, and note hardware needs and run time in the README:
from pathlib import Path
import os
ROOT = Path(__file__).resolve().parents[1] # repository root, from src/pipeline.py
DATA = Path(os.environ.get("DATA_DIR", ROOT / "data"))
counts = DATA / "raw" / "counts.h5ad"A tidy repository
- What counts
- Only what the analysis needs: no system or editor junk, scratch copies, duplicates, large binaries, files your
.gitignoreexcludes, unused functions or committed keys. - Common gaps
- Functions nothing calls, files the analysis never reads (one Warning per repository), or an API key in the code, which only you see.
- How to fix
- Warnings here never change the level. Add a
.gitignore, remove what no result needs, and revoke any committed key and purge it from the history.
Related
- Materials: code, inputs and software
- Automation: one command, locked and seeded
- Pre-submission checklist
See where your project stands
Run the analysis on your paper or repository. Every finding names its area and links back here.
On this page
- At a glance
- Dependencies and versions are listed
- The exact code version is identified
- Versions of tools and models
- Data sources have stable identifiers
- Data preparation steps are documented
- Code is linked to the paper’s results
- Instructions explain how to run the analysis
- Documented commands agree with the code
- Machine assumptions documented
- A tidy repository
- Related