Skip to content
Docs menu

Reproducibility levels

Each level says what the evidence supports, and each builds on the one below. A check names the evidence for every requirement it confirms.

Last updated

The five levels

Today a check can establish L0, L1 and L2. L3 and L4 need evidence from running the code and are locked until RepoReady can produce it.

  1. L4

    Robustness confirmed within tested scope

    Coming soon

    L3 plus a defined, accepted sensitivity-testing protocol. The conclusion is limited to the variations and targets actually tested.

  2. L3

    Results reproduced

    Coming soon

    L2 plus independent execution and agreement with the declared results. A script finishing successfully is not enough.

  3. L2

    Workflow documented

    Available now

    L1 plus complete instructions, versions, preparation and result mapping. Confirmed manuscript contradictions block this level.

  4. L1

    Materials available

    Available now

    The necessary code, inputs and software are accounted for within the assessed scope, with evidence of availability or a credible access route.

  5. L0

    Materials incomplete

    Confirmed material gap

    An essential material is confirmed missing, incomplete or unavailable within the assessed scope. Insufficient evidence alone does not establish L0.

When no level is set

? · Level not yet established

A question mark is not a failure. It means the evidence is not yet enough to set any level, for example when a dataset could not be checked. L0 needs a confirmed gap.

How a level is awarded

Materials and documentation are assessed independently, and the report counts their requirements: confirmed, gaps, needs evidence, not assessed. A level is awarded only when its own requirements and every level below it are confirmed. There is no combined score.

Example · complete documentation, missing input

  • Materials: gap confirmed
  • Documentation: confirmed

L0Materials incomplete

L2 requirements can be complete while L1 is unmet. The documentation result stays visible; L2 remains blocked until the material gap is resolved.

Good documentation stays visible, but it cannot lift a project past a material gap.

Supporting results, such as supplementary figures, count one level down: a missing script for a supplementary figure blocks L2, not L1, and gaps in their documentation never block L2.

What a finding blocks

Each finding says what it blocks. A finding marked Needs evidence is not settled yet; it never lowers a level.

  • Blocks L1A key result lacks its code, an input or the software it needs.
  • Blocks L2The workflow behind a key result, or its agreement with the paper, falls short.
  • Blocks automationRunning every step takes more than one locked, seeded command.
  • WarningWorth fixing, but blocks no level: a placeholder link in the docs, commented-out steps, scripts or files nothing uses, an external service, an undefined environment for the results, parallel workers without reproducible random streams. Journal policy and good practice findings are always Warnings.

Review scopes

Each check runs at a review scope. A Materials check awards at most L1 and marks the L2 requirements “Not assessed in this scope”. A Materials and workflow check also assesses documentation and automation, and compares paper and code: the methods, parameters, statistics and reported numbers of every key result. It can award L2.

Compare review scopes⁠

Levels and areas

Two areas decide the level: Materials for L1 and Documentation for L2. A confirmed contradiction with the manuscript also blocks L2. The other areas are reported beside the level.

All requirements by level⁠

What a check covers

A check reads the code, paper, data and documentation; the code is not executed. Analysis that starts from processed data does not establish how the raw data were prepared. A CSV can hold the scientific result: missing cosmetic plotting code alone is not a gap, but statistics inside a plotting script are part of the analysis.

Related frameworks

ACM and NISO recognise separate artifact and result properties; their badges are not a numbered ladder. These notes explain the overlap, not an equivalence or an external endorsement.

ACMArtifact Review and Badging
What the framework assesses
Artifacts Available; Artifacts Evaluated (Functional or Reusable); Results Validated (Reproduced or Replicated).
How RepoReady relates
L1 and L2 cover related material and documentation questions but do not establish ACM Functional, which requires exercising the artifact. L3 targets independent reproduction. Official guidance
NISORP-31-2021
What the framework assesses
Open Research Objects, Research Objects Reviewed, Results Reproduced and Results Replicated.
How RepoReady relates
L3 concerns the original artifacts; replication with independently developed artifacts is different. Controlled access does not meet the open research object criteria. Official guidance
TOPTransparency and Openness Promotion · 2025
What the framework assesses
Research practices at Disclosed, Shared and Cited, or Certified; computational reproducibility is a separate verification practice.
How RepoReady relates
Useful for journal acceptance policies. TOP’s levels describe policy implementation, not RepoReady levels; its computational verification relates to L3. Official guidance
IEEE AccessReproducibility initiative
What the framework assesses
Code Available and Code Reviewed, with execution in the review process.
How RepoReady relates
Separates access from exercised functionality. A RepoReady L1 or L2 does not confer the venue’s Code Reviewed badge. Official guidance
COSOpen Science Badges
What the framework assesses
Open Data, Open Materials and Preregistered.
How RepoReady relates
These recognise openness and registration. Neither they nor a RepoReady materials finding shows that results reproduce. Official guidance
FAIR4RS & RDA FAIRSoftware principles and data maturity
What the framework assesses
Findability, accessibility, interoperability and reusability; RDA provides indicators for assessing FAIR data.
How RepoReady relates
Inform identification, access, metadata and software reuse. FAIRness is not a reproduction level, and accessible does not always mean public. Official guidance · RDA FAIR
JOSSResearch software review
What the framework assesses
Software documentation, installation, examples, tests and verification of core functionality.
How RepoReady relates
A reference for the separate software reuse area. A usable package does not show that a paper’s experiments reproduce. Official guidance
OpenSSFBest Practices
What the framework assesses
Passing, Silver and Gold criteria for open-source engineering and security practices.
How RepoReady relates
Can inform software quality and reuse. Engineering achievements are not scientific readiness or reproduced findings. Official guidance

L4 is our proposed robustness extension, bound to an approved testing protocol. None of these frameworks offers a universal “robust results” tier that could simply be adopted.

See what your evidence supports

Run the analysis on your paper or repository and review each finding with its source.