A reproducible record explains the environment and evaluation method; a matching prompt alone is insufficient.
Define what repetition should preserve
Reproduction can mean identical output bytes, similar measured performance or the same substantive conclusion. Those goals require different tests. AI systems may vary with model versions, sampling, software, hardware and retrieved information. PyTorch's reproducibility documentation explicitly warns that identical results are not guaranteed across every release or platform. Before comparing two runs, decide which differences matter and what tolerance the task can reasonably accept.
Build a hypothetical test packet
Imagine evaluating an agent that classifies fifty transaction descriptions. Save the exact inputs, expected labels, scoring rules, model identifier, relevant settings and execution date. If it retrieves outside information, preserve permitted source snapshots or stable references with observation times. Hashes can help detect changed files, but the files must remain available to an authorised reviewer. Publishing only a score and a prompt does not provide enough material to inspect the experiment.
Avoid selecting only the best run
Suppose three runs score 42, 45 and 43 correct answers. Reporting only 45 conceals variation. Present the full set or a predeclared summary and identify failures that carry disproportionate consequences. Keep a held-out evaluation set separate from examples used to adjust the system. A model that memorises familiar cases may look consistent without generalising. These are evaluation-design questions; attaching the results to a blockchain does not resolve them.
Make changes visible
Version the test packet and note changes to prompts, tools, datasets and scoring rules. Rerun relevant checks after a meaningful dependency update, and distinguish a reproduced historical result from a fresh assessment of a changed system. Do not put personal or confidential inputs into a public immutable record merely to make a benchmark inspectable. A useful reproducibility claim states exactly what another reviewer can repeat, which materials are available and where environmental variation or access restrictions limit the comparison.
Sources & context
Sources checked 7 October 2026. This is explanatory coverage, not personalised investment advice. Our corrections policy.