Record the question before the query
A reproducible comparison begins with a sentence that defines the population and measure. For example: compare the number of grouped research works tagged error correction across two complete weeks in the configured corpus. This statement fixes the unit, tag, periods, and data boundary before the result is known.
Avoid beginning with a surprising percentage and then searching for a definition that supports it. If the question changes during exploration, save the new definition with the final output.
Capture the snapshot identity
Every built snapshot exposes a run identifier and build identifier in its manifest. The run identifies the collection attempt; the build identifier also reflects the dataset, application, pipeline, and configuration used to render the site. Save both with the date of access.
The content hash identifies normalized records, while the code and collection-code revisions distinguish software used for collection and presentation. If a later export differs, these fields help determine whether new records, relationship corrections, rules, or interface code changed.
Export data and preserve the transformation
Use the Research page to select the date basis, period, topic, material type, and source. Export the resulting CSV and store the visible filter values beside it. For a programmatic analysis, retain the summary, index, and relevant detail shards rather than assuming the live endpoints will remain unchanged forever.
Describe any additional grouping or exclusion in executable code or an unambiguous formula. If you manually inspect records, keep the source URLs and a reason for every exclusion. Do not replace missing dates with the collection date unless your question explicitly concerns discovery.
- Write the research question and unit of analysis.
- Save run_id, build_id, content hash, and access date.
- Record every interface filter or query parameter.
- Retain source URLs and identifiers used to resolve duplicates.
- Report unavailable coverage and undated exclusions.
Explain what can still change
A reproduced output can be correct for its snapshot and still differ from a later site view. Late indexing can add an older work, a reliable DOI link can merge two records, and a classifier update can alter topic membership. These are versioned corrections rather than reasons to hide the difference.
Publish the absolute counts with the percentage, state that topics overlap, and link the methodology used. If the conclusion depends on only a few records, name and inspect them. A reader should be able to repeat the calculation and also see why the available corpus may not represent the entire field.