One work can have several public records

Researchers often release a manuscript on arXiv, update it after feedback, and later publish a journal version. arXiv preserves version history under one identifier, while Crossref commonly exposes a DOI record for the journal article. A search index should retain both sources: the preprint shows early public history and revisions, and the DOI record identifies the version of record and journal metadata.

Counting every source record as a new piece of research would make publication workflows look like new discoveries. Quantum Observatory instead stores each record with its own provenance and groups records into a work when an explicit DOI, arXiv identifier, or reviewed correction connects them. The work’s earliest known public date determines its weekly research count.

Why title matching is not enough

Titles change. A journal title may be shorter than a preprint title, punctuation can differ, and unrelated groups can use nearly identical phrases. Authors also change order or join later versions. A similarity score can suggest candidates for review, but automatically merging on title risks erasing distinct results and attaching the wrong publication history.

The observatory therefore accepts some temporary duplication rather than inventing a relationship. When a source exposes a reliable identifier link, the records can be grouped. A maintainer can also add a documented override after checking the original pages. The uncertainty stays visible until that evidence exists.

Grouping changes totals, not provenance

Suppose an arXiv manuscript appears in June, receives a second revision in July, and receives a DOI in September. The research-work series records one June work. The revision remains an event, and the journal article remains a searchable September source record, but neither becomes a second research work. A publication-focused filter can still find the journal record.

This design prevents double counting without pretending the records are identical. Abstract availability, author spelling, venue, citation metadata, and access links can differ. The detail panel shows linked records so readers can choose the version appropriate to their purpose.

  • Use the earliest explicit public date for the work-level timeline.
  • Retain later revisions as events rather than new works.
  • Keep each source URL and its retrieved metadata intact.
  • Expose manual merges as reviewable configuration.
  • Recalculate historical counts when a reliable link arrives later.

How to read a changing archive

A historical weekly count is not a sealed newspaper edition. An indexing service can add an older paper, or a DOI can reveal that two records represent one work. The archive is recalculated from the latest verified relationships, so an older total may increase or decrease. That is a correction to the model, not evidence that research occurred retroactively.

For reproducible work, export the relevant records and retain the run identifier and build identifier shown in the data manifest. A later reader can then distinguish a changed dataset from a changed interpretation. The linked-version policy makes the total more meaningful, while the retained source records make the transformation inspectable.