# Full-text source verification worksheets

Original BookTranslator observation data and a blank worksheet for keeping a paper's identity, copy version, terms, update checks, and actual file separate. Checked September 9, 2026. No paper text or source PDF is redistributed here.

- [Blank worksheet](full-text-worksheet.csv): one blank row; add one row for each copy, not merely one for each title.
- [Three checked examples](checked-examples.csv): three different scholarly articles, not three versions of one article and not a comparison of their scientific merit.

## Fill the worksheet

1. **Work:** title, authors, DOI or repository identifier. Match the record to the file, not just the search result.
2. **Copy:** preprint revision, accepted manuscript, published version, supplement, or unknown. Put the source of that label in `version_evidence`. Do not substitute a repository metadata revision number for a manuscript revision.
3. **Terms:** record the copy's license/access statement and its URL. A dataset or code license mentioned inside an article may not be the article's license. Record uncertainty rather than inferring reuse permissions.
4. **Updates:** save the source checked and what it actually reported. “Publisher retrieval failed” and “no notice found in the inspected record” are not equivalent to “no corrections exist.”
5. **File:** preserve record URL, actual file URL, retrieval time, local filename, page count and optionally SHA-256. A hash identifies bytes, not authenticity or authorization.
6. **Acceptance:** check the saved file's title/authors and relevant content, including appendices or separate supplements. Record what was checked and what remains open before choosing the next use.

Do not put signed download links, library-session tokens, private account details, or unpublished files into a shared worksheet. Keep private access records local; share stable public record URLs when appropriate.

## What the examples verify

Files were retrieved without account login using Python urllib, checked for a PDF header, opened with pypdf, counted, hashed with SHA-256, and inspected through their first and last rendered pages. We compared labels/identifiers against public repository or publisher records and inspected relevant license links. The main PLOS DOI's Crossref record confirmed its linked correction; Freeth's Crossref record corroborated publication metadata. This is a narrow source-file check, not an exhaustive content comparison, accessibility test, translation test, scientific review, or legal determination.

All three rows retain material limitations. Freeth's publisher page was not retrievable in the browser check, so we do not claim a complete publisher-side update check. The arXiv submission record is not proof that no later journal version exists. PLOS's correction link does not by itself establish which distributed PDFs include its revised figure.

An initial PeerJ candidate could not be downloaded because the public request returned HTTP 403. It was excluded from the three successful examples, not labeled paywalled or unavailable to everyone. Files and URLs may change after the observation date.

For each PDF fingerprint, rerun a download from its recorded public file URL and compute SHA-256 on the bytes. A different hash requires investigation; it is not automatically a substantive revision. These observations are a dated snapshot, not a live availability service.
