Rivul · Legal

Data Sources

Last updated: August 27, 2026

Why this page exists

Rivul does not write your citations. References come from the scholarly metadata systems below, formatted by a deterministic rule, so a citation can be traced to the index that produced it. This page lists every one of those systems, the licence Rivul reads it under, and what Rivul keeps. Nothing here is paid for and no scholarly source charges Rivul for access.

Semantic Scholar (Allen Institute for AI)

Used for literature search, citation counts, abstracts and similar-paper recommendations.

Read under the Semantic Scholar API licence, which requires attribution to Semantic Scholar and a link back to the paper. Every search result card carries that link. The underlying graph data is generally ODC-BY; individual records may carry their own terms, and open-access PDF links must be checked at the source.

Rivul stores title, authors, year, venue, abstract, DOI, arXiv id and citation count for papers you save to your library. Search responses are cached for 24 hours and recommendations for 7 days, then discarded.

arXiv

Used for preprint search and for fetching open-access PDFs.

arXiv metadata is CC0 1.0; the e-prints themselves stay under their authors' own licences, so a PDF Rivul fetches for you is the author's work under the author's terms, not Rivul's. Thank you to arXiv for use of its open access interoperability.

arXiv asks for no more than one request every three seconds across all machines a service controls. Rivul holds itself to that rather than treating each server as a fresh allowance: requests are serialised and each one has to claim a three-second slot from shared state. Query results are cached for a day, which arXiv's manual asks integrators to do.

Rivul stores the metadata, the extracted text of any PDF you fetch (which is what grounds chat and citation checking), and the PDF file itself in Rivul's storage so your library keeps working offline and reopening a paper does not re-download it. Deleting the library entry deletes both.

Crossref

Used for DOI lookup, publisher metadata, and retraction and preprint status.

Crossref asserts no licence over the metadata facts it distributes. Its API asks callers to identify themselves with a contact address, which puts them in the polite pool with a published concurrency limit; Rivul identifies itself and stays inside that limit.

Rivul stores the reference fields a citation needs — DOI, title, authors, year, venue, volume, issue, pages, publisher, type — and a retraction flag on library entries.

Crossref Retraction Watch

Used to flag retracted papers in your library.

Crossref acquired the Retraction Watch database and publishes it as a CC0 dataset, rebuilt every working day. Rivul reads the whole dataset once a night and compares it against the DOIs in its users' libraries, rather than polling for one DOI at a time, so every paper in a library is checked every night regardless of when it was added.

Rivul stores only the resulting flag. The dataset itself is streamed, matched and discarded; no copy is kept.

OpenAlex (OurResearch)

Used as a supplementary index for search, citation counts and abstracts.

OpenAlex data is CC0, which requires no attribution; it is credited here because citing a source is the practice this product is for. The API is credit-metered, so OpenAlex is treated as a bonus rather than a dependency: when its budget is exhausted, search continues from the other sources and the interface says results may be partial instead of showing a false "no matches".

Rivul stores citation counts and abstracts on saved library entries.

DataCite

Used as a fallback for DOIs that Crossref does not hold — arXiv DOIs, datasets and software are registered with DataCite.

DataCite metadata is CC0.

Rivul stores the same reference fields as for Crossref.

Unpaywall (OurResearch)

Used to find a legal open-access copy of a paper by DOI.

Unpaywall's data is CC0. Its API requires a real contact email, and Rivul sends one; when none is configured Rivul skips the lookup rather than presenting an address that does not receive mail.

Rivul stores nothing from Unpaywall beyond a cached answer to "is there an open-access copy of this DOI", kept for a week for a hit and a day for a miss.

What none of these sources receive

Your drafts, your notes, your annotations and your account details are never sent to any of the sources above. What leaves Rivul for them is a search query, a DOI, or an arXiv identifier.

Corrections

If you maintain one of these services and something on this page is wrong or out of date, write to support@rivul.ai and it will be corrected.