---
title: "Related papers derived from category + co-author heuristic"
method: GET
path: "/api/v1/research/arxiv/recommendations"
tags: ["Research"]
---

# Related papers derived from category + co-author heuristic

`GET /api/v1/research/arxiv/recommendations`

Return papers related to a seed paper from arXiv.org (Cornell University). arXiv does not expose a citation graph via the public API, so Sugra compiles recommendations from the seed paper's primary category combined with co-author overlap - two signals commonly used as citation proxies. The seed paper itself is excluded from the result set. Pagination via `limit` and `offset`.

## Query parameters

- `paper_id` string, required — Seed arXiv identifier (modern `YYMM.NNNNN` or legacy `archive/YYMMNNN`).
- `limit` integer — Papers per page (1 to 2000).
- `offset` integer — Zero-based offset into the result set (0 to 10000).

## Response `200`

Related papers derived from the seed paper's category and co-author graph, seed-paper excluded.

- EnvelopeArxivRecommendationsPayload
  - `data` ArxivRecommendationsPayload, required — Response payload for `/recommendations`.
    - `paper_id` string, required — Seed paper identifier echoed from the request.
    - `seed_paper` ArxivPaper — One paper (one Atom `<entry>`) returned by the arXiv API.
      - `arxiv_id` string, nullable — Canonical arXiv identifier without version suffix (e.g. `2303.08774` or legacy `hep-th/9901001`).
      - `version` integer, nullable — Version number (1-based) of this revision (e.g. 1 for the first upload, 2 for v2). Null when not known.
      - `entry_id` string, nullable — Full `<id>` URL for the entry, including version suffix (e.g. `http://arxiv.org/abs/2303.08774v1`).
      - `title` string, nullable — Paper title as submitted, whitespace-normalized.
      - `summary` string, nullable — Abstract text, whitespace-normalized. Contains raw LaTeX math (e.g. `$O(n \log n)$`) when the author used math mode.
      - `authors` ArxivAuthor[] — Ordered list of authors as supplied on submission.
        - `name` string, nullable — Author full name as supplied on submission, whitespace-normalized (case preserved).
        - `affiliation` string, nullable — Author affiliation string when provided on the submission (often null).
      - `primary_category` string, nullable — arXiv category code the author chose as primary (e.g. `cs.AI`, `q-fin.ST`, `hep-th`).
      - `categories` string[] — All arXiv category codes attached to this paper (primary plus cross-lists).
      - `published` string, nullable — ISO 8601 UTC timestamp of the original (v1) submission. Immutable after submission.
      - `updated` string, nullable — ISO 8601 UTC timestamp of the latest revision. Equals `published` when the paper has only v1.
      - `doi` string, nullable — DOI string (e.g. `10.1143/PTP.101.1155`) when the paper has been published in a journal with a registered DOI; null otherwise.
      - `journal_ref` string, nullable — Freeform journal reference string as filed by the author (e.g. `Prog.Theor.Phys.101:1155-1164,1999`).
      - `comment` string, nullable — Author-supplied comment on the submission (e.g. `Accepted at NeurIPS 2023`, `15 pages, 3 figures`); null when absent.
      - `urls` ArxivUrls — External URLs attached to each arXiv entry.
        - `abstract` string, nullable — Canonical HTML abstract page on arxiv.org (e.g. `https://arxiv.org/abs/2303.08774v1`).
        - `pdf` string, nullable — Direct link to the arXiv-hosted PDF of this version.
        - `doi` string, nullable — DOI URL when the paper has a registered DOI (e.g. `https://doi.org/10.1143/PTP.101.1155`); null otherwise.
    - `strategy` string, required — Heuristic used to compile recommendations (e.g. `primary_category+coauthor`).
    - `papers` ArxivPaper[], required — Related papers derived from the seed's primary category and co-author overlap (seed paper itself excluded).
      - `arxiv_id` string, nullable — Canonical arXiv identifier without version suffix (e.g. `2303.08774` or legacy `hep-th/9901001`).
      - `version` integer, nullable — Version number (1-based) of this revision (e.g. 1 for the first upload, 2 for v2). Null when not known.
      - `entry_id` string, nullable — Full `<id>` URL for the entry, including version suffix (e.g. `http://arxiv.org/abs/2303.08774v1`).
      - `title` string, nullable — Paper title as submitted, whitespace-normalized.
      - `summary` string, nullable — Abstract text, whitespace-normalized. Contains raw LaTeX math (e.g. `$O(n \log n)$`) when the author used math mode.
      - `authors` ArxivAuthor[] — Ordered list of authors as supplied on submission.
        - `name` string, nullable — Author full name as supplied on submission, whitespace-normalized (case preserved).
        - `affiliation` string, nullable — Author affiliation string when provided on the submission (often null).
      - `primary_category` string, nullable — arXiv category code the author chose as primary (e.g. `cs.AI`, `q-fin.ST`, `hep-th`).
      - `categories` string[] — All arXiv category codes attached to this paper (primary plus cross-lists).
      - `published` string, nullable — ISO 8601 UTC timestamp of the original (v1) submission. Immutable after submission.
      - `updated` string, nullable — ISO 8601 UTC timestamp of the latest revision. Equals `published` when the paper has only v1.
      - `doi` string, nullable — DOI string (e.g. `10.1143/PTP.101.1155`) when the paper has been published in a journal with a registered DOI; null otherwise.
      - `journal_ref` string, nullable — Freeform journal reference string as filed by the author (e.g. `Prog.Theor.Phys.101:1155-1164,1999`).
      - `comment` string, nullable — Author-supplied comment on the submission (e.g. `Accepted at NeurIPS 2023`, `15 pages, 3 figures`); null when absent.
      - `urls` ArxivUrls — External URLs attached to each arXiv entry.
        - `abstract` string, nullable — Canonical HTML abstract page on arxiv.org (e.g. `https://arxiv.org/abs/2303.08774v1`).
        - `pdf` string, nullable — Direct link to the arXiv-hosted PDF of this version.
        - `doi` string, nullable — DOI URL when the paper has a registered DOI (e.g. `https://doi.org/10.1143/PTP.101.1155`); null otherwise.
    - `pagination` ArxivPagination, required — Pagination counters surfaced by the OpenSearch block in the Atom feed.
      - `total_results` integer, nullable — Total matching papers across the entire arXiv corpus for this query (from `<opensearch:totalResults>`).
      - `start_index` integer, nullable — Zero-based offset of the first result on this page (echoed from upstream).
      - `items_per_page` integer, nullable — Page size used by upstream for this response.
      - `limit` integer, required — Page size requested by the caller.
      - `offset` integer, required — Zero-based offset requested by the caller.
  - `meta` SugraMeta, required — Metadata on a /api/v1/* response envelope built through `helpers.response.sugra_response`, which is how routes are expected to answer. A route that assembles its own `meta` dict carries only the keys it writes itself, so an optional field below can be absent because this response has nothing to report OR because that route does not build its envelope here - the two are not distinguishable from the outside (API-43).
    - `endpoint` string, required — Requested endpoint path.
    - `data_time` string, required — ISO 8601 timestamp the data on this response is stamped with. It is the source's own timestamp whenever the source supplied one this API could read; when it did not, this field falls back to the value of `response_time` and `data_age_days` is omitted, so the PRESENCE of that field is the signal to read - with the one exception named in its own description, a route that substitutes its own current time for a source timestamp it never received. Usually UTC (`Z`), but a source stating its own numeric offset keeps it (2026-04-16T14:30:00+09:00) rather than being converted a second time. For a source that publishes by period this is the period's START (see `period`) and for one that publishes by calendar day it is that day's midnight - in neither case a moment at which anything was observed or released.
    - `response_time` string, required — ISO 8601 UTC timestamp when this response was produced.
    - `provider` string, required — API name and version.
    - `data_age_days` number, nullable — Age of the data in days at the moment this response was produced, i.e. `response_time` minus `data_time`. Present ONLY when the timestamp this response is stamped with is a clock time that could be read as a real instant. It is ABSENT - never 0 - in every other case. Absent when no readable source timestamp was supplied, because `data_time` then repeats `response_time` and a zero age would assert that the data is current precisely where its true age is unknown. Absent when the source names a calendar day, a month, a quarter or a year (see `period`): the instant is then a boundary this API anchored at midnight, and time since a day or a quarter BEGAN is a different quantity from the age of the data - a daily series is out by up to a day, a quarterly one by up to a quarter. A midnight counts as such a boundary whichever zone it is stated in, and whether the source stated it or this API anchored it. The one case this field cannot see is a route that substitutes its own current time for a source timestamp it never received: the substituted value is a real, readable instant and is indistinguishable from one the source stated, so the age reads as roughly 0. The shared cache-and-fetch helper behind most routes stopped doing that (API-43), but the presence of this field is a statement about the timestamp the response carries, not a guarantee about the route that supplied it. Rounded to 0.001 day (86.4 seconds), so 0.0 is a real measured age anywhere within roughly +/-43 seconds and not a stand-in for unknown; a source stamping an instant in the future reports a negative value (-0.001 or less) rather than being clamped. Sources publish on very different cadences, so a non-zero age is normal, not an error. Preserve absence in client code: a generated client that materialises a missing optional number as its numeric default turns 'age unknown' back into 'age zero', which is the exact confusion this field exists to remove.
    - `source` string, nullable — Identifier of the primary upstream source used for this response.
    - `attribution` string, nullable — Human-readable attribution mandated by an upstream source (e.g. a securities regulator or self-regulatory organization). Present only on responses whose source requires the owner and source to be clearly identified. Do not remove or alter it when using the response.
    - `fallback_used` boolean, nullable — True when the primary source failed and a fallback produced the data.
    - `fallback_chain` string[], nullable — Ordered list of sources attempted, in the order they were tried.
    - `cached` boolean, nullable — True when this response was served from the internal cache.
    - `stale` boolean, nullable — True when the cached response was returned after the upstream rate-limited or errored. Clients can use this to detect degraded data.
    - `period` string, nullable — Unit of observation, when the source publishes by period rather than by instant. `data_time` carries the period's START instant so it stays machine-readable; this field preserves what that instant used to mean, which the conversion would otherwise erase. Present only for such sources, and only when the source hands the API the label itself - a client that converts the period to its start instant before building the envelope loses the label, though not the age exclusion, which is decided by the instant. Note that `data_age_days` is omitted whenever this is present, because an age measured from a period start is not a freshness figure.
    - `notes` string, nullable — Data-quality caveat about THIS response - how old the underlying report is, a chokepoint AIS lower-bound, or that a source-reported `data_time` could not be read and the response time is shown instead. Distinct from `attribution`, which is a licensing obligation. Multiple caveats are joined with ' | '. Present only when there is one.

## Other responses

- `401` — Missing or invalid `x-api-key` header. JSON body with a stable `code` distinguishing `missing_api_key` (no header sent) from `invalid_api_key` (header sent, key not accepted); any other 401 source carries the generic `unauthorized` with its detail as `reason`. Plus `hint`. `plan` is always null on 401 - an unauthenticated request has no plan; quota exhaustion is 429, not 401.
- `422` — Validation Error
- `429` — Daily rate limit exceeded. Check `X-RateLimit-Reset` for the next window.
- `503` — Upstream source is temporarily unavailable. Retry after a short delay.

## Changes

- **2026-09-01** `328d061c12ca` — 3 info
  - added the optional property `meta/data_age_days` to the response with the `200` status
  - added the optional property `meta/notes` to the response with the `200` status
  - added the optional property `meta/period` to the response with the `200` status
- **2026-08-08** `4c4530760ba1` — 12 info
  - added the optional property `code` to the response with the `401` status
  - added the optional property `code` to the response with the `429` status
  - added the optional property `code` to the response with the `503` status
  - added the optional property `hint` to the response with the `401` status
  - …8 more

[Change history](https://skmtc.dev/sugra/apis/sugra-api/changes/api/v1/research/arxiv/recommendations/get.md)

---

[API](https://skmtc.dev/sugra/apis/sugra-api.md) · [All operations](https://skmtc.dev/sugra/apis/sugra-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/sugra/sugra-api/revisions/cdcc60731935/schema)
