List Datasets
List the datasets an integration instance exposes, most-recently-seen first.
A dataset is one queryable unit of the integration — an index pattern for Elasticsearch/OpenSearch, an application for AppDynamics, a table for ServiceNow. Each entry carries when we last saw it and whether data has arrived for it recently.
Paginated because the count is unbounded in practice: the largest instance in production owns over 100,000 datasets. Use total_count for an "N indexes" label rather than the length of datasets.
search filters server-side on a case-insensitive substring of the dataset's stored key, narrowing total_count alongside the page. It does not match the rendered display_name, which is derived from the stored details and has no column to filter on; the two coincide for most integrations but not AppDynamics, whose key holds an application id while the row displays the application name.
dataset_term narrows to the rows carrying one source-system noun and narrows total_count too. It takes a key from this integration type's dataset_terms (see the integration catalog), which is also what each row reports as dataset_term_key — never the displayed wording, so the filter survives a rewording. log_group returns just the CloudWatch Logs half of an instance that also holds metric namespaces. A valid key belonging to another integration type selects nothing, and that empty page is the honest answer rather than an error.
status narrows to datasets that are available (data observed within the last 7 days) or not_available, and narrows total_count too. Availability is computed from the label-value store, not stored on the dataset row, so this filter is unknowable when that store cannot be reached — the same 503 as an unfiltered listing. The filters compose.
Permissions: users with integration-setup permission on the owning org.
Returns HTTP 404 if the instance does not exist in the resolved organization, and HTTP 503 (DATASET_AVAILABILITY_UNAVAILABLE) if the label-value store cannot be reached — availability is unknowable then, and reporting every dataset as not_available would be a false statement about the customer's data rather than about our own reachability.
Path parameters
Query parameters
A noun one integration's datasets are known by, in the source system.
The single identity of a noun: what a variant's term returns, what keys the selector that filters by it, and the value a client sends to filter by it. Each member carries its own spellings, so a noun cannot exist without them.
The enum value is the key rather than the singular, so filtering never depends on display wording and no value needs escaping in a query string.
Whether a dataset has label values in Doris within the freshness window.
NOT_AVAILABLE deliberately conflates "the source is empty" with "no population job has run for it": both present as an absence of Doris rows and telling them apart would need per-dataset job history we do not keep. A dataset's last_checked_at is the only hint — a months-old value points at the latter.
Response
Successful Response