---
title: "Pull Dataset From Hub"
method: POST
path: "/felix/datasets/pull-from-hub"
tags: ["felix", "datasets"]
---

# Pull Dataset From Hub

`POST /felix/datasets/pull-from-hub`

Pull a dataset from HuggingFace Hub and save it locally.

User-authenticated requests require hf_token. Sandbox run keys omit it and
import public repos anonymously. Gated or private repos must be imported
in the Pioneer UI with a user token.
Optionally accepts session_id for SSE log streaming.

## Request body

- HuggingFacePullRequest — Request to pull dataset from HuggingFace Hub.
  - `repo_id` string, required — HuggingFace repo ID to pull from
  - `config_name` string, nullable — HuggingFace dataset configuration, distinct from the local Pioneer name
  - `revision` string, nullable — Specific revision/branch to pull
  - `name` string, nullable — Name for the local dataset (defaults to repo name)
  - `hf_token` string, nullable — HuggingFace API token. Required for user-authenticated imports. Sandbox run keys must omit this field and can only import public repos. Gated or private repos must be imported in the Pioneer UI.
  - `session_id` string, nullable — Session ID for SSE log streaming
  - `column_mapping` object, nullable — Column mapping from source to standard names
  - `dataset_type` string, nullable — Dataset type (ner, classification, custom)
  - `type` 'training' | 'evaluation' | 'benchmark', nullable — Dataset purpose: 'training' (trainable), 'evaluation' (not trainable), 'benchmark' (system-managed, evaluation-only; cannot be created via this endpoint). This model is also bound to POST /felix/datasets/preview-from-hub, which ignores this field entirely -- nothing is persisted on a preview.

## Response `200`

Successful Response

- DatasetResponse — Response model for a single dataset.
  - `id` string, required
  - `user_id` string, required
  - `dataset_name` string, required
  - `dataset_path` string, required
  - `dataset_type` string, required
  - `size` integer, nullable
  - `sample_size` integer, nullable
  - `train_ratio` number, nullable — Train split ratio for this dataset version. Left-to-right split with no shuffle; validation is the tail.
  - `created_at` string, required
  - `updated_at` string, required
  - `version_number` string
  - `root_dataset_id` string, nullable
  - `project_id` string, nullable
  - `schema` object, nullable
  - `schema_warnings` string[], nullable
  - `validation` object, nullable
  - `annotation_status` 'none' | 'in_progress' | 'completed', nullable
  - `annotation_config` object, nullable
  - `annotation_progress` object, nullable
  - `status` 'initialized' | 'uploading' | 'converting' | 'validating' | 'ready' | 'failed' | 'generating' | 'queued', nullable — Dataset status: initialized/uploading/converting/validating/ready/failed/generating/queued
  - `processing_error` string, nullable — Error message if status is failed
  - `type` string — Dataset purpose tag: 'training', 'evaluation', or 'benchmark'
  - `visibility` string — Dataset visibility: 'private' or 'public'
  - `is_competition` boolean — Whether this dataset is a competition benchmark
  - `labels` string[], nullable — Label names (entity types for NER, class labels for classification)
  - `generation_type` string, nullable — Canonical operation that created this dataset version.
  - `provenance` DatasetProvenance — Record how a dataset was created and which inputs produced it. Attributes: schema_version: Contract version for future migrations. method: Canonical dataset generation type. sources: External source descriptors. fallback: Fallback reason and attempt count, when one was required. source_dataset_ids: Dataset inputs combined or transformed. parent_dataset_id: Immediate dataset version or transform parent. transform_context: Bounded details about a transform. generator_context: Bounded generator configuration or agent context. synthesis_session_id: Synthesis-log session shared with the dataset row.
    - `schema_version` 1
    - `method` 'synthesize' | 'upload' | 'external' | 'grow' | 'augment' | 'version' | 'merge' | 'auto_relabel' | 'manual_relabel' | 'evaluation_suite' | 'agent_curated', required
    - `sources` DatasetProvenanceSource[]
      - `url` string, nullable
      - `revision` string, nullable
      - `license` string, nullable
      - `retrieved_at` string, date-time, nullable
      - `raw_hash` string, nullable
    - `fallback` DatasetProvenanceFallback — Describe a fallback used after the preferred source or strategy failed. Attributes: reason: Why the preferred path could not be used. attempts: Number of attempts made before the fallback succeeded.
      - `reason` string, required
      - `attempts` integer, required
    - `source_dataset_ids` string[]
    - `parent_dataset_id` string, nullable
    - `transform_context` JsonObjectOutput
    - `generator_context` JsonObjectOutput
    - `synthesis_session_id` string, nullable
  - `is_seed` boolean, nullable — Whether this dataset is a seed dataset (small set for review before full expansion)
  - `synthesis_session_id` string, nullable — UUID of the synthesis log session for this dataset, used to restore creation workflow on resume
  - `column_mapping` object, nullable — Column mapping from original to standard names

## Other responses

- `422` — Validation Error

## Changes

- **2026-09-24** `1cffaad2a921` — 1 warning, 6 info
  - removed the optional property `detail` from the response with the `422` status
  - added the new optional request property `config_name`
  - added the new optional request property `type`
  - the request property `hf_token` became optional
  - …3 more

[Change history](https://skmtc.dev/pioneer/apis/brain-api/changes/felix/datasets/pull-from-hub/post.md)

---

[API](https://skmtc.dev/pioneer/apis/brain-api.md) · [All operations](https://skmtc.dev/pioneer/apis/brain-api/llms.txt) · [OpenAPI document](https://skmtc.dev/pioneer/apis/brain-api/revisions/1cffaad2a921?raw)
