---
title: "Augment Dataset"
method: POST
path: "/felix/dataset/augment"
tags: ["felix"]
---

# Augment Dataset

`POST /felix/dataset/augment`

Augment a dataset by removing duplicates/outliers and generating synthetic samples.

This endpoint creates a NEW dataset with the augmentations applied.
Operations include:
- **remove_duplicates**: Remove exact duplicate samples.
- **remove_outliers**: Remove samples with anomalous lengths.
- **balance**: Generate synthetic samples for underrepresented classes/entities.

The original dataset is preserved.

## Request body

- DatasetAugmentationRequest — Request for dataset augmentation.
  - `task_type` 'ner' | 'classification' | 'custom', required — Task type of the dataset
  - `dataset_name` string, nullable — Name of stored dataset to augment
  - `dataset_version` string, nullable — Dataset version (latest if omitted)
  - `dataset` object[], nullable — Inline dataset to augment (if dataset_name not provided)
  - `operations` AugmentationOperation[], required — List of augmentation operations to perform
    - `type` 'remove_duplicates' | 'remove_outliers' | 'balance', required — Type of augmentation operation
    - `enabled` boolean — Whether this operation is enabled
  - `target_distribution` object, nullable — Target class/entity distribution as percentages (must sum to 1.0)
  - `domain_description` string, nullable — Domain description for synthetic sample generation
  - `new_dataset_name` string, nullable — Name for the new augmented dataset (required if update_in_place=False)
  - `labels` string[], nullable — Labels for generation (required if balance operation used)
  - `update_in_place` boolean — If True, creates new version with same name and soft-deletes old version

## Response `200`

Successful Response

- DatasetAugmentationResponse — Response containing augmentation results.
  - `success` boolean, required
  - `original_dataset_name` string, nullable
  - `original_dataset_version` string, nullable
  - `new_dataset` DatasetResponse, required — Response model for a single dataset.
    - `id` string, required
    - `user_id` string, required
    - `dataset_name` string, required
    - `dataset_path` string, required
    - `dataset_type` string, required
    - `size` integer, nullable
    - `sample_size` integer, nullable
    - `train_ratio` number, nullable — Train split ratio for this dataset version. Left-to-right split with no shuffle; validation is the tail.
    - `created_at` string, required
    - `updated_at` string, required
    - `version_number` string
    - `root_dataset_id` string, nullable
    - `project_id` string, nullable
    - `schema` object, nullable
    - `schema_warnings` string[], nullable
    - `validation` object, nullable
    - `annotation_status` 'none' | 'in_progress' | 'completed', nullable
    - `annotation_config` object, nullable
    - `annotation_progress` object, nullable
    - `status` 'initialized' | 'uploading' | 'converting' | 'validating' | 'ready' | 'failed' | 'generating' | 'queued', nullable — Dataset status: initialized/uploading/converting/validating/ready/failed/generating/queued
    - `processing_error` string, nullable — Error message if status is failed
    - `type` string — Dataset purpose tag: 'training', 'evaluation', or 'benchmark'
    - `visibility` string — Dataset visibility: 'private' or 'public'
    - `is_competition` boolean — Whether this dataset is a competition benchmark
    - `labels` string[], nullable — Label names (entity types for NER, class labels for classification)
    - `generation_type` string, nullable — Canonical operation that created this dataset version.
    - `provenance` DatasetProvenance — Record how a dataset was created and which inputs produced it. Attributes: schema_version: Contract version for future migrations. method: Canonical dataset generation type. sources: External source descriptors. fallback: Fallback reason and attempt count, when one was required. source_dataset_ids: Dataset inputs combined or transformed. parent_dataset_id: Immediate dataset version or transform parent. transform_context: Bounded details about a transform. generator_context: Bounded generator configuration or agent context. synthesis_session_id: Synthesis-log session shared with the dataset row.
      - `schema_version` 1
      - `method` 'synthesize' | 'upload' | 'external' | 'grow' | 'augment' | 'version' | 'merge' | 'auto_relabel' | 'manual_relabel' | 'evaluation_suite' | 'agent_curated', required
      - `sources` DatasetProvenanceSource[]
        - `url` string, nullable
        - `revision` string, nullable
        - `license` string, nullable
        - `retrieved_at` string, date-time, nullable
        - `raw_hash` string, nullable
      - `fallback` DatasetProvenanceFallback — Describe a fallback used after the preferred source or strategy failed. Attributes: reason: Why the preferred path could not be used. attempts: Number of attempts made before the fallback succeeded.
        - `reason` string, required
        - `attempts` integer, required
      - `source_dataset_ids` string[]
      - `parent_dataset_id` string, nullable
      - `transform_context` JsonObjectOutput
      - `generator_context` JsonObjectOutput
      - `synthesis_session_id` string, nullable
    - `is_seed` boolean, nullable — Whether this dataset is a seed dataset (small set for review before full expansion)
    - `synthesis_session_id` string, nullable — UUID of the synthesis log session for this dataset, used to restore creation workflow on resume
    - `column_mapping` object, nullable — Column mapping from original to standard names
  - `modifications` ModificationSummary, required — Summary of modifications applied during augmentation.
    - `duplicates_removed` integer
    - `outliers_removed` integer
    - `samples_generated` integer
    - `original_count` integer
    - `final_count` integer
  - `distribution_comparison` DistributionComparison[], required
    - `label` string, required
    - `before_count` integer, required
    - `before_percentage` number, required
    - `after_count` integer, required
    - `after_percentage` number, required
  - `message` string, required
  - `updated_in_place` boolean — True if the original dataset was replaced (soft-deleted)

## Other responses

- `422` — Validation Error

## Changes

- **2026-09-24** `1cffaad2a921` — 1 warning, 2 info
  - removed the optional property `detail` from the response with the `422` status
  - added the optional property `new_dataset/provenance` to the response with the `200` status
  - added the required property `error` to the response with the `422` status

[Change history](https://skmtc.dev/pioneer/apis/brain-api/changes/felix/dataset/augment/post.md)

---

[API](https://skmtc.dev/pioneer/apis/brain-api.md) · [All operations](https://skmtc.dev/pioneer/apis/brain-api/llms.txt) · [OpenAPI document](https://skmtc.dev/pioneer/apis/brain-api/revisions/1cffaad2a921?raw)
