---
title: "Split text on a regex"
method: POST
path: "/ai/async-chunking/regex-splitter/{MODEL_ID}"
tags: ["Split content into chunks"]
---

# Split text on a regex

`POST /ai/async-chunking/regex-splitter/{MODEL_ID}`

The regex-splitter chunker (chunking strategy) splits the submitted text based on the specified regex (regular expression), according to the conventions employed by the `re` python package. For more information about the `re` operations, see https://docs.python.org/3/library/re.html.

## Headers

- `Authorization` string, required
- `Content-Type` string

## Request body

- RegexRequest
  - `batch` object[] — The batch of key:value pairs used in the chunking request.
    - `text` string — The content to be split into chunks based on the chunking strategy (`chunker`) sent in the request. The maximum text size allowed for input is approximately 1 MB.
  - `modelConfig` ModelConfig — Provides fields and values that specify ranges for tokens.
    - `vectorQuantizationMethod` string — Vector quantization compresses data size, as well as reducing memory usage. The methods are: * `min-max` - Creates tensors of the text and converts it to uint8 by normalizing it to the range [0, 255]. * `max-scale` - Finds the maximum absolute value for the encoded text, normalizes it by scaling the text to a range of -127 to 127, and then returns the quantized text as an 8-bit integer tensor.
    - `dimReductionSize` integer — Used to reduce vector size while maintaining good quality. This field allows any integer above 0, but less than or equal to the vector dimension of the model. If you send a vector dimension larger than the model, a 400 Bad Request error is returned. Not every model is designed to support this parameter. In this scenario, a warning message is generated that indicates quality can decrease.
  - `useCaseConfig` UseCaseConfigChunking
    - `dataType` string — This optional parameter enables model-specific handling in the Async Chunking API to help improve model accuracy. Use the most applicable fields based on available dataTypes and the dataType value that best aligns with the text sent to the Async Chunking API. The string values to use are: "dataType": "query" for the query. "dataType": "passage" for fields searched at query time.
  - `chunkerConfig` RegexSplitterChunkerConfig — The regex-splitter chunker (chunking strategy) splits the submitted text based on the specified regex (regular expression), according to the conventions employed by the `re` python package. This is the default chunker configuration if nothing is passed. For more information about the `re` operations, see https://docs.python.org/3/library/re.html.
    - `regex` string — This field sets the regular expression used to split the provided text.

## Response `200`

OK

- POSTresponse — This is the response to the POST chunking request submitted for a specific `chunker` and `modelId`.
  - `chunkingId` string, uuid — The universal unique identifier (UUID) returned in the POST request. This UUID is required in the GET request to retrieve results.
  - `status` string — The current status of the request. Allowed values are: * SUBMITTED - The POST request was successful and the response has returned the `chunkingId` and `status` that is used by the GET request. * ERROR - An error was generated when the GET request was sent. * READY - The results associated with the `chunkingId` are available and ready to be retrieved. * RETRIEVED - The results associated with the `chunkingId` are returned successfully when the GET request was sent.

## Other responses

- `4XX` — The error varies based on the issue encountered regarding the `chunkingId` or related information.

---

[API](https://skmtc.dev/lucidworks/apis/rules-rewrites-api.md) · [All operations](https://skmtc.dev/lucidworks/apis/rules-rewrites-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/lucidworks/rules-rewrites-api/revisions/f2d3747848e8/schema)
