gateway.openapi_Gateway

CRUD APIs for deployment shape. Create Deployment Shape

post/v1/accounts/{account_id}/deploymentShapes

Path parameters

account_idstring required

The Account Id

Query parameters

deploymentShapeIdstring

The ID of the deployment shape. If not specified, a random ID will be generated.

Request body

namestring
displayNamestring

Human-readable display name of the deployment shape. e.g. "My Deployment Shape" Must be fewer than 64 characters long.

descriptionstring

The description of the deployment shape. Must be fewer than 1000 characters long.

createTimestring date-time

The creation time of the deployment shape.

updateTimestring date-time

The update time for the deployment shape.

baseModelstring required
modelTypestring

The model type of the base model.

parameterCountstring int64

The parameter count of the base model .

acceleratorCountinteger

The number of accelerators used per replica. If not specified, the default is the estimated minimum required by the base model.

acceleratorType'ACCELERATOR_TYPE_UNSPECIFIED' | 'NVIDIA_A100_80GB' | 'NVIDIA_H100_80GB' | 'AMD_MI300X_192GB' | 'NVIDIA_A10G_24GB' | 'NVIDIA_A100_40GB' | 'NVIDIA_L4_24GB' | 'NVIDIA_H200_141GB' | 'NVIDIA_B200_180GB' | 'AMD_MI325X_256GB' | 'AMD_MI350X_288GB'
precision'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
disableDeploymentSizeValidationboolean

If true, the deployment size validation is disabled.

enableAddonsboolean

If true, LORA addons are enabled for deployments created from this shape.

draftTokenCountinteger

The number of candidate tokens to generate per step for speculative decoding. Default is the base model's draft_token_count.

draftModelstring

The draft model name for speculative decoding. e.g. accounts/fireworks/models/my-draft-model If empty, speculative decoding using a draft model is disabled. Default is the base model's default_draft_model. this behavior.

ngramSpeculationLengthinteger

The length of previous input sequence to be considered for N-gram speculation.

enableSessionAffinityboolean

Whether to apply sticky routing based on user field.

numLoraDeviceCachedinteger
maxContextLengthinteger

The maximum context length supported by the model (context window). If set to 0 or not specified, the model's default maximum context length will be used.

presetType'PRESET_TYPE_UNSPECIFIED' | 'MINIMAL' | 'FAST' | 'THROUGHPUT' | 'FULL_PRECISION' | 'AGENTIC_CODING' | 'CHAT' | 'SUMMARIZATION'

Response

A successful response.

namestring
displayNamestring

Human-readable display name of the deployment shape. e.g. "My Deployment Shape" Must be fewer than 64 characters long.

descriptionstring

The description of the deployment shape. Must be fewer than 1000 characters long.

createTimestring date-time

The creation time of the deployment shape.

updateTimestring date-time

The update time for the deployment shape.

baseModelstring required
modelTypestring

The model type of the base model.

parameterCountstring int64

The parameter count of the base model .

acceleratorCountinteger

The number of accelerators used per replica. If not specified, the default is the estimated minimum required by the base model.

acceleratorType'ACCELERATOR_TYPE_UNSPECIFIED' | 'NVIDIA_A100_80GB' | 'NVIDIA_H100_80GB' | 'AMD_MI300X_192GB' | 'NVIDIA_A10G_24GB' | 'NVIDIA_A100_40GB' | 'NVIDIA_L4_24GB' | 'NVIDIA_H200_141GB' | 'NVIDIA_B200_180GB' | 'AMD_MI325X_256GB' | 'AMD_MI350X_288GB'
precision'PRECISION_UNSPECIFIED' | 'FP16' | 'FP8' | 'FP8_MM' | 'FP8_AR' | 'FP8_MM_KV_ATTN' | 'FP8_KV' | 'FP8_MM_V2' | 'FP8_V2' | 'FP8_MM_KV_ATTN_V2' | 'NF4' | 'FP4' | 'BF16' | 'FP4_BLOCKSCALED_MM' | 'FP4_MX_MOE'
disableDeploymentSizeValidationboolean

If true, the deployment size validation is disabled.

enableAddonsboolean

If true, LORA addons are enabled for deployments created from this shape.

draftTokenCountinteger

The number of candidate tokens to generate per step for speculative decoding. Default is the base model's draft_token_count.

draftModelstring

The draft model name for speculative decoding. e.g. accounts/fireworks/models/my-draft-model If empty, speculative decoding using a draft model is disabled. Default is the base model's default_draft_model. this behavior.

ngramSpeculationLengthinteger

The length of previous input sequence to be considered for N-gram speculation.

enableSessionAffinityboolean

Whether to apply sticky routing based on user field.

numLoraDeviceCachedinteger
maxContextLengthinteger

The maximum context length supported by the model (context window). If set to 0 or not specified, the model's default maximum context length will be used.

presetType'PRESET_TYPE_UNSPECIFIED' | 'MINIMAL' | 'FAST' | 'THROUGHPUT' | 'FULL_PRECISION' | 'AGENTIC_CODING' | 'CHAT' | 'SUMMARIZATION'

Changes

Changed in 6 of the 27 revisions of this API.1716

  • 3c3ed322dd5233See the full diff
    • ●

      added the new AGENTIC_CODING enum value to the response property for the response status

      response-property-enum-value-added

    • ●

      added the new CHAT enum value to the response property for the response status

      response-property-enum-value-added

    • ●

      added the new SUMMARIZATION enum value to the response property for the response status

      response-property-enum-value-added

    • ○

      added the new AGENTIC_CODING enum value to the request property

      request-property-enum-value-added

    • ○

      added the new CHAT enum value to the request property

      request-property-enum-value-added

    • ○

      added the new SUMMARIZATION enum value to the request property

      request-property-enum-value-added

    • ○

      added the new optional request property

      new-optional-request-property

    • ○

      added the optional property to the response with the status

      response-optional-property-added

    This revision also has 5 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog

  • 4e1bdf0cae5611See the full diff
    • ●

      added the new AMD_MI350X_288GB enum value to the response property for the response status

      response-property-enum-value-added

    • ○

      added the new AMD_MI350X_288GB enum value to the request property

      request-property-enum-value-added

  • 4506ec9d12ff126See the full diff
    • ▲

      removed the enum value REINFORCEMENT_FINE_TUNING of the request property

      request-property-enum-value-removed

    • ●

      deleted the query request parameter disableSizeValidation

      request-parameter-removed

    • ●

      added the new FULL_PRECISION enum value to the response property for the response status

      response-property-enum-value-added

    • ○

      added the new optional request property

      new-optional-request-property

    • ○

      the request optional property became not read-only

      request-optional-property-became-not-read-only

    • ○

      added the new FULL_PRECISION enum value to the request property

      request-property-enum-value-added

    • ○

      added the optional property to the response with the status

      response-optional-property-added

    • ○

      the response optional property became not read-only for the status

      response-optional-property-became-not-read-only

    • ○

      removed the REINFORCEMENT_FINE_TUNING enum value from the response property for the response status

      response-property-enum-value-removed

    • ●

      added the new REINFORCEMENT_FINE_TUNING enum value to the response property for the response status

      response-property-enum-value-added

    • ○

      added the new REINFORCEMENT_FINE_TUNING enum value to the request property

      request-property-enum-value-added

    • ○

      the endpoint scheme security BearerAuth was added to the API

      api-security-added

    • ○

      api tag gateway.openapi_Gateway added

      api-tag-added

    • ○

      api tag Gateway removed

      api-tag-removed

    This revision also has 2 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog