Pods

Create a pod

Compute-fabric pod (BYO container image + SSH keys): a hardened GPU container on gpuCount whole native-driver GPUs of one node. Admission claims free device-request GPU units (see listPodCapacity for what is claimable), snapshots the on-demand rate immutably onto the pod, and drives the guest workflow. A valid request the fleet has no room for right now answers 503 with a capacity code; a GPU count no node could ever hold for the org answers 400 GPU_COUNT_NOT_ALLOCATABLE; and an org without the pod tier answers 409 FEATURE_DISABLED.

post/v1/pods

Request body

computeClass'on_demand' | 'reserved' | 'interruptible'

purchase option. on_demand (pay per second, no commitment) is the only class sold; reserved is a deprecated alias that is stored as on_demand. interruptible is retired and rejected with 400 COMPUTE_CLASS_RETIRED; it remains listed only until the deprecation window closes.

envVarsobject
gpuCountinteger

how many whole GPUs this pod gets. All of them are claimed on ONE node and passed to the container together; the pod is priced for gpuCount GPUs. Omitted = 1. A count no single node could ever hold for your org is refused with 400 GPU_COUNT_NOT_ALLOCATABLE; one that some node could hold but none has free right now is refused with 503 INSUFFICIENT_GPU_CAPACITY (maxGpuCount on /v1/pods/capacity is the largest a create can take now).

gpuModelIdstring
httpPortinteger

in-guest HTTP port exposed on the endpoint

imagestring required

tenant container image ref (required)

namestring
publicboolean

true = open endpoint (no data-plane auth); default false requires an org API key

shmSizeGbinteger

size of the pod's private /dev/shm in GiB (docker --shm-size). Optional; the maximum is half the pod's memory limit on its placement node (itself the pod's GPU-proportional share of node RAM), and a larger value is refused with 400 POD_ADMISSION_FAILED. Omitted lets the platform pick a safe default: that same maximum for a pod that owns its whole node, 2 GiB per GPU (capped at the maximum) for a shared multi-GPU pod, and docker's default for a shared single-GPU pod. Raise this when a multi-GPU NCCL / PyTorch DataLoader workload needs more shared memory than the default.

sshKeyIdsstring[] required

Org SSH key ids injected into the guest. At least one is required and every id must name a live key of this org: the data-plane gateway authorizes guest SSH against exactly this list, so a pod created without one has no SSH access. Attaching or detaching a key later takes effect within about a minute, with no restart.

volumeGbinteger

size in GB of the pod's persistent volume, mounted read-write at /workspace. Omitted = 50. The volume lives on the pod's node and is kept across stop, start, restart, container crashes, node reboots and interruptions; it is deleted only when the pod is deleted (or after 7 days interrupted with its node gone). Everything outside /workspace is the container's own filesystem and is reset whenever the container is recreated. Fixed at create; there is no resize.

Response

Created

capabilityTierstring

compute-fabric capability tier (e.g. "pod"); empty for a plain VM

computeClassstring

compute-fabric purchase option (on_demand; pods created before the rename carry the legacy reserved or interruptible values); empty for a plain VM

createdAtstring
diskSizeGbinteger
endpointUrlstring
gpuCountinteger
gpuManufacturerstring

GPU vendor from the catalog, the same value as GpuModel.manufacturer (NVIDIA, AMD); absent on a CPU-only row

gpuModelIdstring
gpuModelNamestring
idstring required
imageUrlstring

tenant container image (pods)

managedBystring

platform owner of this row ("serving" = inference serving replica, read-only on customer surfaces); empty for a customer-launched VM

namestring required
organizationIdstring required
pricePerHourCentsinteger
provisioningStagestring
publicboolean

true = open endpoint (no data-plane auth); false = require an org API key

resourceSizestring
serviceTypestring

managed-service tag; empty for a plain VM

statusstring required

deploying, running, stopped, failed, terminated, or interrupted; a row can also carry the queued status pending. interrupted means it served and is not serving now: it keeps its node, GPU and disk, and is not billed. A VM is interrupted while its node is offline and resumes when the node returns; a pod is interrupted for that and for losing its supervisor or its job on a live node, which the platform recovers by re-submitting it. After 7 days interrupted a VM is stopped (disk kept) and a pod is deleted.

statusReasonstring

Cause of the current status. The machine-readable values are "node_lost" (the host went offline and is expected back, so the disk is preserved), "node_removed" (the host was decommissioned, so the disk is gone and a restart deploys fresh elsewhere), "vm_exited", "user_stopped", "user_terminated", "user_restarted", "spot_reclaimed_by_capacity_owner", "spot_preempted_for_reservation_tenant", "auto_stopped_insufficient_balance", and the prefix "container failed to start: " followed by the container runtime's own message. Anything else is free text for a human to read.

tierstring
volumeGbinteger

pods only: size in GB of the pod's persistent /workspace volume.

Changes

Changed in 8 of the 36 revisions of this API.1110

    • ○

      added the non-success response with the status

      response-non-success-status-added

    • ○

      added the optional property to the response with the status

      response-optional-property-added

  • 74204a876c1a111See the full diff
    • ▲

      added the new required request property

      new-required-request-property

    • ●

      removed the request property

      request-property-removed

    • ○

      added the non-success response with the status

      response-non-success-status-added

    This revision also has 65 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog

    • ○

      added the new optional request property

      new-optional-request-property

    • ○

      added the optional property to the response with the status

      response-optional-property-added

    • ○

      added the new optional request property

      new-optional-request-property

    • ○

      the request property default value changed from interruptible to on_demand

      request-property-default-value-changed

    • ○

      added the new on_demand enum value to the request property

      request-property-enum-value-added

    • ○

      added the new optional request property

      new-optional-request-property

    • ○

      added the optional property to the response with the status

      response-optional-property-added