Provisioning & Agents

Provision pods on demand — or let agents do it.

Deploy a pod onto a GPU slice in seconds: empty for interactive work, or with an AI model from your model directory. Auto-provisioning picks the GPU and carves the slice; AI agents can request and release capacity on their own.

Empty or with a model

Provision a bare slice for interactive work, or deploy a model onto it in a single step.

Auto or manual

Let the platform pick a GPU and carve the slice, or target a specific GPU yourself.

Self-driving agents

AI agents request capacity when they need it and release it when idle — no tickets, no waiting.

Features

Provisioning that fits the work

Auto-provisioning

The platform finds an online GPU with enough free VRAM, carves a right-sized slice, and launches the pod.

Manual placement

Pin a workload to a specific GPU and pod when you need precise control over placement.

Empty pods

Spin up a ready slice without a model for bring-your-own or interactive use, then release it when done.

AI agent capacity

Agents provision and free pods programmatically, so capacity flows with demand instead of sitting idle.

Use cases

Capacity that moves with demand

Bursty agentic workloads

Agents scale capacity up under load and hand it back when the burst passes.

Interactive experiments

Grab an empty pod, do the work, and release the slice — no long-lived reservation.

One-step model serving

Provision a slice and deploy a model from the directory together in a single flow.

Targeted placement

Pin a latency-sensitive workload to a specific GPU for predictable performance.

FAQ

Common questions

What is a pod?

A workload that occupies one VRAM slice on a GPU. A pod can be empty (a ready slice) or run an AI model deployed onto it.

Auto-provisioning vs manual?

Auto-provisioning chooses an online GPU with enough free VRAM and carves the slice for you. Manual lets you target a specific GPU and pod.

How do AI agents get capacity?

Agents request a pod through the platform when they need to run work, and release it when idle — the fleet allocates and reclaims capacity automatically.

Where do the models come from?

Models are deployed from the managed model directory, so a pod-with-model can be launched without per-deploy downloads.

Ready to run your GPU fleet?

Slice your data-center GPUs into right-sized pods, provision on demand, and keep every gigabyte of VRAM working — all from one console.