GPU Fleet

See and manage every GPU in your data center.

Organize nodes into groups, watch utilization and health in real time, and slice each GPU into right-sized VRAM partitions — so no card ever sits half-idle.

Nodes in groups

Group hosts by data center, rack, or purpose. Connect a node and its GPUs report in automatically.

Real-time health

Per-GPU utilization, VRAM usage, and temperature across the fleet, refreshed continuously.

No oversubscription

A per-GPU budget guarantees the sum of active slices never exceeds the physical VRAM.

Features

Built to run a GPU fleet

Fine-grained VRAM slicing

Carve a GPU into arbitrary-sized slices (e.g. 80 GB → 40 + 20 + 20) and run isolated workloads side by side.

Node grouping & onboarding

Create groups and connect new hosts; the node agent dials out and reports inventory and health.

Allocation tracking

Every slice records who holds it and what runs on it, backed by an append-only audit trail.

Health & operational alerts

Alerts for offline GPUs, overheating, downed nodes, and exhausted capacity.

Use cases

One fleet, many teams

Consolidate small services

Pack many small inference services onto shared GPUs instead of dedicating one card each.

Internal multi-team sharing

Allocate slices to teams and track usage for transparent internal chargeback.

Maximize expensive cards

Keep H100 and A100 capacity working by filling unused VRAM with additional pods.

Capacity at a glance

See free versus allocated VRAM across every node and GPU in one view.

FAQ

Common questions

How does VRAM slicing work?

Each GPU has a VRAM budget; you carve slices of any size as long as their sum stays within the total. Workloads run with a memory cap per slice. Stronger hardware isolation (MIG/MPS) is on the roadmap.

How do I add a node?

Create or pick a group, then connect the host. Its agent dials out to the control plane and reports the GPUs it sees, along with live health.

Can multiple workloads share one GPU?

Yes — that is the point. Several slices can run on one GPU at once, as long as their combined VRAM stays within the card’s budget.

What happens if a GPU goes offline?

Heartbeats stop, the node is marked offline, and an operational alert is raised so you can act before workloads are affected.

Ready to run your GPU fleet?

Slice your data-center GPUs into right-sized pods, provision on demand, and keep every gigabyte of VRAM working — all from one console.