Model Directory

One managed directory of models, ready to deploy.

Stage models in a managed directory and deploy any of them onto a GPU slice. What is in the directory is what is deployable — consistent across every node and team.

Managed directory

Models live in a directory the platform manages; the directory is the single source of what can be deployed.

One-step deploy

Pick a model and provision a pod for it in a single flow — onto a fresh slice or a GPU you choose.

Consistent across the fleet

The same model set is available to every node and every team, with no per-host setup.

Features

A model directory built for operators

Staged, ready-to-run models

Operators curate the directory; staged models are immediately available for provisioning.

Deploy onto any slice

Place a model on a right-sized VRAM slice on any GPU in the fleet.

Choose the serving runtime

Select the serving runtime for a deployment to match the model and the workload.

No per-deploy downloads

Models are staged once in the directory, so provisioning does not wait on large downloads.

Use cases

Standardize how models ship

One approved model set

Give every team the same curated directory instead of scattered, ad-hoc model files.

Fast provisioning

Skip download time at deploy — the model is already staged and ready.

Multiple models per GPU

Run different models on different slices of the same card.

Predictable rollouts

Promote a model into the directory once and deploy it consistently across the fleet.

FAQ

Common questions

What is the model directory?

A managed directory the platform deploys from. Models staged there are what operators can provision onto GPU slices.

How do models get into the directory?

Operators stage model data into the managed directory; from there it is available to provision across the fleet.

Can I run several models on one GPU?

Yes. Each model runs on its own VRAM slice, so multiple models can share a single GPU within its budget.

Does the platform train models?

No. NUSAPOD serves models that are already staged in the directory; building or training models is out of scope.

Ready to run your GPU fleet?

Slice your data-center GPUs into right-sized pods, provision on demand, and keep every gigabyte of VRAM working — all from one console.