Docs

Run GPU workloads

Run an application with GPU access on Northflank cloud or your own infrastructure. Services and jobs can use GPUs. Use a service for a continuous process and a job for work that ends.

Requirements

You will need the following to get started:

  • A container image that supports the intended GPU workload
  • The GPU type, memory, and resource requirements of the application
  • Access to a deployment location with the required GPU capacity

1. Choose where the workload runs

Hosting choiceFirst stepDetailed guide
Northflank cloudCreate a GPU-enabled projectDeploy GPUs on Northflank cloud
Your cloudPrepare a cluster and GPU node poolsDeploy GPUs in your own cloud

Review availability, requirements, and supported configuration in the selected guide before creating resources. The hosting paths have different GPU configuration constraints.

For your cloud, review the provider requirements and node pool configuration. Make sure that the cluster has GPU capacity before deploying the workload.

2. Configure the service or job

Use the selected hosting guide to choose the image and GPU resources. The shared configuration guidance is under Compute → GPUs.

Add the application's runtime variables and secrets. If the image needs a different start command, override its command or entrypoint.

For an HTTP service, configure its port. A batch job does not need public HTTP access to perform its work.

3. Run a representative task

Examine the container logs and the application's own GPU diagnostics. Make sure that the application can use the selected device and complete its intended task.

Use metrics to review resource use. For your cloud, also review node pool capacity and cluster operation in the provider guide.

Next steps

For model serving, continue with Host a model on GPUs. For tasks that end or use a schedule, read Run background tasks.

© 2026 Northflank Ltd. All rights reserved.

northflank.com / Terms / Privacy / feedback@northflank.ai