Run an application with GPU access on Northflank cloud or your own infrastructure. Services and jobs can use GPUs. Use a service for a continuous process and a job for work that ends.
You will need the following to get started:
- A container image that supports the intended GPU workload
- The GPU type, memory, and resource requirements of the application
- Access to a deployment location with the required GPU capacity
1. Choose where the workload runs
| Hosting choice | First step | Detailed guide |
|---|---|---|
| Northflank cloud | Create a GPU-enabled project | Deploy GPUs on Northflank cloud |
| Your cloud | Prepare a cluster and GPU node pools | Deploy GPUs in your own cloud |
Review availability, requirements, and supported configuration in the selected guide before creating resources. The hosting paths have different GPU configuration constraints.
For your cloud, review the provider requirements and node pool configuration. Make sure that the cluster has GPU capacity before deploying the workload.
2. Configure the service or job
Use the selected hosting guide to choose the image and GPU resources. The shared configuration guidance is under Compute → GPUs.
Add the application's runtime variables and secrets. If the image needs a different start command, override its command or entrypoint.
For an HTTP service, configure its port. A batch job does not need public HTTP access to perform its work.
3. Run a representative task
Examine the container logs and the application's own GPU diagnostics. Make sure that the application can use the selected device and complete its intended task.
Use metrics to review resource use. For your cloud, also review node pool capacity and cluster operation in the provider guide.
Next steps
For model serving, continue with Host a model on GPUs. For tasks that end or use a schedule, read Run background tasks.