

How to optimise GPU utilisation and reduce cloud costs
- GPU cost optimization starts with measuring utilisation, identifying idle capacity, and matching GPU types and quantities to workload requirements.
- Scheduling, right-sizing, auto-shutdown, and shared GPU pools can increase utilisation and reduce wasted capacity.
- Showback and chargeback make GPU costs visible across teams and help organisations control shared infrastructure spend.
- Northflank GPU workloads run on per-second billing with no minimum, autoscale by GPU utilisation metrics, and support BYOC into your own cloud account where reserved capacity and committed spend apply.
Northflank provides GPU workloads on H100, H200, A100, L4, L40S, B200 and more with per-second billing, autoscaling, and bring your own cloud (BYOC). Get started (self-serve) or book a demo.
GPU compute is expensive. An H100 at $2.74 per hour running at 15% SM utilisation costs the same as one running at 95%. The difference between those two numbers is not a pricing problem. It is a utilisation problem. Most organisations running GPU workloads are paying for more compute than their workloads actually use, and standard cloud cost reports often fail to show where that waste comes from.
GPU cost optimization requires more than CPU right-sizing or autoscaling. Model loading, memory constraints, distributed training, and idle development environments create GPU-specific sources of waste. This guide covers how to measure GPU utilisation, identify wasted capacity, and improve GPU efficiency.
CPU utilisation is a relatively direct measure of whether a workload is using the compute it was allocated. Low CPU utilisation on a running instance usually means the workload is idle, making it easier to identify capacity that can be downsized or shut down.
GPU utilisation is more complex. GPU memory can determine whether a workload can run independently of compute utilisation, while large models can keep GPUs allocated during model loading or between requests. Distributed training adds another constraint because all GPUs required by a job may need to be available at the same time. These characteristics mean standard cloud cost reports can miss significant GPU waste, making GPU-specific monitoring and optimisation controls necessary.
Improving GPU efficiency requires more than simply increasing utilisation. You need to understand what is consuming GPU capacity, match workloads to appropriate hardware, schedule workloads efficiently, and reclaim capacity when it is no longer needed.
The metrics that matter for GPU cost optimization are not the same as the metrics cloud provider dashboards surface by default.
SM (streaming multiprocessor) utilisation shows how much of the GPU's compute capacity is actively being used. GPU memory utilisation shows how much GPU memory is allocated, while memory bandwidth utilisation helps identify workloads that are limited by memory rather than compute. Idle time shows how long allocated GPUs are sitting without doing useful work.
These metrics should be considered together. A GPU can have low SM utilisation while using most of its memory, such as when a model is loaded and waiting for requests. Looking at compute utilisation alone can therefore make an expensive GPU appear more efficient than it actually is.
Match the GPU type and quantity to the workload instead of defaulting to the most powerful hardware available. Measure GPU memory and compute requirements on representative workloads, then choose the smallest configuration that meets your performance requirements.
Also separate baseline capacity from peak demand. Predictable workloads can use committed or reserved capacity, while interruptible workloads such as batch processing and checkpointed training can use more flexible capacity. Right-sizing the GPU and the number of GPUs can reduce costs before you make any changes to scheduling or autoscaling.
How workloads are placed across your GPU infrastructure directly affects utilisation. Bin-packing can consolidate workloads onto fewer nodes, leaving unused nodes available for scale-down instead of spreading workloads across partially utilised nodes.
For smaller workloads, GPU partitioning or time-slicing can allow multiple workloads to share a physical GPU. Distributed training introduces another challenge: jobs often need multiple GPUs to become available at the same time. Gang scheduling can prevent partially allocated jobs from holding GPUs idle while waiting for the rest of their capacity.
Idle GPUs are one of the clearest sources of wasted spend. Development environments, notebooks, completed jobs, and inference services can continue consuming GPU capacity long after useful work has stopped.
Use idle-time monitoring, TTL policies, scheduled scale-down, and automatic shutdown to reclaim unused capacity. For production inference, scale GPU replicas based on actual demand while accounting for model loading time and latency requirements. This avoids keeping more GPU capacity running than the workload needs.
Shared GPU pools can improve utilisation by allowing teams to use capacity when they need it instead of maintaining separate dedicated GPUs. Resource quotas prevent one team or project from consuming the entire pool.
Use project or team-level quotas and workload priorities to balance development, training, and production workloads. For large training jobs, planned reservation windows can also help coordinate GPU capacity without allowing one workload to block the rest of the pool.
GPU cost allocation connects infrastructure spending to the teams using it. Showback makes GPU costs visible without directly charging them to team budgets, while chargeback allocates those costs to the responsible team or project.
For either model, track GPU usage by team, project, workload, GPU type, and effective rate. Metrics such as cost per training run, cost per inference request, and idle GPU cost make the data actionable and help teams identify inefficient workloads.
Together, these controls turn GPU cost optimization from a pricing exercise into an ongoing utilisation practice. Instead of simply buying cheaper GPUs, teams can reduce waste by ensuring the GPUs they already pay for are being used effectively.
Northflank GPU workloads run on H100, H200, A100, L4, L40S, and B200 hardware with per-second billing and no minimum allocation. The billing model eliminates the idle cost of hourly minimums: a job that runs for 8 minutes bills for 8 minutes.
- Autoscaling: GPU services on Northflank support autoscaling by CPU, memory, requests per second, or custom Prometheus metrics. For inference services, scaling based on request queue depth or P95 latency scales replicas in response to actual load rather than predicted load. Scale-to-zero can eliminate GPU compute billing during idle periods when the workload has a mechanism to detect new demand and restart capacity.
- Scheduled jobs: GPU jobs on Northflank run to completion and release GPU allocation automatically. There is no ongoing cost after the job finishes. Scheduled training jobs, batch inference runs, and evaluation jobs run on their schedule and terminate, without holding GPU capacity between runs.
- Spot GPU orchestration: BYOC node pools on Northflank support spot instance configuration. Checkpointable training workloads can target spot GPU nodes at significantly reduced cost. Northflank orchestrates workload placement across available capacity and handles spot interruption by rescheduling on available nodes.
- BYOC for reserved capacity: Northflank BYOC deploys into your own cloud account, where GPU reserved instances, committed use discounts, and capacity reservations negotiated with your cloud provider apply directly to Northflank-managed workloads. Your existing GPU cost commitments reduce the effective rate Northflank workloads pay.
- Observability: GPU utilisation metrics SM utilisation, memory utilisation, memory bandwidth are available through the Northflank observability layer alongside service logs, build metrics, and infrastructure events. Log sinks export to your existing observability stack for fleet-wide GPU cost dashboards.
- RBAC and project isolation : GPU workloads are scoped to projects, which maps to team boundaries for cost attribution. Project-level resource controls define which teams can launch GPU workloads and at what scale, providing the quota enforcement layer for multi-tenant GPU pools.
Request GPU capacity for planned allocations. Get started on Northflank (self-serve) or book a demo.
GPU cost optimization is primarily a utilisation problem, not a pricing problem. Measure utilisation, right-size workloads, schedule GPUs efficiently, eliminate idle capacity, and use quotas and cost allocation to make shared infrastructure more efficient.
Northflank helps teams put these practices into place with flexible GPU workloads, autoscaling, per-second billing, and BYOC support. This makes it easier to match GPU capacity to workload demand, reclaim unused resources, and manage GPU infrastructure costs as your workloads grow.
GPU utilisation measures how much of a GPU's compute capacity is actively being used by a workload. SM utilisation is the primary compute metric, but it does not tell the whole story. A GPU can have low SM utilisation while consuming most of its memory, such as when a model is loaded and waiting for requests. For accurate monitoring, consider SM utilisation alongside GPU memory utilisation, memory bandwidth, and idle time.
Low GPU utilisation can come from several sources, including idle development environments, oversized GPU allocations, workloads waiting on data, inference services running below their expected traffic levels, and distributed training jobs waiting for enough GPUs to start. Each requires a different solution, such as auto-shutdown, right-sizing, better workload placement, autoscaling, or gang scheduling.
Showback makes GPU costs visible to each team without directly charging those costs to their budget. Chargeback goes further by allocating GPU costs to the team, project, or business unit responsible for the usage. Both approaches require accurate cost allocation data, such as GPU-hours, GPU type, workload, and effective rate.
Use MIG when workloads need stronger isolation and predictable performance. MIG divides a supported GPU into isolated instances with dedicated memory and compute resources. Time-slicing allows multiple workloads to share a GPU without the same level of isolation and is better suited to development, experimentation, and workloads that can tolerate variable performance.
Multiply the number of GPUs used by the number of hours the workload ran and the effective GPU rate. For example, eight H100s running for 12 hours at $2.74 per hour would cost $263.04. For spot workloads, use the effective spot rate instead of the on-demand rate when calculating the actual cost.


