Kubernetes HPA Scale-to-Zero: When Idle Savings Break Latency and Reliability

September 30, 2026

A practical look at scale-to-zero in Kubernetes 1.37: metric requirements, end-to-end wake-up latency, the real conditions for savings, KEDA and managed-platform comparisons, failure modes, and a validation plan before production rollout.

Changing minReplicas from 1 to 0 is technically straightforward. The complexity begins after the first idle period: a job arrives, but there is nothing available to process it. Component availability now depends not only on the application and Kubernetes, but also on metric freshness, the HPA loop, the scheduler, node availability, image pulls, and readiness checks.

This is not a routine replica-count optimization. Scale-to-zero changes the failure model: the demand signal becomes part of the critical path, and cold starts become part of user-visible latency.

In Kubernetes 1.37, the HPAScaleToZero feature reached Beta and is enabled by default. That does not make it a universal way to reduce cluster costs. A durable queue consumer that can tolerate a cold start is a strong candidate. A synchronous HTTP API, where the first request must both signal demand and wait for the application to become ready, is not.

Scale-to-zero starts with the service model, not HPA

Before configuring autoscaling, answer two questions:

  1. Where does work reside while no Pods are ready?

  2. How long can it remain there without violating the SLA?

For an asynchronous worker, the answer may be a durable queue. The job has already been accepted and persisted, and it can wait for a consumer. A standard Kubernetes Service has no built-in queue that retains an HTTP request until a Ready Pod appears.

A Kubernetes Service does not become a queue when Ready Pods disappear. It does not retain a request until the application starts. Therefore, scale-to-zero for an HTTP endpoint without an external buffer, a proxy with explicit waiting semantics, or an always-running replica creates a window for errors and timeouts. This is fundamentally different from platforms where request acceptance and instance startup are combined into a managed model (Kubernetes Blog, KEP-2021).

It is useful to distinguish three different outcomes:

  • zero Pods — the application no longer consumes requested resources at the scheduler level;

  • fewer nodes — autoscaling or consolidation has actually removed paid infrastructure;

  • serverless behavior — the platform accepts an event or request, persists it, and orchestrates runtime startup itself.

Native HPA provides the first mechanism. The second depends on node autoscaling and the placement of other workloads. The third must be designed separately.

What HPA in Kubernetes 1.37 actually adds

HPA can use minReplicas: 0 only when at least one Object or External metric is configured. A configuration that uses only CPU or memory is rejected by the API server (HPA documentation).

This is not a cosmetic restriction: with zero Pods, there is no CPU or memory consumption from which to calculate whether startup is needed. The signal must exist independently of the workload being scaled. Typical examples include queue depth, the age of the oldest job, or another external demand indicator.

Automatic zero and manual pause are different states

Historically, replicas: 0 could mean that a workload had been manually stopped. To prevent HPA from unexpectedly restarting it, Kubernetes distinguishes between:

  • zero reached by HPA itself;

  • zero set manually or by an older controller.

When HPA scales a workload to zero automatically, it sets the ScaledToZero=True condition. If the workload is at replicas: 0 without that condition, HPA preserves the semantics of a manual pause and is not required to scale it back up.

This matters for troubleshooting. The existence of an HPA object and an external metric does not prove that a component will automatically recover from zero. An operational runbook should check both the actual replica count and HPA conditions.

Control-plane component version skew

During an upgrade, the API server may already accept an HPA with minReplicas: 0 while the older kube-controller-manager cannot yet scale a workload up from zero. The old controller interprets replicas: 0 as a manual pause.

A safe rollback or feature-gate disablement therefore requires preparatory steps:

  1. change affected HPAs to minReplicas >= 1;

  2. ensure every workload has scaled up to at least one replica;

  3. only then roll back the controller or disable the feature.

The documentation describes the risk itself, but managed Kubernetes still requires an upgrade and rollback rehearsal: access to feature gates and control-plane flags may depend on the provider.

Four autoscaling models solve different problems

Model Transition from zero What happens to the first request or job Failure characteristics
HPA with minReplicas: 1 Not required Processed by a warm replica Ongoing cost, but fewer dependencies in the critical path
Native HPA 1.37 Based on an Object or External metric Must wait in an external buffer or queue A metric failure may leave the workload at zero
KEDA The operator manages 0↔1; HPA manages 1↔N Depends on the event source and its semantics Activation thresholds, fallback behavior, periodic polling, and activity notifications
Cloud Run or Lambda Managed by the platform The platform handles part of the waiting and startup process Provider-specific limits, pricing, and redelivery rules apply

Why minReplicas: 1 remains a strong default

For many SaaS components, one small always-running replica costs less than a complex metrics path. It removes cold starts from the normal request path, reduces dependence on the metrics adapter, and shrinks the operational recovery surface.

Scale-to-zero is especially low-value when idle periods are frequent and short. The component has time to terminate, only to start again soon afterward; latency and operational complexity increase, while the node may remain billable anyway.

Native HPA and KEDA

Native HPA requires an Object or External metric to already be published through the Kubernetes Metrics API. When Prometheus is the source, this usually requires an adapter that translates the HPA query into PromQL.

KEDA uses a different architecture. Its operator determines whether a source is active and manages 0↔1 transitions. The HPA object created by KEDA then scales the workload in the 1↔N range. KEDA provides many ready-made event-source integrations as well as an external gRPC scaler (KEDA concepts, External scalers).

With KEDA, you can separate:

  • the activation threshold — whether demand is sufficient to leave zero;

  • the scaling threshold — how many replicas are required after startup.

For example, with a scaling threshold of 10 and an activation threshold of 50, a value of 40 leaves the workload at zero. This can avoid starting an expensive worker for a single job, but a poorly chosen threshold can delay work even when a normal HPA calculation would already recommend scaling up.

KEDA does not become obsolete with Kubernetes 1.37. Native HPA standardizes scale-to-zero, while KEDA remains useful for ready-made integrations, push activation, source-specific thresholds, and fallback behavior when metrics fail.

Managed platforms

Cloud Run can scale to zero by default and independently connects request queueing with instance startup. When the maximum instance count is reached, a request waits for a limited time and may then receive a 429 response. Minimum instances reduce cold-start latency, but they are billable and remain a best-effort mechanism rather than an absolute guarantee of available capacity (Cloud Run autoscaling, minimum instances).

AWS Lambda illustrates a different warm-reserve model: Provisioned Concurrency initializes execution environments in advance for more predictable latency. Even so, an SQS integration still requires coordination among the function timeout, visibility timeout, retries, and DLQ (Lambda lifecycle, SQS event source mapping).

These platforms remove some work from the team, but they do not eliminate the need to design delivery and retry behavior. They are also not direct replacements for Kubernetes: their execution model, constraints, and pricing differ.

End-to-end wake-up latency

The documented default HPA control-loop interval is 15 seconds, configured through --horizontal-pod-autoscaler-sync-period. But a 15-second cold start should not be assumed.

A practical decomposition looks like this:

T_wake ≈
  T_metric_freshness
  + T_wait_next_HPA_sync
  + T_API/Deployment
  + T_schedule
  + T_image/init
  + T_readiness
  + T_node_provision, if no suitable node is available

This is an engineering model, not a Kubernetes SLA.

It includes:

  • time until the source metric is updated;

  • waiting for the next HPA cycle;

  • changing the desired replica count through the API and Deployment controller;

  • Pod scheduling;

  • image pulls and init container execution;

  • application, JVM, or model initialization;

  • readiness probe completion;

  • provisioning of a new node when resources are insufficient.

There is no universal p95 or p99 for this chain. It depends on image size, registry location, scheduler contention, probe configuration, application initialization, and node-pool state. It must be measured in the target cluster.

Why the metrics path often dominates

If HPA receives metrics through Prometheus, latency begins before the HPA cycle. Prometheus's global scrape_interval defaults to one minute unless configured otherwise (Prometheus configuration). The Prometheus Adapter then translates an External metric request into PromQL according to externalRules and metricsQuery.

Errors in labels, namespace mapping, or the query can result in a missing or incorrect signal (External Metrics in Prometheus Adapter). As a result, HPA may not receive a scale-up signal and the Pod may not start even though jobs are already accumulating.

So, “HPA checks every 15 seconds” does not mean “the worker wakes up in 15 seconds.” The metric may become visible later, while node creation and application initialization may add most of the latency.

KEDA also polls scalers periodically through pollingInterval. An external scaler that sends activity notifications can use a long-lived StreamIsActive call to report activation outside the regular polling interval. Metric caching reduces load on the source and helps meet its query-rate limits, but changes the trade-off between signal freshness and metrics-system resilience (ScaledObject specification).

Zero Pods does not guarantee zero cost

The scheduler and node autoscaling primarily rely on resource requests, not actual CPU and memory usage. Removing a Pod releases requested resources, but saves money only when that makes it possible to remove or replace a billable node (Node Autoscaling).

There may be no economic benefit if:

  • the node pool has a nonzero minimum;

  • other Pods remain on the node;

  • DaemonSets or placement constraints prevent consolidation;

  • released resources are too fragmented;

  • a GPU node cannot be removed or returned quickly;

  • paid capacity is reserved regardless of utilization.

For evaluation, it is useful to separate cost into three parts:

TotalCost = IdleCapacityCost + ActivationCost + ReliabilityCost

Where:

  • IdleCapacityCost is the cost of Pods, nodes, or GPUs kept continuously during idle time;

  • ActivationCost is the cost of cold starts and temporary capacity for bursts;

  • ReliabilityCost is the cost of SLO violations, retries, duplicates, and operational support.

Compare minReplicas: 0 and minReplicas: 1 using the real distribution of idle intervals and bursts, not average utilization. You need prices for the actual resources, minimum node-group sizes, DaemonSet overhead, and real consolidation history.

The metrics path becomes part of the availability loop

With a single external metric, its unavailability can leave a workload at zero. HPA sets ScalingActive=False, for example with the FailedGetExternalMetric reason, but that status does not itself create reserve capacity.

If HPA uses multiple metrics and one cannot be retrieved, it does not scale down when the available metrics recommend a reduction. It can still scale up if at least one available metric requires growth. This is useful protection against incorrect scale-down, but it does not solve the problem of a single external signal while the workload is at zero.

For such a component, Prometheus, the adapter, and the corresponding query effectively become part of the production path even though they do not process jobs themselves. They should be observed as production dependencies:

  • source metric freshness;

  • HPA query failures against the Metrics API;

  • adapter latency and errors;

  • age of the oldest queued job;

  • divergence between queue depth and replica count;

  • duration of ScalingActive=False.

When its source fails, KEDA keeps the current instance count by default. For supported trigger types, after a configured number of consecutive errors, the fallback mechanism can synthetically set HPA to a fixed replica count; fallback is not supported for CPU or memory triggers. This does not guarantee continuously available capacity, and its behavior must be verified for the selected trigger type, installed KEDA version, and corresponding CRD. With native HPA, equivalent protection must be built separately or manual recovery must be planned (KEDA troubleshooting).

Scale-to-zero does not solve delivery or interrupted work

Removing a Pod is safe only if the worker can stop correctly. During termination, Kubernetes sends the process a stop signal, waits for the grace period, and may then terminate it forcibly (Pod Lifecycle).

A correct queue worker should:

  1. stop accepting new jobs once termination begins;

  2. complete or roll back the current operation;

  3. persist the result in durable storage;

  4. acknowledge the message only afterward;

  5. tolerate redelivery without repeating irreversible effects.

The typical semantics of such systems are at-least-once, not exactly-once. For example, when SQS/Lambda batch processing fails, messages are returned for redelivery. Without partial batch response, records that have already completed successfully may also be processed again.

Depending on broker semantics and the consumer protocol, a practical pattern may include:

  • a durable queue;

  • a visibility timeout or lease aligned with processing time;

  • acknowledgement after the result is committed;

  • an idempotency key;

  • a retry policy;

  • a DLQ;

  • graceful shutdown handling on SIGTERM.

HPA controls replica count. Broker semantics and the consumer protocol determine whether data is lost or duplicated.

Which components are suitable for scale-to-zero

Component type Assessment Primary condition
Batch processing with long idle periods Strong candidate Jobs are persisted and can tolerate waiting
Report generation Strong candidate The user does not wait for the result in a synchronous request
Video processing or GPU workloads Strong candidate Image, model, and, if necessary, GPU node startup are acceptable
Asynchronous integrations Conditionally safe A durable queue, retries, and idempotency exist
Internal administrative jobs Depends on SLA Delay before the first start is acceptable
Webhook receiver Unsafe without a buffer The incoming request must be accepted quickly and reliably
Public synchronous API Usually unsuitable A Kubernetes Service does not hold the request until a Pod is ready
Authentication or checkout Unsuitable without a warm reserve Cold starts move into the critical user path
Synchronous inference with strict p99 Unsuitable without a warm reserve Model and node loading create an unpredictable latency tail

A representative safe scenario is an asynchronous GPU video worker. The user request completes after the job is accepted durably, not after processing finishes. The queue exists independently of Pods. In this case, three separate latencies must be verified: HPA response, image and model loading, and GPU node provisioning.

The opposite case is an authentication endpoint or public synchronous API. Here, the first request cannot safely serve as the wake-up signal unless a separate always-running layer in front of the application can accept and retain the work.

How to validate the decision before enabling it in production

Start not with a broad HPA change, but with one component and an explicit hypothesis: which billable capacity is expected to be released, and what latency budget is acceptable for the first job.

1. Define suitability criteria

Before the experiment, establish:

  • the maximum acceptable age of a job;

  • the acceptable delay before the first start after idle time;

  • behavior when the metric is unavailable;

  • required reserve capacity;

  • whether reprocessing is acceptable;

  • the resource that can actually be removed: a Pod, node, or GPU.

2. Measure cold wake-up by stage

Run series of starts after different idle periods and collect p50, p95, and p99 for:

  • a job appearing in the source;

  • metric update;

  • the HPA or KEDA decision;

  • Pod creation;

  • assignment to a node;

  • container start;

  • successful readiness probe completion;

  • processing of the first job beginning.

Repeat the experiment separately with no free suitable node available. Otherwise, the measurement captures only a warm cluster and hides T_node_provision.

3. Run failure scenarios

At a minimum, test:

  • the metric source is unavailable;

  • the adapter returns an error or empty result;

  • the queue grows while the replica count remains zero;

  • the container receives SIGTERM during processing;

  • termination exceeds the grace period;

  • a message is delivered again;

  • the control plane is upgraded or rolled back while HPAs with minReplicas: 0 exist.

For native HPA, maintain a runbook for manually restoring capacity. For KEDA, verify the actual fallback behavior for the selected trigger type, installed operator version, and CRD—not only the current documentation.

4. Observe more than HPA

Useful signals for alerting include:

  • ScaledToZero and time spent at zero;

  • ScalingActive=False;

  • External Metrics API, adapter, or KEDA errors;

  • age of the oldest job;

  • failed readiness probes;

  • fallback activation;

  • forced Pod terminations;

  • the share of messages processed more than once;

  • time from demand appearing to the first Ready replica.

5. Verify actual savings

After the experiment, compare two periods: one with a continuously running replica and one with a zero minimum. Check not only Pod runtime hours, but also:

  • whether nodes were removed;

  • whether billable GPU or VM hours declined;

  • how many additional cold starts occurred;

  • whether retries and SLO violations increased;

  • how much operational time the metrics path required.

If the Pod disappears but the node remains billable, scale-to-zero has improved resource placement, not necessarily the budget. If one warm replica fits within the cost envelope and removes the long latency tail, minReplicas: 1 may be the more economical choice once reliability is included.

The practical rule is simple: use HPA scale-to-zero where demand is stored outside Pods, cold starts fit within a measured SLA, and released requests genuinely make it possible to reduce billable capacity. In all other cases, zero replicas may not be a saving at all, but rather a transfer of infrastructure risk into user latency and operations.

— we'll discuss the project, estimate timelines and suggest a format.
Need help with your project?

We'll figure it out together—
and show you how to solve
the problem quickly and effectively

Kubernetes HPA Scale-to-Zero: When Idle Savings Break Latency and Reliability