Kubernetes Probes: When to Wait, Route Traffic, or Restart
Startup, readiness, and liveness probes answer different operational questions. Use a decision matrix, explicit endpoint contracts, and a proposed staging failure-test plan to choose when an application should wait, leave traffic rotation, or restart.
A hypothetical API process is running, but initialization has not finished. Later, its database goes offline. On a different day, a local deadlock stops requests from completing. Each condition might produce a failed health check. They should not automatically produce the same response.
Decide what would help. Allow more startup time? Remove this instance from traffic? Restart its container? Kubernetes gives those choices different probe semantics, as the official probe documentation explains.
Good probe design starts with a recovery argument. This guide covers a conventional HTTP application in a Deployment behind a normal Service. Direct Pod access, workers, specialized Service configurations, and custom routing need separate analysis. The configuration is illustrative and unexecuted.
Kubernetes readiness vs liveness vs startup: three different decisions
The probe determines what Kubernetes does with the endpoint signal. Review Kubernetes probe behavior before choosing the check.
| Probe | Question | Failure consequence at the configured threshold | Suitable evidence |
|---|---|---|---|
| Startup | Has initialization finished? | Container termination, then restart policy applies | Required local initialization completed |
| Readiness | Should this instance receive this traffic? | Container becomes unready, affecting Pod readiness and matching Service traffic | Ability to serve the scoped requests |
| Liveness | Is local recovery by restart appropriate? | Container termination, then restart policy applies | A local failure that restart can resolve |
Startup protects initialization
A startup probe holds back liveness and readiness checks until it succeeds. That gives initialization its own allowance without weakening ongoing checks. There is still a limit: repeated startup failure terminates the container. The startup configuration guidance describes this bounded wait.
Readiness controls traffic eligibility
Failed readiness does not itself restart anything. It changes readiness and eligibility for matching Service traffic. Pod readiness also requires the other containers and any configured readiness gates to be ready, as the Pod lifecycle documentation explains. Background work does not stop just because the instance becomes unready, and a failed check does not prove every route is unusable. See the probe documentation for the Service consequence; shutdown needs a separate contract.
Liveness asks whether restarting can help
At its failure threshold, liveness triggers termination of the specific container, not replacement of the entire Pod. Restart policy governs what follows, including the backoff described in the Pod lifecycle documentation. A running process can be stuck. A process waiting on an unavailable external service may have nothing locally wrong with it.
The Wait, Route, Restart framework for probe decisions
Wait, Route, Restart is an original Optijara editorial decision aid, not a Kubernetes feature, industry standard, or verified client method. Use it to review the recovery argument before implementing a health check.
Match the failure to an action
Every scenario below is hypothetical. Test the acceptance conditions against the actual workload.
| Failure condition | Useful signal | Probe action | Expected recovery | Acceptance evidence |
|---|---|---|---|---|
| Initialization incomplete | Local initialization state | Wait through bounded startup allowance | Initialization completes | Requests begin only after readiness succeeds |
| Essential request dependency unavailable | Scoped dependency status | Consider readiness failure, not automatic liveness failure | Dependency recovers | Required routes recover without unexplained restarts |
| Local deadlock | Progress signal tied to the request-serving path | Consider liveness failure | Restart restores local progress | Replacement process serves requests |
| Optional telemetry unavailable | Feature-specific failure | Keep useful routes eligible | Telemetry recovers separately | Core routes remain usable |
Shared dependencies complicate the routing decision. If every replica checks the same unavailable backend, withdrawing one instance can become withdrawing all of them. Colin Breck's practitioner analysis examines this failure-domain problem. Restarting replicas cannot, by itself, repair their shared dependency.
Record evidence and recovery ownership
Record what remains useful, what restart would change, and the recovery owner beside the configuration. This JSON is a planning record, not executable policy.
{
"framework": "wait-route-restart",
"wait": "bounded-initialization-allowance",
"route": "scoped-request-usefulness",
"restart": "tested-local-recovery-mechanism",
"acceptance": "observed-state-and-real-requests"
}Leaving out liveness can be reasonable when there is no justified restart signal. Kubernetes documents that option. The objection deserves equal attention: readiness withdrawal alone can leave a stuck instance unavailable indefinitely. Breck's analysis shows why that design still needs someone or something responsible for restoring progress.
Readiness dependency checks without withdrawing useful capacity
Separate essential dependencies from optional features
The official guidance permits readiness checks for strict backend dependencies. Henning Jacobs warns about shared dependency checks, while Breck examines mixed request dependencies. Ask what withdrawal accomplishes.
Take a hypothetical API that needs its database for every authorized request. Marking the instance unready might accurately describe its capacity. For another hypothetical API, read routes work despite a telemetry outage. Withdrawal would discard useful capacity. Degraded operation must still respect authorization and other safety requirements.
Where routes depend on different services, consider route-level failure handling or separate workloads with distinct readiness contracts. The failure boundary must justify the added deployment and maintenance overhead.
Bound checks and preserve degraded behavior
Set time budgets for dependency checks and specify what happens on timeout. Cached health state needs a maximum age and a rule for unknown or stale results. Yesterday's successful check is not evidence that the service can accept requests now.
Readiness can oscillate. Kubernetes offers consecutive success and failure thresholds, with semantics set out in the configuration guide. Choose those thresholds from observed behavior. Any extra smoothing in application code needs testing too.
Keep dependency failures out of liveness unless testing explains how local restart repairs them. The practitioner posts include historical controller examples; validate the installed ingress path rather than assuming those examples describe its current behavior.
An illustrative HTTP probe configuration and endpoint contract
Give startup, readiness, and liveness separate meanings
This fragment belongs inside a container specification. It is not a complete runnable Deployment. It assumes an application listening on port 8080 with the endpoints shown. The timing values are unbenchmarked examples, not tuning recommendations.
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /startupz
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 24
successThreshold: 1
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
successThreshold: 2
livenessProbe:
httpGet:
path: /livez
port: http
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
successThreshold: 1Under this proposed contract, /startupz reports completed initialization. /readyz reports whether the instance should receive its scoped traffic. /livez detects a local condition that restart can address. The separate paths make this example easier to review; Kubernetes does not require separate URLs.
HTTP probes check response status, not business-operation success. The official configuration guide defines successful status codes as 200 through 399. Redirect handling has exceptions worth inspecting: kubelet follows same-host redirects, but a different-host redirect or 11 or more redirects is treated as success with a ProbeWarning event. Prefer an explicit health response over a login redirect. Keep payloads small and exclude secrets or sensitive dependency details.
Choose timing from workload evidence
periodSeconds sets probe frequency. timeoutSeconds bounds an individual probe. failureThreshold counts consecutive failures, while successThreshold counts consecutive successes after failure. The latter must be 1 for startup and liveness. Readiness may run more frequently while the container is unready. These are documented field semantics, not a tuning recipe.
Measure initialization under the startup conditions you intend to support, including relevant cold starts. Review termination grace as well. Kubelet honors applicable grace during probe-triggered termination; probe-level grace is available for startup and liveness, not readiness. See the probe documentation.
Multiplying thresholds does not produce an exact recovery deadline. Scheduling, timeouts, shutdown, restart backoff, and initialization also affect elapsed time. A runnable guide would need a complete fixture with locked versions and retained execution output. This fragment has not been executed on a cluster.
A staging playbook: test failure, inspect evidence, then roll out
Capture the baseline and test one failure at a time
These are proposed tests, not experimental results. Before the first test:
- Inventory endpoint behavior, framework health defaults, and routing assumptions.
- Save the current probe configuration and baseline request, readiness, and restart evidence.
- Assign an expected action and recovery owner to each failure.
- Exercise one failure at a time in an isolated staging workload.
- Compare observations with the contract before approving a bounded rollout.
- Retain the reviewed application configuration and deployment rollback path.
Use a disposable application for deliberate stalls. Do not inject deadlocks into shared production services.
| Proposed test | Expected readiness and restarts | Useful request outcome | Recovery trigger | Evidence to retain |
|---|---|---|---|---|
| Delayed startup within allowance | Unready until startup and readiness succeed; no probe-triggered restart | No premature Service traffic | Initialization completes | Startup logs, conditions, requests |
| Essential dependency outage | Contract-defined unready; no automatic liveness restart | Required operations fail explicitly | Dependency restored | Dependency state, events, restart counts |
| Optional dependency outage | Useful routes stay ready; no probe-triggered restart | Core operations remain available | Optional feature restored | Route-specific requests and errors |
| Readiness recovery | Probe succeeds after configured successes; Pod readiness still requires other conditions; no restart required | Intended Service path works again | Readiness contract restored | EndpointSlice state and request timeline |
| Controlled local stall | Readiness follows its contract; liveness termination only if justified | Restarted process resumes useful work | Local process restart | Probe events, previous logs, requests |
Check both Kubernetes state and real requests
Read Pod conditions and container restart counts. Inspect events and previous-container logs where available, then correlate matching Service EndpointSlice readiness with requests through the intended Service or ingress. The probe and lifecycle documentation explain the state transitions. Accepted configuration proves little about recovery on its own.
EndpointSlice changes reach client watches and caches at different times, as the official EndpointSlice documentation explains. It also documents an exception: publishNotReadyAddresses makes endpoint ready true regardless of Pod readiness. This guide assumes that option is disabled. Test the installed routing path and long-lived requests; do not assume every controller immediately stops new routing or that readiness failure closes established connections.
For inference services, compare probe evidence with AI inference observability. A healthy process does not establish response latency, request success, or output quality. The OpenTelemetry GenAI tracing playbook adds context for model and tool calls. It complements probe testing.
Measure recovery, not just green indicators.
| Measurement | Capture method | Acceptance question |
|---|---|---|
| Startup behavior | Timestamp initialization and readiness transitions | Does the allowance cover intended startup conditions? |
| Restart behavior | Compare counts, events, and previous logs | Did restart resolve the local failure? |
| Traffic eligibility | Correlate Pod conditions and EndpointSlice state | Did only the intended instances withdraw and recover? |
| User-visible behavior | Exercise representative routes through actual routing | Did useful requests behave as the contract specifies? |
| Recovery and check overhead | Record recovery timeline and health-handler resource use | Is recovery explained without excessive checking cost? |
Define recovery and rollback before production
Reject the change if unrelated replicas withdraw, restarts increase without recovery, useful routes fail, or recovery remains unexplained. Set acceptance criteria for the workload instead of borrowing universal latency or availability thresholds.
Rollback must restore the reviewed endpoint contract and probe configuration through the normal deployment process. Check readiness, restart behavior, and actual requests afterward. A successful rollout command does not prove recovery.
What teams get wrong, and what probes cannot guarantee
Avoid shared failure signals and unjustified restarts
Identical checks deserve scrutiny when their failure consequences differ. Other mistakes include probing a management listener that says little about the request-serving path, or setting aggressive thresholds without startup measurements. Jacobs discusses management-port blind spots and restart risks.
Insisting on different URLs is no better as an absolute rule. Kubernetes describes shared-endpoint designs with different thresholds. Review the signal and its consequence. Endpoint naming cannot settle whether a restart will help.
Keep disruption, shutdown, and durable work separate
A PodDisruptionBudget constrains supported voluntary eviction operations. It does not coordinate liveness-triggered container restarts or guarantee an availability floor. The official disruption documentation defines the scope. The historical community discussion helps explain the confusion; it is not a current specification.
Readiness withdrawal does not provide graceful termination, connection draining, or worker shutdown. The Pod lifecycle documentation treats termination separately. Those responsibilities need explicit designs and tests.
Restarting a container does not establish authoritative job state or deduplicate external writes either. For agent workloads, review durable workflow state and recovery alongside container health.
Probes observe the signals you chose. They cannot establish complete application correctness, security, or inference quality. Health handlers also consume resources, and cached state can become stale. The runbook should explain these limits and how to recognize recovery.
Key Takeaways
- 1Choose the recovery action before designing the health signal: wait, route, or restart.
- 2Startup probes gate liveness and readiness until success, but repeated startup failure can still terminate the container.
- 3Readiness controls traffic eligibility, not process recovery or background-worker lifecycle.
- 4Only use restart-triggering liveness checks when evidence supports local recovery by restart.
- 5Shared dependency checks require a review of useful routes and replica-wide failure behavior.
- 6Validate Pod state, EndpointSlice readiness, and real request outcomes together in staging.
- 7Treat disruption budgets, graceful shutdown, durable state, and application correctness as separate responsibilities.
Conclusion
Before changing production probes, choose one representative workload and work through the staging matrix. Be able to explain why an instance should leave traffic rotation and what a restart would repair. Keep the endpoint contract with its recovery owner and rollback evidence. For Kubernetes-hosted AI services, Optijara can help review that reasoning and the failure-test evidence alongside the wider application architecture.
Frequently Asked Questions
What is the difference between Kubernetes readiness and liveness probes?
Readiness controls eligibility for matching Service traffic; failure does not itself restart the container. Liveness failure at its threshold triggers container termination, then restart policy applies. Neither check proves complete application correctness. See https://kubernetes.io/docs/concepts/workloads/pods/probes/
When should I use a startup probe instead of a longer liveness delay?
Use startup for a separate, bounded initialization allowance. It blocks readiness and liveness checks until success; repeated startup failure can still terminate the container. Choose timing from observed initialization, not a universal delay. See https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
Should a readiness probe check the database?
Only when database availability determines whether the instance can safely serve its scoped traffic. Check shared-outage behavior and routes that remain useful. Database loss should not trigger liveness failure unless local restart has a tested recovery purpose. See https://kubernetes.io/docs/concepts/workloads/pods/probes/ and https://blog.colinbreck.com/kubernetes-liveness-and-readiness-probes-looking-for-more-feet/
Why can a liveness probe cause a restart loop?
Without a startup probe protecting initialization, liveness can fail before startup completes. Overload or an external outage can also fail a poorly scoped check without a restart repairing the cause. Inspect events, restart counts, previous logs, timing, and real requests. CrashLoopBackOff indicates restart backoff, not proof of a probe-caused failure. See https://kubernetes.io/docs/concepts/workloads/pods/probes/ and https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/
Do readiness failures or liveness restarts coordinate shutdown and disruption protection?
No. Readiness does not stop background jobs or guarantee that existing connections close. Shutdown and routing propagation need explicit handling. PodDisruptionBudgets constrain supported voluntary evictions, not liveness-triggered container restarts. See https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/ and https://kubernetes.io/docs/concepts/workloads/pods/disruptions/
Sources
- https://kubernetes.io/docs/concepts/workloads/pods/probes/
- https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
- https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/
- https://kubernetes.io/docs/concepts/workloads/pods/disruptions/
- https://srcco.de/posts/kubernetes-liveness-probes-are-dangerous.html
- https://blog.colinbreck.com/kubernetes-liveness-and-readiness-probes-looking-for-more-feet/
- https://github.com/kubernetes/website/issues/16607
- https://kubernetes.io/docs/concepts/services-networking/endpoint-slices/
Written by
Hamza DiazHamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.
