The idea: a pod is a record that components keep acting on

A pod is one or more containers that share a network namespace (one IP) and run together on one node. But in Kubernetes a pod is first of all an object in the API server, stored in etcd: a spec (what should run) and a status (what is running). No single program "runs" the pod. Several components each watch the API server, and each does one step when the object changes. Then it writes its result back to the API, where the next component sees it.

The API server in the middle, connected only to kubectl (1 apply Deployment web), the ReplicaSet controller (2 create Pod), kube-scheduler (3 Binding nodeName node-1), the kubelet on node-1 (4 status Running, Ready) and the EndpointSlice controller (5 add the pod's IP)
Every component only watches and writes objects in the API server; the pod starts as each one, in turn, does its single step.

The canvas shows the control plane at the top: kubectl, the API server, the scheduler, the controller-manager and the EndpointSlices. Below it are two nodes, each with 4 CPUs already partly used by other pods, a kubelet and the container runtime containerd. On the right is the pod object, as kubectl get pod -o yaml would print it. Under them are the pod phase and the app container's state as two small state machines, then the live kubectl get pods table and the latest events. Every pod has one colour, used for its box on the node, its table row and its events. The clock t counts seconds.

Who does what

ComponentWatchesDoes
kubectl—sends your YAML to the API server (apply, delete, scale)
API server + etcd—validates and stores every object; serves the watches
Deployment / ReplicaSet / Job controllerits objects and their podscreates and deletes pods so that the count matches the spec
kube-schedulerpods with no nodeNamefilters and scores the nodes, writes a Binding
kubelet (one per node)pods bound to its nodedrives the runtime over CRI, runs the probes, reports status
containerd + CNI plugin— (called by the kubelet)sandbox, network namespace and IP, image pulls, containers
EndpointSlice controllerpods and Serviceslists each pod IP with ready / terminating
node lifecycle + taint-eviction controllersnode Leasesmark silent nodes Unknown, taint them, evict their pods

Follow one healthy pod in Demo 1. kubectl creates a Deployment, and the ReplicaSet controller creates the pod. The scheduler binds it to a node. The kubelet makes the sandbox, runs the init container, pulls the image and starts the app. The probes pass, the pod becomes Ready, and the EndpointSlice lists it.

Phase, container states and conditions

A pod's status has three layers. The phase is a coarse summary:

PhaseMeaning
Pendingaccepted, but not every container has started: not scheduled yet, init containers running, image pulling
Runningbound to a node, containers created, at least one is running, starting or restarting
Succeededall containers exited 0 and will not be restarted
Failedall containers have ended and at least one failed
Unknownthe pod's state could not be obtained, usually because its node cannot be reached

Each container is Waiting (with a reason: ContainerCreating, PodInitializing, ErrImagePull, ImagePullBackOff, CrashLoopBackOff), Running, or Terminated (with a reason and an exit code: Completed 0, Error, OOMKilled 137). Its lastState keeps the previous termination, so you can still see why it crashed after a restart.

The conditions are the milestones, each True or False with a time: PodScheduled → PodReadyToStartContainers (sandbox and network ready) → Initialized (all init containers done) → ContainersReady → Ready (also needs any readiness gates). Only Ready decides whether a Service sends traffic to the pod.

The STATUS column of kubectl get pods is none of these. kubectl picks the most telling string: Terminating if deletionTimestamp is set, Init:0/1 while init containers run, otherwise the waiting or terminated reason of a container, otherwise the phase. So a crash-looping pod shows CrashLoopBackOff while its phase is still Running (Demo 4).

Six steps of a healthy pod's start: scheduling, sandbox and IP, init container, image pull, app start, Ready. Phase is Pending for the first four and Running for the last two; kubectl STATUS goes Pending, Init:0/1, PodInitializing, Running 0/1, Running 1/1; conditions PodScheduled, PodReadyToStartContainers, Initialized and finally ContainersReady and Ready turn true along the way
Phase, STATUS and conditions are three different views of the same start; only the last condition, Ready, puts the pod behind its Service.

Scheduling: filter, score, bind

The scheduler takes one pod without a node at a time. Filter plugins drop the nodes that cannot run it: not enough free cpu or memory requests (the sum of the requests of the pods already there, not the actual use), taints the pod does not tolerate, node selectors and affinity, ports. Score plugins rank the nodes that are left; the page uses only "most free cpu" (LeastAllocated). The result is a Binding, which sets spec.nodeName. The scheduler never contacts the node.

If no node passes, the pod stays Pending with PodScheduled=False, reason Unschedulable, and an event such as 0/2 nodes are available: 2 Insufficient cpu. It waits in the unschedulable queue and is tried again when something in the cluster changes, for example when another pod ends (Demo 2). In a real cluster a cluster autoscaler would add a node at this point, or a higher-priority pod could preempt lower ones.

Starting: sandbox, init containers, image

The kubelet first asks the runtime for the pod sandbox (RunPodSandbox): a tiny pause container that owns the network namespace. The CNI plugin connects it and gives the pod its IP (see The Life of an HTTP Request in Kubernetes). Then the init containers run one after another, each to completion, typically to wait for a dependency or prepare files. Then the app's image is pulled, unless it is already on the node and imagePullPolicy is IfNotPresent.

If the pull fails (a typo in the tag, a private registry without credentials), the container waits with ErrImagePull and then ImagePullBackOff. The kubelet retries after 10 s, 20 s, 40 s and so on, up to 5 minutes. The phase stays Pending, because no container has ever run (Demo 3).

Probes: startup, readiness, liveness

The kubelet itself runs the probes, from the node: an HTTP GET, a TCP connect, a gRPC health check or a command in the container. Each probe has periodSeconds, timeoutSeconds and failureThreshold. The three kinds do different things when they fail:

ProbeQuestionWhen it fails failureThreshold times
startuphas the app finished booting?container killed and restarted; until it passes, the other two do not run
readinesscan it serve requests right now?Ready=False, removed from the Service's endpoints; not restarted
livenessis the process stuck?container killed (SIGTERM, then SIGKILL) and restarted

The page uses startup every 2 s × 15 (a slow app gets 30 s), readiness every 5 s × 3 and liveness every 10 s × 3. When a database goes down, the readiness probe fails and the pod leaves the EndpointSlice. When the database is back, it rejoins without a restart (Demo 6). A deadlocked app fails both probes: first it stops getting traffic, then the liveness probe has it killed (Demo 7). A liveness probe that checks a dependency is a classic mistake: when the database is down, it restarts every pod and fixes nothing.

Restarts and CrashLoopBackOff

restartPolicyContainer exits 0Container exits ≠ 0Used by
Always (default)restartrestartDeployments, StatefulSets, DaemonSets
OnFailurepod SucceededrestartJobs
Neverpod Succeededpod FailedJobs, one-off pods

A restart happens inside the same pod: same name, same node, same IP. Only restartCount goes up. The kubelet waits before each restart, and the wait doubles: 10 s, 20 s, 40 s, 80 s … up to 300 s. It goes back to 10 s after the container has run fine for 10 minutes. During the wait the container is Waiting with reason CrashLoopBackOff (Demo 4). The exit code tells you why it stopped: 1 is the program's own error, 137 is 128 + 9 (SIGKILL), 143 is 128 + 15 (SIGTERM). OOMKilled means the kernel killed the process because it went over its limits.memory, a cgroup limit (Demo 5). The kubelet is only the messenger.

A controller makes a new pod only when a pod is deleted or evicted. A crash-looping pod is not replaced: the ReplicaSet sees one pod, as the spec wants.

Termination: preStop, SIGTERM, grace period, SIGKILL

kubectl delete pod does not remove the object. It sets metadata.deletionTimestamp with a grace period (terminationGracePeriodSeconds, default 30 s). From that moment two things run in parallel:

  • The EndpointSlice controller marks the endpoint terminating and not ready, and every kube-proxy and load balancer removes it. This takes a moment to reach every node.
  • The kubelet runs the preStop hook, then sends SIGTERM to PID 1 of each container. If a container is still running when the grace period (counted from the delete, preStop included) ends, it gets SIGKILL.

Because of this race, an app that exits at once on SIGTERM can still get new requests and refuse them. A short preStop: sleep gives the endpoint removal time to spread (Demo 8). An app that ignores SIGTERM, often because it runs under a shell that does not forward signals, is always killed hard after the full grace period (Demo 9). When every container has stopped, the kubelet removes the sandbox and deletes the object with grace 0. Since Kubernetes 1.27 the pod first gets a final phase: Succeeded if every container exited 0, otherwise Failed.

The ReplicaSet does not wait for any of this. It counts only pods that are not terminating, so the replacement is created and scheduled while the old pod is still shutting down. For a moment both exist.

When the node disappears

Each kubelet renews a Lease object every 10 s. If a node goes silent (power, network, a frozen kubelet), the control plane waits node-monitor-grace-period (50 s; it was 40 s before Kubernetes 1.32), then sets the node's condition Ready=Unknown. It also adds the taints node.kubernetes.io/unreachable with effect NoSchedule and NoExecute. The pods on that node get Ready=False and leave their Services. Their phase still says Running, because that is the kubelet's last report.

Every pod has a default toleration for this taint with tolerationSeconds: 300. When those 5 minutes run out, the taint-eviction controller deletes the pod and its ReplicaSet makes a replacement on a healthy node. The old object stays Terminating: only its own kubelet can confirm the containers stopped. When the node comes back, the kubelet sees the deletion, kills the containers and the object disappears (Demo 10). Until then, a partitioned node may still be running the old containers. That is why StatefulSets never start the replacement while the old pod object exists.

Pods that finish: Jobs

A Job's pods use restartPolicy: Never or OnFailure. With Never, a container that exits 0 makes the pod Succeeded (STATUS Completed), and one that exits 1 makes it Failed (STATUS Error). Both phases are final, and the pod no longer counts against the node's cpu (Demo 11). The Job controller then decides: count a success, or make a new pod until backoffLimit failures (the page uses backoffLimit: 0). Finished pods stay in the API so their logs can be read, until the Job is deleted or ttlSecondsAfterFinished cleans them up.

See also The Life of an HTTP Request in Kubernetes (what happens to traffic once the pod is Ready) and How a Program Is Loaded and Run (what starting a process means inside the container).

What the page leaves out

QoS classes and eviction under node memory or disk pressure, priority and preemption, sidecar containers (init containers with restartPolicy: Always), ephemeral debug containers, readiness gates, PodDisruptionBudgets and kubectl drain, in-place resize of cpu and memory, static pods, postStart hooks, real timings (the scheduler and the kubelet react in milliseconds, not in whole seconds), and the probes that keep running in the background while the pod is healthy. The page also simplifies the Deployment: each kind of pod has its own Deployment and ReplicaSet, and there are no rolling updates.