Debugging production, running incidents and writing them up.
5 items at advanced level · all topics
After editing the kube-apiserver static Pod manifest, kubectl fails with a connection refused error. Which approach diagnoses this?
With the API server down, kubectl is useless. Use crictl on the control plane node to find the container and read its logs, and check the kubelet journal for manifest errors.
A Pod writes a large volume of logs. kubectl logs returns only recent output, and the earlier lines you need are missing. Why, and what does that imply?
The kubelet rotates container logs and kubectl logs only reads the latest file. Keeping history means shipping logs off the node to a cluster-level logging system.
Pods on one node show status Evicted and the node reports the DiskPressure condition. Which explanation is correct?
The kubelet evicts Pods when a node-level resource crosses an eviction threshold. DiskPressure comes from node or image filesystem thresholds, whose defaults are 10% and 15% available.
A Pod has been Terminating for fifteen minutes after a delete. What is the correct sequence of things to consider?
Check whether the process is ignoring SIGTERM within its grace period, whether the node is unreachable, and whether a finalizer is holding the object. Force deletion is a last resort.
A Pod that was Running is suddenly gone, and events show the scheduler removed it to make room for a higher-priority Pod. Which mechanism is this, and what protects against it?
This is preemption, driven by PriorityClass. The scheduler removes lower priority Pods so a pending higher priority Pod can be scheduled, and preemptionPolicy: Never opts a Pod out of preempting others.