Observed Signal · May 17, 2026 · Incident & Recovery Guide · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
etcd NOSPACE recovery guide for on‑prem Kubernetes
A DEV.to post documents a production incident where an on‑prem Kubernetes control plane became unresponsive because etcd hit its storage limit and raised a NOSPACE alarm. The author describes diagnosing the issue by inspecting etcd logs and per‑node disk usage, then recovering the cluster without kubectl by SSHing into master nodes, using crictl to exec into the etcd container, and running etcdctl compact, defrag and alarm disarm on each member. The guide explains the difference between compaction and defragmentation, notes default etcd size limits (2GB) and that auto‑compaction is often not configured by kubeadm, and recommends adding --auto-compaction-retention=1h to static pod manifests to prevent recurrence.
Practical operational runbook for recovering on‑prem Kubernetes control planes after etcd storage exhaustion; relevant to infrastructure and SRE teams but not industry‑shifting.
Track Real-Time Infrastructure / Core IT (Kubernetes etcd incident) Signals & Market Shifts
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- An on‑prem Kubernetes control plane became degraded after etcd raised a NOSPACE alarm when the database exceeded its storage limit (default ~2GB).
- Recovery was performed without kubectl by SSHing into master nodes, using crictl to access the etcd container, then running etcdctl compact, defrag and alarm disarm on each etcd member.
- Compaction removes old revisions logically; defragmentation rewrites the DB file to actually reclaim disk space; both steps are required to reduce disk usage.
- Many kubeadm clusters do not enable automatic compaction by default; the author recommends adding --auto-compaction-retention=1h to /etc/kubernetes/manifests/etcd.yaml.
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Practical Kubernetes Production Troubleshooting Workflow
A practical how-to describing a repeatable sequence for troubleshooting Kubernetes workloads in production. The author prescribes a baseline flow (kubectl get pods -A → kubectl describe pod → kubectl logs --previous → kubectl top → kubectl get events) and a failure-classification approach that maps observed symptoms to targeted diagnosis and recovery actions. Five common scenarios are documented: ImagePullBackOff, CrashLoopBackOff, Pending pods, ingress 502/503 with healthy pods, and cluster DNS/CoreDNS failures. For each scenario the post lists typical root causes, concrete kubectl commands for diagnosis and recovery, and prevention tactics (CI image pinning, startup probes, capacity planning, smoke tests). The guidance emphasizes reading events and previous logs before restarting to preserve crash context.
Fix AKS Memory Eviction with Azure Disk CSI
This technical guide explains a cascade that causes stateful pods on Azure Kubernetes Service (AKS) to be evicted under node MemoryPressure, leaving Azure Disk CSI VolumeAttachment objects stuck in Terminating and blocking pod rescheduling. It details immediate remediation (force-removing the VolumeAttachment finalizer after verifying the disk is not attached), root-cause fixes (set explicit resource requests/limits and use Guaranteed QoS), kubelet eviction tuning via AKS KubeletConfig, and operational controls such as PodDisruptionBudgets, OPA/Gatekeeper admission policies, Checkov IaC scanning, and node-pool sizing. The article provides CLI and manifest examples to apply each fix and CI/CD prevention measures to avoid recurrence.
How Kubernetes Storage Works for Sysadmins
This technical guide explains how Kubernetes provides persistent storage for pods through a sequence of abstractions: Pod → PVC → CSI → external storage → PV. It describes the roles of PersistentVolumeClaims (PVCs), PersistentVolumes (PVs), StorageClasses, and CSI drivers (Controller and Node plugins), and explains access modes (ReadWriteOnce, ReadWriteMany, ReadOnlyMany), reclaim policies (Delete vs Retain), and provisioning modes (static vs dynamic). The article outlines how kubelet, the CSI Node plugin, and mount paths operate on nodes, and provides a step-by-step debugging checklist (kubectl commands and where to check CSI logs) for diagnosing storage failures.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
