Making GPU Failure Invisible: A Zero-Cost Fallback for LLM Inference on KubernetesHow three things Kubernetes already does for free replaced a failover system I didn't want to writeAug 3, 2026·4 min read