前面的文章已經介紹過 Deployment、Service、Ingress、NetworkPolicy 等 Kubernetes 資源,也在實作過程中使用過 kubectl get、kubectl describe、kubectl logs 等指令查看資源狀態。
但實際使用 Kubernetes 時,遇到的問題通常不會直接告訴我們是哪個地方設定錯誤。
例如:
Pod 明明顯示 Running,為什麼 Service 還是連不上?
這時就需要從不同 Kubernetes 資源之間的關係,一層一層找出問題的位置。
這篇會故意建立一個設定錯誤的 Service,實際走一次簡單的故障排查流程。
首先建立一個 Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: debug-demo
spec:
replicas: 2
selector:
matchLabels:
app: debug-demo
template:
metadata:
labels:
app: debug-demo
spec:
containers:
- name: nginx
image: nginx
ports:
- containerPort: 80
接著建立 Service:
apiVersion: v1
kind: Service
metadata:
name: debug-service
spec:
selector:
app: debug-test
ports:
- port: 80
targetPort: 80
套用到Kubernetes。
這份設定其實故意留下了一個錯誤。
先不要急著找錯在哪裡,接下來按照實際排查問題的方式一步一步檢查。
服務無法使用時,可以先確認 Pod 是否正常:
kubectl get pods
可以看到類似:

兩個 Pod 都是:
Running
代表目前 Container 已經正常啟動。
如果這裡出現:
Pending
CrashLoopBackOff
ImagePullBackOff
就應該先處理 Pod 本身的問題。
但這次 Pod 看起來沒有異常,因此問題可能發生在其他地方。
接著查看 Service:
kubectl get svc
可以看到:
Service 也成功建立了。
但「Service 存在」不代表它後面一定有可以接收流量的 Pod。
Service 需要透過 Selector 找到符合 Label 的 Pod,因此接下來要確認 Service 實際找到哪些後端。
可以查看:
kubectl get endpoints debug-service
這時可能會看到:

ENDPOINTS 顯示:
<none>
這就是一個重要線索。
代表 Service 雖然存在,但目前沒有找到可以接收流量的後端 Pod。
現在問題範圍就縮小到:
Service
↓
找不到
↓
Pod
接下來就要檢查 Service 的 Selector 與 Pod 的 Label。
在較新的 Kubernetes 環境中,Service 後端主要由 EndpointSlice 管理,也可以使用
kubectl get endpointslices查看;這裡使用kubectl get endpoints方便直接觀察 Service 是否找到後端。
查看 Pod 的 Label:
kubectl get pods --show-labels

可以看到 Pod 使用:
app=debug-demo
接著查看 Service:
kubectl describe service debug-service

其中可以看到:
Selector: app=debug-test
這時就找到問題了:
Pod Label
app=debug-demo
≠
Service Selector
app=debug-test
Service 使用 Selector 尋找 Pod,但兩邊的 Label 不一致,因此 Service 找不到任何 Pod。
整個問題可以整理成:
Service
│
│ selector: app=debug-test
▼
找不到符合條件的 Pod
Pod
└── label: app=debug-demo
將 Service 修改成:
apiVersion: v1
kind: Service
metadata:
name: debug-service
spec:
selector:
app: debug-demo
ports:
- port: 80
targetPort: 80
重新套用後再次查看:
kubectl get endpoints debug-service
這次應該可以看到
代表 Service 已經成功找到兩個 Pod。
接著可以在 Cluster 內建立一個暫時的 Pod 測試 Service:
kubectl run test-client \
--image=busybox:1.36 \
--restart=Never \
--rm -it \
-- wget -qO- http://debug-service
如果設定正確,就可以取得 NGINX 回傳的網頁內容。
這次的問題其實不複雜,但重點並不是記住某一個指令,而是知道發生問題時應該從哪裡開始找。
例如服務無法連線時,可以按照:
服務無法連線
│
▼
Pod 是否正常?
│
├── No
│ ↓
│ 檢查 Pod 狀態
│ describe / logs
│
└── Yes
↓
Service 是否存在?
│
▼
Service 是否有後端?
│
├── No
│ ↓
│ 檢查 Selector / Label
│
└── Yes
↓
Port / targetPort 是否正確?
│
▼
NetworkPolicy 是否阻擋?
│
▼
Ingress 設定是否正確?
不同問題會需要不同的排查方式。
如果 Pod 本身無法正常啟動,就應該先查看 Pod 的狀態、Events 或 Logs。
如果 Pod 正常,但 Service 找不到後端,就應該檢查 Label、Selector 與 Endpoint。
如果 Service 在 Cluster 內可以正常使用,但從外部無法存取,問題則可能出現在 Ingress、Service 類型或其他網路設定。
Kubernetes 中的應用程式通常不是由單一資源組成,而是由多個資源互相配合:
Ingress
↓
Service
↓
Pod
↓
Container
因此當服務發生問題時,不一定代表 Pod 本身出錯。
這篇的案例中:
Pod → Running
Service → 正常建立
Endpoint → <none>
Service Selector → 與 Pod Label 不一致
透過逐層檢查,最後找到 Service Selector 設定錯誤。
前面的文章比較著重在「如何建立與使用 Kubernetes 資源」,到了實際維運時,更重要的是理解這些資源之間的關係。當服務發生問題時,依照 Pod、Service、Endpoint、Ingress 等層級逐步縮小範圍,會比直接猜測是哪個 YAML 寫錯更容易找到問題。