iT邦幫忙

2026 iThome 鐵人賽

DAY 15
0

昨天學了 Liveness Probe,讓 K8s 在服務掛掉時自動重啟 Pod。今天來看另一種情況:服務明明還沒準備好,K8s 卻已經把流量送過來了,這是 Readiness Probe 要解決的問題 !


Liveness 和 Readiness 的差異

這兩個 Probe 長得很像,但做的事情完全不同:

Liveness Probe Readiness Probe
檢查失敗時 重啟 Pod 從 Service 的 Endpoints 移除
目的 服務死掉了,重啟它 服務還沒準備好,先不要送流量
典型情境 死鎖、記憶體洩漏 啟動中、資料庫連線還沒建好

關鍵差異:Readiness 失敗不會重啟 Pod,只會把它從 Service 的流量池裡移除。等 Pod 準備好了,Readiness 檢查通過,才會重新加回流量池。


為什麼需要 Readiness Probe

有兩個典型情境:

情境一:服務啟動需要時間

FastAPI 啟動的時候需要:

  • 載入應用程式設定
  • 建立資料庫連線池
  • 執行資料庫 migration

這個過程可能要幾秒甚至幾十秒。如果 Readiness Probe 還沒設好,Pod 一啟動 K8s 就把它加進 Service 的 Endpoints,這段時間進來的請求會收到錯誤。

情境二:Rolling Update 的時候

Day 07 學的 Rolling Update,新版本 Pod 啟動的時候,K8s 要知道它什麼時候準備好才能把流量切過去、把舊版本 Pod 關掉。

沒有 Readiness Probe,K8s 只能靠 initialDelaySeconds 猜「等這麼久應該準備好了吧」,不夠精確。有了 Readiness Probe,K8s 能精確知道「這個 Pod 現在可以服務流量了」。


實際操作

今天目標:幫 todo-api 設定 Readiness Probe,透過故意設定錯誤路徑觀察 Pod 停在 0/1、Endpoints 變空、請求回傳 503,確認 Readiness 失敗不重啟只擋流量的行為。

Step1: 加上 Readiness Probe

更新 todo-api-deployment.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: todo-api
spec:
  replicas: 2
  selector:
    matchLabels:
      app: todo-api
  template:
    metadata:
      labels:
        app: todo-api
        tier: backend
    spec:
      containers:
        - name: api
          image: yourname/todo-app-api:v1.0.0
          ports:
            - containerPort: 8000
          envFrom:
            - configMapRef:
                name: todo-api-config
            - secretRef:
                name: todo-api-secret
          livenessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 15
            periodSeconds: 10
            timeoutSeconds: 5
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /ready
              port: 8000
            initialDelaySeconds: 5
            periodSeconds: 5
            timeoutSeconds: 3
            failureThreshold: 3
            successThreshold: 1
kubectl apply -f todo-api-deployment.yaml

Step2: 觀察 Pod 啟動的狀態變化

kubectl get pods -w

Pod 剛啟動的時候 READY 欄位是 0/1:

NAME                       READY   STATUS    RESTARTS   AGE
todo-api-xxx               0/1     Running   0          5s    ← 還沒 Ready
todo-api-xxx               1/1     Running   0          12s   ← Ready 了

0/1 的時候這個 Pod 不在 Service 的 Endpoints 裡,流量不會送過來。

Step3: 觀察 Readiness 失敗時的行為

故意把todo-api-deployment.yaml 中 path 改成不存在的路徑,模擬服務還沒準備好的情境:

readinessProbe:
  httpGet:
    path: /nonexistent    # 故意打一個不存在的路徑
    port: 8000
  initialDelaySeconds: 5
  periodSeconds: 5
  failureThreshold: 3

先把所有 Pod 砍掉,確保重建的 Pod 都套用新設定:

kubectl apply -f todo-api-deployment.yaml
kubectl scale deployment todo-api --replicas=0
kubectl scale deployment todo-api --replicas=2

⚠️ 直接 apply 不夠,因為 Rolling Update 機制會保留舊版本 Pod 繼續服務流量,導致實驗結果不乾淨。scale to 0 再 scale back 可以確保所有 Pod 都是新設定。

同時開三個 terminal 觀察:

Terminal 1:port-forward Ingress Controller

kubectl port-forward service/ingress-nginx-controller 8888:80 -n ingress-nginx

⚠️ 注意:要 forward Ingress Controller 而不是直接 forward todo-api-service。因為 Ingress 有設定 rewrite-target,直接打 Service 會繞過 rewrite,FastAPI 收到的路徑是 /api/todos 而非 /todos,導致 404。

Terminal 2:觀察 Pod 狀態

kubectl get pods -w

你會看到 READY 一直停在 0/1,而且 RESTARTS 不會增加:

NAME                       READY   STATUS    RESTARTS   AGE
todo-api-xxx               0/1     Running   0          10s   ← 一直停在這裡
todo-api-xxx               0/1     Running   0          30s   ← 沒有重啟
todo-api-xxx               0/1     Running   0          50s   ← 還是沒有重啟

Terminal 3:觀察流量是否被擋住

while ($true) { try { (Invoke-WebRequest -Uri http://localhost:8888/api/todos -UseBasicParsing).StatusCode } catch { $_.Exception.Response.StatusCode.value__ }; Start-Sleep 2 }

這段時間打進來的請求全部回傳 503,因為沒有任何 Pod 在 Endpoints 裡:

https://ithelp.ithome.com.tw/upload/images/20260921/20183863vGVpfSNAkC.png

對比 Day 14 的 Liveness 失敗——RESTARTS 會一直增加。Readiness 失敗的特徵是:Pod 活著、沒有重啟,但就是不接流量。

Step4: 確認 Endpoints 的變化

kubectl describe service todo-api-service

Readiness 失敗的時候,Endpoints 欄位是空的:

https://ithelp.ithome.com.tw/upload/images/20260921/20183863gqG28omnTT.png

Step5: 恢復正常

確認效果後,todo-api-deployment.yaml 把 path 改回 /ready 並重新 apply:

readinessProbe:
  httpGet:
    path: /ready
    port: 8000
  initialDelaySeconds: 5
  periodSeconds: 5
  timeoutSeconds: 3
  failureThreshold: 3
  successThreshold: 1
kubectl apply -f todo-api-deployment.yaml

再觀察 Terminal 2,流量會自動恢復正常。

兩個 Probe 的參數建議

livenessProbe:
  initialDelaySeconds: 15    # 等應用程式完全啟動
  periodSeconds: 10          # 每 10 秒檢查一次,不需要太頻繁
  failureThreshold: 3        # 連續失敗 3 次才重啟,避免誤判

readinessProbe:
  initialDelaySeconds: 5     # 盡快開始檢查,確認準備狀態
  periodSeconds: 5           # 比 Liveness 頻繁,更即時感知狀態
  failureThreshold: 3        # 連續失敗 3 次才移出 Endpoints
  successThreshold: 1        # 成功 1 次就算準備好(預設值)

延伸筆記:Startup Probe

除了 Liveness 和 Readiness,K8s 還有第三種 Probe:Startup Probe。

它解決的問題是:有些應用程式(例如 Java 服務、大型資料庫)啟動非常慢,可能要 1-2 分鐘。如果 initialDelaySeconds 設太短,Liveness Probe 會在服務還沒啟動完就開始檢查,導致不斷重啟;設太長,服務真的掛掉了也要等很久才會重啟。

Startup Probe 的邏輯是:在 Startup Probe 成功之前,Liveness 和 Readiness Probe 都不會啟動。

startupProbe:
  httpGet:
    path: /health
    port: 8000
  failureThreshold: 30    # 最多等 30 * 10 = 300 秒(5 分鐘)
  periodSeconds: 10

Startup Probe 成功之後,才開始跑 Liveness 和 Readiness Probe。這讓啟動慢的服務有足夠的時間啟動,啟動完成後又能快速感知健康狀態。

Todo App 的 FastAPI 啟動很快,不需要用到 Startup Probe。但如果未來有遇到啟動慢的服務,可以使用Startup Probe。


小結

今天學了 Readiness Probe,總結一下 K8s 的三種 Probe 各自解決不同問題:

  • Startup Probe:服務啟動太慢 → 暫緩 Liveness 和 Readiness 的檢查,Startup 成功前兩者都不會啟動
  • Liveness Probe:服務死掉了 → 重啟 Pod
  • Readiness Probe:服務還沒準備好 → 暫時移出流量池,不重啟

三者的觸發順序是:Startup → Liveness / Readiness 同時運行。

Liveness 和 Readiness 最關鍵的差異在於失敗後的處理方式:一個重啟 Pod,一個只擋流量。實際部署通常兩個一起設定。

明天會學 Resource Request & Limit,設定每個 Pod 能用多少 CPU 和記憶體,避免單一 Pod 吃光 Node 的資源 !


上一篇
Day 14|Liveness Probe
系列文
從零學 K8s|30 天核心概念 × 實作,新手也能真正掌握 Kubernetes 共 15 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言