到目前為止,我們已經做到:
Git Push
│
▼
CI/CD
│
▼
Docker Image
│
▼
Kubernetes
│
▼
Deployment
│
▼
Service
│
▼
Ingress
│
▼
User
而且還有:
HPA
→ 自動增加 Pod
Cluster Autoscaler
→ 自動增加 Node
看起來好像已經非常完整。
但是一個真正的 Production System,還會面臨很多問題:
API 變慢了,我怎麼知道?
某個 Pod 一直 Error,我怎麼查?
CPU 快爆了,誰會通知我?
Database 掛掉怎麼辦?
資料被刪掉怎麼復原?
Secret 洩漏怎麼處理?
某一次部署到底是誰做的?
第三方 API 變慢怎麼發現?
昨天晚上發生問題,但現在已經恢復,怎麼追?
因此 Production 不只是:
讓 Application 跑起來
還要做到:
可以觀察
可以監控
可以告警
可以復原
可以追蹤
可以保護
最基本的就是 Log。
例如 FastAPI 收到 Request:
GET /api/products
可以記錄:
2026-09-25 10:01:23
GET /api/products
status=200
duration=120ms
如果出錯:
2026-09-25 10:02:10
POST /api/orders
status=500
Database connection failed
Log 可以幫助我們回答:
發生了什麼事情?
如果只有一台 Server:
Server A
└── /var/log/app.log
出問題時:
tail -f /var/log/app.log
就能看。
但是 Kubernetes 裡可能有:
Pod 1
Pod 2
Pod 3
Pod 4
Pod 5
而且 Pod 隨時可能被:
重啟
刪除
重新建立
如果 Log 只存在 Pod 裡:
Pod X
↓
刪除
Log X
歷史紀錄就可能一起消失。
因此 Production 通常需要:
集中式 Logging
架構:
Pod 1 ─┐
Pod 2 ─┤
Pod 3 ─┼──→ Log System
Pod 4 ─┤
Pod 5 ─┘
常見組合例如:
Elasticsearch
+
Logstash / Fluent Bit
+
Kibana
或者:
Loki
+
Grafana
這樣工程師就不需要:
kubectl logs pod-1
kubectl logs pod-2
kubectl logs pod-3
逐台找。
而是直接在中央平台搜尋:
request_id = abc123
或:
status = 500
Production Log 最好不要只是:
something error
而應該包含有結構的資訊。
例如:
{
"timestamp": "2026-09-25T10:00:00",
"level": "ERROR",
"service": "backend",
"request_id": "abc123",
"path": "/api/orders",
"status": 500,
"duration_ms": 235,
"message": "database timeout"
}
這樣就比較容易搜尋:
所有 ERROR
所有 /api/orders
duration > 1000ms
特定 request_id
假設一個 Request 經過:
Ingress
│
▼
Backend
│
▼
Redis
│
▼
Database
如果每一層都有自己的 Log,很難知道哪些 Log 是同一次 Request。
因此可以產生:
request_id = abc123
然後一路帶下去:
Ingress
request_id=abc123
Backend
request_id=abc123
Database Error
request_id=abc123
這樣就可以把同一次 Request 串起來。
Log 告訴我們:
發生了什麼?
Metrics 則告訴我們:
整體系統現在狀況怎麼樣?
例如:
CPU Usage
Memory Usage
Requests Per Second
Response Time
Error Rate
Pod Count
Database Connections
Redis Memory
這些都是 Metrics。
Kubernetes 世界裡非常常見的 Metrics 系統是:
Prometheus
Prometheus 會定期收集:
Application
Kubernetes
Node
Database
的 Metrics。
架構:
FastAPI Pods ─┐
Kubernetes ───┤
Nodes ────────┼──→ Prometheus
Redis ────────┤
Database ─────┘
例如收集:
http_requests_total
http_request_duration_seconds
cpu_usage
memory_usage
Prometheus 主要負責:
收集與查詢 Metrics
Grafana 則負責:
視覺化
例如 Dashboard:
API Request / second
████████████████
Average Response Time
320ms
Error Rate
1.2%
CPU
72%
Memory
63%
所以常見組合:
Prometheus
│
▼
Grafana
可以理解成:
Prometheus
→ 收資料
Grafana
→ 看資料
Production Monitoring 可以先關注幾個非常重要的指標。
例如:
Traffic
Latency
Errors
Saturation
Traffic:
現在每秒多少 Request?
Latency:
Response 多久?
Errors:
多少 Request 失敗?
Saturation:
CPU / Memory / Connection
是不是快用完?
例如:
Request/sec = 500
P95 latency = 800ms
5xx rate = 3%
CPU = 92%
這些資訊通常比單純問:
Server 活著嗎?
更有價值。
例如平均 Response Time:
200ms
看起來很快。
但是可能:
90% Request
100ms
10% Request
2 秒
平均值可能掩蓋問題。
所以 Production 常會看:
P50
P95
P99
例如:
P50 = 100ms
P95 = 500ms
P99 = 2s
意思是:
99% Request
在 2 秒內完成
這通常比單純 Average 更能反映使用者體驗。
Monitoring 只是:
你打開 Grafana
就可以看到問題
但 Production 不可能要求工程師:
24 小時盯著 Dashboard
所以需要:
Alert
例如:
Error Rate > 5%
持續 5 分鐘
就通知:
Slack
Email
PagerDuty
Teams
或者:
CPU > 90%
持續 10 分鐘
發出警報。
如果設定:
CPU > 50%
→ Alert
可能每天一直通知。
久了工程師會變成:
又來了
不用管
這叫:
Alert Fatigue
所以真正有價值的 Alert 應該是:
需要有人採取行動
例如:
5xx Error Rate
持續超過 5%
Database Connection
快用完
Pod Crash
持續發生
Disk
剩不到 10%
服務完全無法使用
如果系統很簡單:
Browser
→ Backend
→ Database
Log 通常就夠用。
但如果變成:
Frontend
│
▼
API Gateway
│
▼
Order Service
│
├── Payment Service
│
├── Inventory Service
│
└── Notification Service
一個 Request 可能經過很多 Service。
現在使用者說:
結帳要 5 秒
問題到底在哪?
可能:
Order Service
50ms
Inventory
100ms
Payment
4.5 秒
Notification
200ms
這時就需要:
Distributed Tracing
一個 Request 可以被賦予:
trace_id
例如:
trace_id = XYZ123
然後看到:
Ingress
20ms
Backend
100ms
Payment API
3200ms
Database
50ms
馬上就能發現:
Payment API
是瓶頸。
常見技術包括:
OpenTelemetry
Jaeger
Tempo
因此 Production 常講:
Observability
常見三個主要部分:
Logs
Metrics
Traces
簡單記:
Logs
→ 發生什麼?
Metrics
→ 系統現在多健康?
Traces
→ 一個 Request 到底走了哪裡?
例如:
Metrics
發現 latency 突然變高
│
▼
Tracing
發現 Payment API 很慢
│
▼
Logs
找到 Payment timeout error
三者可以互相搭配。
Production 最大的風險之一不是:
Pod 掛掉
因為 Pod 可以重建。
真正危險的是:
Data 掛掉
例如:
Database 被誤刪
資料表 DROP
Storage 損壞
勒索軟體
程式 Bug 大量更新錯誤資料
所以 Production 必須有:
Backup
例如 MariaDB / MySQL:
Production Database
│
▼
Daily Backup
│
▼
Backup Storage
可能包含:
每日 Full Backup
Incremental Backup
Binlog
但最重要的不是:
有沒有 Backup
而是:
能不能 Restore
假設系統每天都說:
Backup Success
結果真的需要復原時:
Backup File corrupted
那等於沒有 Backup。
所以 Production 應該定期做:
Restore Test
例如:
Backup
│
▼
建立測試 Database
│
▼
Restore
│
▼
驗證資料
談 Backup 時通常還會遇到兩個概念。
RPO
Recovery Point Objective。
簡單說:
最多可以接受丟多少資料?
例如:
RPO = 1 hour
代表最多接受:
遺失最近一小時資料
RTO
Recovery Time Objective。
代表:
系統掛掉後,多久內要恢復?
例如:
RTO = 30 minutes
代表:
30 分鐘內
要把服務恢復
所以:
RPO
→ 可以丟多少資料?
RTO
→ 可以停多久?
前面我們已經讓 Backend 有:
Pod 1
Pod 2
Pod 3
但如果:
所有 Pod
都放在 Node 1
結果:
Node 1 掛掉
三個 Pod 一起消失。
因此 Production 會希望:
Pod 分散到不同 Node
甚至:
不同 Availability Zone
例如:
Zone A
└── Pod 1
Zone B
└── Pod 2
Zone C
└── Pod 3
這樣單一機房或 Node 出問題時,仍然有服務可以運作。
Backend 有十個 Pod:
Pod × 10
如果最後全部連:
一台 Database
那 Database 就是:
Single Point of Failure
也就是單點故障。
架構:
10 Backend Pods
│
▼
Database
X
Database 一掛:
Backend 全部不能用
所以 Production Database 可能需要:
Primary
Replica
Failover
例如:
Primary DB
│
▼
Replica DB
Primary 掛掉時,可以進行 Failover。
Production 通常還需要處理:
TLS
Secrets
IAM
Network Policy
Container Security
Dependency Vulnerability
例如 Password 不應該:
寫死在 Git
API Key 不應該:
放進 Docker Image
而應該使用:
Secret Manager
Kubernetes Secret
Cloud Secret Service
Production 權限應該遵守:
Least Privilege
也就是:
只給真正需要的權限。
例如 Backend 只需要:
讀寫自己的 Database
就不應該給:
整個 Cluster Admin
CI/CD 只需要:
更新某個 Deployment
也不一定需要:
刪除整個 Kubernetes Cluster
這樣可以降低帳號或 Token 洩漏時的影響範圍。
Production API 暴露到 Internet 後,也要防止:
大量惡意 Request
程式 Bug 無限重試
爬蟲
暴力登入
所以可以設定:
100 requests / minute
超過:
429 Too Many Requests
Rate Limit 可以做在:
Ingress
API Gateway
Backend
Redis
Production 很重要但容易忽略的是:
Timeout
假設 Backend 呼叫第三方:
Payment API
但是 Payment API 卡住。
如果沒有 Timeout:
Backend Request
一直等
一直等
一直等
Connection 會越積越多。
所以應該設定:
HTTP Timeout
Database Timeout
Redis Timeout
例如:
Payment API
最多等 5 秒
超過就 Fail。
如果 Request Fail:
Retry
可能有幫助。
但是如果所有 Pod 同時:
失敗
→ Retry
→ 失敗
→ Retry
就可能造成:
Retry Storm
例如原本:
1000 Request
全部重試三次:
3000 Request
反而讓已經過載的服務更慘。
所以 Production 常搭配:
Timeout
Retry
Exponential Backoff
Circuit Breaker
假設 Payment Service 已經掛了。
如果 Backend 還一直:
Request
Request
Request
Request
只會浪費資源。
Circuit Breaker 的概念是:
失敗太多
│
▼
暫時停止呼叫
等一段時間後:
再試看看
很像家裡的:
電路斷路器
故障時先切斷,避免問題繼續擴大。
Kubernetes Scale In:
Pod 5
→ 刪除
但是 Pod 5 此時可能正在處理:
Order Request
如果直接 Kill:
Request 中斷
所以 Application 應該支援:
Graceful Shutdown
概念:
Kubernetes
通知 Pod 要關閉
│
▼
停止接新 Request
│
▼
等待目前 Request 完成
│
▼
Shutdown
這對:
Rolling Update
HPA Scale In
Node Maintenance
都非常重要。
假設現在:
3 Pods
維護 Node 時,如果一次把:
3 Pods 全部停止
服務就中斷。
Kubernetes 可以使用:
PodDisruptionBudget
例如要求:
至少保持 2 個 Pod 可用
這樣維護時就不會一次全部停止。
Production Deployment 也不一定只有:
Rolling Update
還有其他方式。
例如:
Blue-Green Deployment
Blue
v1.0
Green
v1.1
先把 Green 完整部署好:
Traffic
→ Blue
測試完之後直接切:
Traffic
→ Green
如果出問題:
切回 Blue
另一種是:
Canary Deployment
例如:
95% Traffic
→ v1.0
5% Traffic
→ v1.1
先觀察:
Error Rate
Latency
Business Metrics
如果沒問題,再:
25%
50%
100%
逐步放大。
這樣可以降低一次部署新版造成大規模事故的風險。
Production 通常也不會只有:
Production
而是:
Development
Staging
Production
例如:
dev.example.com
staging.example.com
example.com
不同環境應該有各自:
Config
Secret
Database
Redis
避免:
測試程式
直接操作 Production Database
前面只有:
/health
但 Production 可以分:
/liveness
/readiness
例如:
/liveness
只檢查:
Application Process
還活著嗎?
而:
/readiness
檢查:
現在是否適合接流量?
例如 Database 暫時無法使用:
Liveness
200
Readiness
503
代表:
程式沒死
但暫時不要送 Request 給它
Production 不只需要監控技術指標。
例如:
CPU
Memory
HTTP 500
還要看:
Order Success Rate
Payment Success Rate
Booking Success Rate
Login Failure Rate
例如:
CPU = 30%
HTTP 200 = 正常
看起來系統很好。
但:
Payment Success
99%
↓
40%
這才可能是真正嚴重的 Production Incident。
所以 Monitoring 最終還要包含:
Business Metrics
可以變成:
Users
│
▼
Ingress
│
▼
Service
│
▼
FastAPI Pods
│ │ │
┌───────┘ │ └────────┐
▼ ▼ ▼
Logs Metrics Traces
│ │ │
▼ ▼ ▼
Log System Prometheus Trace System
│ │ │
└────────────┼─────────────┘
│
▼
Grafana
│
▼
Alert
│
▼
Engineer / On-call
除了 Observability:
Application
│
┌───────────┼───────────┐
▼ ▼ ▼
Redis Database Storage
│
▼
Backup
│
▼
Restore Testing
另外:
High Availability
Auto Scaling
Rate Limiting
Timeout
Retry
Circuit Breaker
Graceful Shutdown
一起提高整體可靠性。
這裡有一個很重要的觀念。
用了 Kubernetes:
不代表 Production 就完成了
Kubernetes 解決的主要是:
Container Orchestration
例如:
Pod Scheduling
Self Healing
Scaling
Rolling Update
Service Discovery
但是:
資料備份
安全性
監控
告警
Log
Tracing
Database HA
災難復原
仍然需要另外設計。
所以 Kubernetes 只是 Production Architecture 中的一部分。
到這裡,可以把 Production 需要的能力整理成:
1. Deployment
CI/CD
Docker
Kubernetes
2. Scalability
Load Balancer
HPA
Cluster Autoscaler
3. Observability
Logs
Metrics
Traces
4. Reliability
Health Check
Retry
Timeout
Circuit Breaker
5. Availability
Multi Pod
Multi Node
Database Replica
6. Data Protection
Backup
Restore
RPO
RTO
7. Security
Secret
IAM
TLS
Least Privilege
8. Operations
Alert
Rollback
Incident Response
這些東西加起來,才比較接近真正完整的 Production System。
最早:
User
│
▼
Nginx
│
▼
FastAPI
│
▼
Database
逐漸演進成:
Users
│
▼
DNS
│
▼
Load Balancer
│
▼
Ingress
│
▼
Service
│
▼
┌── FastAPI Pods ──┐
│ │
▼ ▼
Redis Database
│
▼
Backup
部署流程:
Developer
│
▼
Git
│
▼
CI/CD
│
▼
Docker Image
│
▼
Registry
│
▼
Kubernetes
監控:
Application
│
┌──┼──┐
▼ ▼ ▼
Log Metric Trace
│ │ │
└───┼────┘
▼
Observability
│
▼
Alert
擴充:
Traffic
│
▼
HPA
│
▼
More Pods
│
▼
Cluster Autoscaler
│
▼
More Nodes
所以一個 Production System 真正要做到的,不只是:
「我的程式有跑起來。」
而是:
「我的程式可以被穩定部署、被觀察、被擴充、發生問題能被發現,而且資料與服務都可以被恢復。」
這才是從「部署 Application」進一步走向「營運 Production System」。