前面我們知道 Production 通常需要三種 Observability Data:
Logs
Metrics
Traces
但它們不是裝一套 Grafana 就會自動全部出現。
Application 本身必須先產生這些資料。
以 FastAPI 為例:
FastAPI
│
├── Logs
│
├── Metrics
│
└── Traces
接著再把資料送出去。
例如:
FastAPI
│
├── Logs ─────→ Loki / Elasticsearch
│
├── Metrics ──→ Prometheus
│
└── Traces ───→ OpenTelemetry Collector
│
▼
Jaeger / Tempo
最後再透過:
Grafana
或其他平台查看。
FastAPI 最基本可以直接使用 Python:
import logging
logger = logging.getLogger(__name__)
logger.info("application started")
例如 API:
@app.get("/products")
async def products():
logger.info("get products")
return {
"data": []
}
呼叫:
GET /products
就可能看到:
INFO get products
但 Production 如果只記:
get products
資訊還遠遠不夠。
例如:
GET /api/products
至少可以考慮記:
timestamp
level
method
path
status_code
duration
request_id
service
例如:
{
"timestamp": "2026-09-26T20:30:00Z",
"level": "INFO",
"service": "backend-api",
"request_id": "52ee61ad",
"method": "GET",
"path": "/api/products",
"status_code": 200,
"duration_ms": 83
}
這種:
JSON Log
比:
Request successful
更適合 Production。
因為後面可以搜尋:
status_code >= 500
duration_ms > 1000
request_id = 52ee61ad
path = /api/products
FastAPI Middleware 會在每一個 HTTP Request 進入 Path Operation 前執行,也能在 Response 回來後再次處理,因此非常適合統一記錄 Request duration、狀態碼等資訊。
例如:
import logging
import time
import uuid
from fastapi import FastAPI, Request
app = FastAPI()
logger = logging.getLogger("backend")
@app.middleware("http")
async def request_logging_middleware(
request: Request,
call_next
):
request_id = str(uuid.uuid4())
start_time = time.perf_counter()
response = await call_next(request)
duration = (
time.perf_counter() - start_time
) * 1000
logger.info(
"request completed",
extra={
"request_id": request_id,
"method": request.method,
"path": request.url.path,
"status_code": response.status_code,
"duration_ms": round(duration, 2),
},
)
response.headers["X-Request-ID"] = request_id
return response
現在所有 Request:
GET /products
POST /orders
GET /users
都會經過:
Request Logging Middleware
而不需要在每一支 API 裡重複寫:
logger.info(...)
前面的:
request_id = str(uuid.uuid4())
非常重要。
例如使用者呼叫:
POST /orders
我們產生:
request_id
=
52ee61ad
接下來所有相關 Log 都可以帶:
52ee61ad
例如:
52ee61ad
receive order request
52ee61ad
query customer
52ee61ad
call payment API
52ee61ad
payment timeout
52ee61ad
return HTTP 500
這樣工程師只要搜尋:
52ee61ad
就能看到同一次 Request 發生什麼事情。
假設架構:
User
│
▼
Ingress
│
▼
Order API
│
▼
Payment API
可以將:
X-Request-ID
放在 HTTP Header。
例如:
X-Request-ID: 52ee61ad
Order API 呼叫 Payment API 時,也把這個 Header 傳過去。
於是:
Ingress
request_id=52ee61ad
│
▼
Order API
request_id=52ee61ad
│
▼
Payment API
request_id=52ee61ad
不同服務的 Log 就能被串起來。
除了正常 Request,也要記 Exception。
例如:
try:
result = await payment_service.pay()
except Exception:
logger.exception(
"payment failed",
extra={
"request_id": request_id
}
)
raise
logger.exception() 會把 Exception Stack Trace 一起記錄。
這樣 Production 發生:
HTTP 500
才有辦法找到真正原因,例如:
ConnectionTimeout
DatabaseError
ValidationError
傳統 Server 常常會:
/var/log/backend/app.log
但 Container / Kubernetes 通常更適合讓 Application:
寫到 stdout / stderr
例如:
FastAPI Container
│
▼
stdout
Kubernetes 就能看到:
kubectl logs <pod-name>
然後另外由:
Fluent Bit
Promtail
其他 Log Agent
把每個 Pod 的 Log 收走。
Production 架構就可能是:
Pod 1 ─┐
Pod 2 ─┤
Pod 3 ─┼──→ Fluent Bit
Pod 4 ─┘
│
▼
Elasticsearch
或
Loki
然後:
Kibana
或:
Grafana
負責查詢。
因此 Application 不需要知道:
Elasticsearch 在哪裡?
Application 只負責:
產生正確的 Structured Log
Log Collector 再負責收集。
Logging 通常是一個 Event 一筆資料。
例如:
Request A → 200
Request B → 200
Request C → 500
但是如果我要知道:
過去五分鐘
500 Error 比例是多少?
一直掃 Log 並不是最適合的方法。
所以需要 Metrics。
例如:
http_requests_total
http_request_duration_seconds
http_errors_total
Prometheus 官方 Python client 可以建立一個 ASGI metrics application,再掛到 FastAPI 的 /metrics;官方也特別提供 FastAPI + Gunicorn 的做法。
安裝:
pip install prometheus-client
然後:
from fastapi import FastAPI
from prometheus_client import make_asgi_app
app = FastAPI()
metrics_app = make_asgi_app()
app.mount("/metrics", metrics_app)
現在:
GET /metrics
就會回傳 Prometheus 格式的 Metrics。
例如想記錄:
API 被呼叫幾次
可以建立 Counter:
from prometheus_client import Counter
REQUEST_COUNT = Counter(
"http_requests_total",
"Total HTTP requests",
[
"method",
"path",
"status",
],
)
Middleware:
@app.middleware("http")
async def metrics_middleware(
request: Request,
call_next
):
response = await call_next(request)
REQUEST_COUNT.labels(
method=request.method,
path=request.url.path,
status=response.status_code,
).inc()
return response
現在:
GET /products
GET /products
GET /users
POST /orders
Prometheus 就可以累積:
http_requests_total{
method="GET",
path="/products",
status="200"
} 2
Counter 適合:
只會增加的數值
例如:
Request Count
Login Count
Order Count
Error Count
例如:
0
↓
1
↓
2
↓
3
通常不會:
3
↓
2
Production 更重要的通常是:
API 到底有多慢?
可以使用 Histogram。
例如:
from prometheus_client import Histogram
REQUEST_DURATION = Histogram(
"http_request_duration_seconds",
"HTTP request duration",
["method", "path"],
)
Middleware:
@app.middleware("http")
async def metrics_middleware(
request: Request,
call_next
):
start = time.perf_counter()
response = await call_next(request)
duration = time.perf_counter() - start
REQUEST_DURATION.labels(
method=request.method,
path=request.url.path,
).observe(duration)
return response
之後 Prometheus 可以計算:
P50
P95
P99
Response Time。
例如這種設定:
user_id=123
user_id=124
user_id=125
...
如果有一百萬使用者,就可能建立大量不同的 time series。
因此 Metrics Label 應該盡量使用有限種類的值,例如:
method
status
route
service
不要輕易使用:
user_id
email
order_id
request_id
這些高基數資料。
這也是 Logs 和 Metrics 的差異之一。
/metrics?FastAPI:
GET /metrics
只是把資料暴露出來。
真正的 Prometheus Server 會定期:
GET /metrics
也就是:
FastAPI Pod
▲
│ scrape
│
Prometheus
例如:
每 15 秒
抓一次。
資料最後存在 Prometheus:
Prometheus
│
▼
Time Series Database
Grafana 再去查 Prometheus:
FastAPI
│
▼
/metrics
▲
│
Prometheus
│
▼
Grafana
假設:
Pod 1
Pod 2
Pod 3
每個 Pod 都有:
/metrics
Prometheus 會分別收集:
Pod 1 /metrics
Pod 2 /metrics
Pod 3 /metrics
然後彙整。
這樣即使 HPA:
3 Pods
↓
10 Pods
Prometheus 仍然可以持續監控新 Pod。
如果單一 Container 裡使用:
Gunicorn
而且:
workers > 1
Metrics 就要注意 multiprocessing。
Prometheus Python client 官方的 FastAPI + Gunicorn 文件提供了 CollectorRegistry 搭配 MultiProcessCollector 的方式來收集多 worker metrics。
概念就是:
Container
Gunicorn
├── Worker 1
├── Worker 2
├── Worker 3
└── Worker 4
│
▼
合併 Metrics
│
▼
/metrics
否則可能只看到部分 Worker 的資料。
Logging 可以看到:
Payment timeout
Metrics 可以看到:
P95 latency
500ms → 3s
但是還是不一定知道:
這三秒到底花在哪?
這時就需要 Tracing。
現在常見做法是使用:
OpenTelemetry
OpenTelemetry Python 可以產生:
Traces
Metrics
Logs
目前官方列出的 Python 狀態中,Tracing 與 Metrics 為 Stable,而 Logs SDK 仍標示為 Development。
因此一個常見策略是:
Logging
→ Python logging + Log Collector
Metrics
→ Prometheus
Tracing
→ OpenTelemetry
之後也可以逐步往統一的 OpenTelemetry 架構演進。
基本套件:
pip install \
opentelemetry-api \
opentelemetry-sdk
OpenTelemetry 官方也提供不同 Library 的 instrumentation package,例如 HTTPX 可以透過 instrumentation library 自動建立相關 telemetry。
FastAPI 專案常會再安裝:
opentelemetry-instrumentation-fastapi
opentelemetry-instrumentation-httpx
opentelemetry-instrumentation-sqlalchemy
這些分別可以協助追蹤:
FastAPI Request
Outbound HTTP
Database Query
假設:
POST /orders
整個 Request 是一個:
Trace
裡面可能包含很多:
Span
例如:
POST /orders
Trace
├── validate_order
│ 20ms
│
├── query_database
│ 60ms
│
├── call_payment
│ 1500ms
│
└── save_order
50ms
因此馬上可以看出:
call_payment
花了最多時間。
概念上可以:
from fastapi import FastAPI
from opentelemetry.instrumentation.fastapi import (
FastAPIInstrumentor,
)
app = FastAPI()
FastAPIInstrumentor.instrument_app(app)
這樣 HTTP Request 就可以自動建立 Span。
例如:
GET /products
產生:
HTTP GET /products
Span。
Instrumentation library 的用途就是避免我們為每個框架或第三方 library 全部手寫 observability code。
自動 Instrumentation 可以知道:
POST /orders
花了 2 秒。
但不知道你的 Business Logic 哪一段最慢。
這時可以手動建立 Span。
例如:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
async def create_order():
with tracer.start_as_current_span(
"validate_order"
):
await validate_order()
with tracer.start_as_current_span(
"process_payment"
):
await process_payment()
with tracer.start_as_current_span(
"save_order"
):
await save_order()
OpenTelemetry 官方 Python SDK 就是透過 TracerProvider、Tracer 與 Span 建立這類 Trace。
最後 Trace:
POST /orders
│
├── validate_order
│
├── process_payment
│
└── save_order
會更加有意義。
假設:
Order Service
│
▼
Payment Service
如果兩邊都有 OpenTelemetry instrumentation:
Order API
Trace ABC
│
▼
Payment API
Trace ABC
Trace Context 可以透過 HTTP Header 傳遞。
於是觀察平台就知道:
Order Service
跟
Payment Service
屬於同一次 Request。
這是 Distributed Tracing 最重要的能力之一。
例如 FastAPI 會:
async with httpx.AsyncClient() as client:
await client.get(
"https://payment.example.com"
)
OpenTelemetry 官方文件說明可以為 HTTPX 安裝 instrumentation library,自動對 HTTP client request 建立 telemetry。
概念就會變成:
FastAPI
│
├── POST /orders
│
└── HTTP GET payment.example.com
不需要每一次:
client.get()
旁邊都手動建立 Span。
Application 產生 Span 之後,需要 Export。
Production 常見架構是:
FastAPI
│
│ OTLP
▼
OpenTelemetry Collector
│
▼
Trace Backend
OpenTelemetry 官方建議 Production 將 telemetry 傳給 OpenTelemetry Collector,再由 Collector Export 到實際 backend;OTLP 可以透過 gRPC 或 HTTP 傳輸。
例如:
FastAPI
│
▼
OTel Collector
│
├──→ Jaeger
│
└──→ Tempo
也可以:
FastAPI
→ Jaeger
但是加入 Collector 後:
FastAPI
│
▼
OTel Collector
│
┌───┼────────┐
▼ ▼ ▼
Jaeger Tempo Vendor
Application 就不需要知道後面到底用哪一套平台。
未來:
Jaeger
→ Tempo
Application 不一定需要修改。
Collector 還可以負責:
Batch
Filter
Sampling
Enrichment
Routing
所以架構會更有彈性。
OpenTelemetry 常使用:
OTLP
也就是:
OpenTelemetry Protocol
例如 Collector:
4317
→ OTLP/gRPC
4318
→ OTLP/HTTP
OpenTelemetry 官方文件目前就是以這兩種方式作為主要 OTLP exporter 配置。
Kubernetes 裡可能:
FastAPI Pod
│
▼
otel-collector:4317
真正好用的 Observability 並不是三套彼此無關的系統。
假設 Grafana 發現:
P95 latency
300ms
↓
3000ms
第一步:
Metrics
告訴你:
系統變慢了
接著查 Trace:
POST /orders
│
├── DB
│ 50ms
│
└── Payment
2700ms
發現:
Payment
特別慢。
最後再利用:
trace_id
request_id
去查 Logs:
payment timeout
connection pool exhausted
third-party returned 503
最後找到真正原因。
完整流程就是:
Metrics
│
│ 發現問題
▼
Tracing
│
│ 定位問題在哪個 Service
▼
Logging
│
│ 找詳細 Error
▼
Root Cause
如果目前 Span:
trace_id
=
ABC123
Log 最好也包含:
{
"level": "ERROR",
"message": "payment timeout",
"request_id": "REQ789",
"trace_id": "ABC123"
}
這樣在 Trace UI 找到:
ABC123
之後,可以直接搜尋:
trace_id=ABC123
取得相關 Log。
這就是 Observability 資料互相串接的關鍵。
完整架構可以變成:
User
│
▼
Ingress
│
▼
Service
│
▼
FastAPI Pods
│
┌──────────────┼──────────────┐
│ │ │
▼ ▼ ▼
Logs Metrics Traces
│ │ │
▼ ▼ ▼
Fluent Bit Prometheus OTel Collector
│ │ │
▼ │ ▼
Loki / ES │ Jaeger / Tempo
│ │ │
└──────────────┼──────────────┘
▼
Grafana
Application 不應該負責:
保存所有 Log
保存所有 Metrics
保存所有 Traces
FastAPI 應該負責:
產生 Log
暴露 /metrics
產生 Trace
而 Infrastructure:
Log Collector
Prometheus
OpenTelemetry Collector
負責收集。
Backend 就可以保持比較單純:
FastAPI
┌────────────┼────────────┐
│ │ │
▼ ▼ ▼
stdout /metrics OTLP
│ │ │
▼ ▼ ▼
Logs Metrics Traces
如果是一個剛開始進 Production 的 FastAPI 專案,不一定一開始就要全部做完。
可以分階段。
第一階段:
Structured Logging
Request ID
HTTP Status
Response Time
第二階段:
Prometheus
/metrics
Request Count
Error Rate
Latency
第三階段:
OpenTelemetry
FastAPI Tracing
HTTP Client Tracing
Database Tracing
第四階段:
Logs
Metrics
Traces
透過
request_id / trace_id
彼此關聯
這樣會比一次導入所有 Observability 工具更容易掌握。
Logging:
FastAPI
│
▼
Python Logging
│
▼
stdout
│
▼
Fluent Bit / Promtail
│
▼
Loki / Elasticsearch
Metrics:
FastAPI
│
▼
/metrics
▲
│
Prometheus
│
▼
Grafana
Tracing:
FastAPI
│
▼
OpenTelemetry
│
│ OTLP
▼
OTel Collector
│
▼
Jaeger / Tempo
三者的功能:
Logging
→ 這次到底發生什麼事情?
Metrics
→ 整個系統現在健康嗎?
Tracing
→ 這一次 Request 的時間花在哪裡?
真正的 Production 排障流程則可以是:
Alert
│
▼
Metrics
│
│ 發現 Error / Latency 異常
▼
Tracing
│
│ 找到慢在哪個 Service
▼
Logging
│
│ 找到 Exception / Error Detail
▼
確認 Root Cause
│
▼
修正
│
▼
CI/CD Deploy
│
▼
Metrics 驗證是否恢復
這時 Observability 就不只是:
「有裝 Grafana」
而是形成一套真正可以用來:
發現問題
定位問題
理解問題
驗證修復
的 Production 維運能力。