iT邦幫忙

2026 iThome 鐵人賽

DAY 12
0
Kubernetes

探討k8s部署方式系列 第 12 篇

如何將log實際接入到後端?以FastAPI為例[Day12]

  • 分享至 

  • xImage
  •  

二百一十二、Logging、Metrics、Tracing 怎麼真正接進 FastAPI?

前面我們知道 Production 通常需要三種 Observability Data:

Logs
Metrics
Traces

但它們不是裝一套 Grafana 就會自動全部出現。

Application 本身必須先產生這些資料。

以 FastAPI 為例:

FastAPI
 │
 ├── Logs
 │
 ├── Metrics
 │
 └── Traces

接著再把資料送出去。

例如:

FastAPI
 │
 ├── Logs ─────→ Loki / Elasticsearch
 │
 ├── Metrics ──→ Prometheus
 │
 └── Traces ───→ OpenTelemetry Collector
                         │
                         ▼
                   Jaeger / Tempo

最後再透過:

Grafana

或其他平台查看。


二百一十三、先從 Logging 開始

FastAPI 最基本可以直接使用 Python:

import logging

logger = logging.getLogger(__name__)

logger.info("application started")

例如 API:

@app.get("/products")
async def products():
    logger.info("get products")

    return {
        "data": []
    }

呼叫:

GET /products

就可能看到:

INFO get products

但 Production 如果只記:

get products

資訊還遠遠不夠。


二百一十四、每一次 Request 應該記錄什麼?

例如:

GET /api/products

至少可以考慮記:

timestamp
level
method
path
status_code
duration
request_id
service

例如:

{
  "timestamp": "2026-09-26T20:30:00Z",
  "level": "INFO",
  "service": "backend-api",
  "request_id": "52ee61ad",
  "method": "GET",
  "path": "/api/products",
  "status_code": 200,
  "duration_ms": 83
}

這種:

JSON Log

比:

Request successful

更適合 Production。

因為後面可以搜尋:

status_code >= 500

duration_ms > 1000

request_id = 52ee61ad

path = /api/products

二百一十五、FastAPI Middleware 統一記錄 Request

FastAPI Middleware 會在每一個 HTTP Request 進入 Path Operation 前執行,也能在 Response 回來後再次處理,因此非常適合統一記錄 Request duration、狀態碼等資訊。

例如:

import logging
import time
import uuid

from fastapi import FastAPI, Request

app = FastAPI()

logger = logging.getLogger("backend")


@app.middleware("http")
async def request_logging_middleware(
    request: Request,
    call_next
):
    request_id = str(uuid.uuid4())

    start_time = time.perf_counter()

    response = await call_next(request)

    duration = (
        time.perf_counter() - start_time
    ) * 1000

    logger.info(
        "request completed",
        extra={
            "request_id": request_id,
            "method": request.method,
            "path": request.url.path,
            "status_code": response.status_code,
            "duration_ms": round(duration, 2),
        },
    )

    response.headers["X-Request-ID"] = request_id

    return response

現在所有 Request:

GET /products

POST /orders

GET /users

都會經過:

Request Logging Middleware

而不需要在每一支 API 裡重複寫:

logger.info(...)

二百一十六、Request ID

前面的:

request_id = str(uuid.uuid4())

非常重要。

例如使用者呼叫:

POST /orders

我們產生:

request_id
=
52ee61ad

接下來所有相關 Log 都可以帶:

52ee61ad

例如:

52ee61ad
receive order request

52ee61ad
query customer

52ee61ad
call payment API

52ee61ad
payment timeout

52ee61ad
return HTTP 500

這樣工程師只要搜尋:

52ee61ad

就能看到同一次 Request 發生什麼事情。


二百一十七、Request ID 最好能一路傳下去

假設架構:

User
 │
 ▼
Ingress
 │
 ▼
Order API
 │
 ▼
Payment API

可以將:

X-Request-ID

放在 HTTP Header。

例如:

X-Request-ID: 52ee61ad

Order API 呼叫 Payment API 時,也把這個 Header 傳過去。

於是:

Ingress

request_id=52ee61ad
      │
      ▼
Order API

request_id=52ee61ad
      │
      ▼
Payment API

request_id=52ee61ad

不同服務的 Log 就能被串起來。


二百一十八、Error Log

除了正常 Request,也要記 Exception。

例如:

try:
    result = await payment_service.pay()

except Exception:
    logger.exception(
        "payment failed",
        extra={
            "request_id": request_id
        }
    )

    raise

logger.exception() 會把 Exception Stack Trace 一起記錄。

這樣 Production 發生:

HTTP 500

才有辦法找到真正原因,例如:

ConnectionTimeout
DatabaseError
ValidationError

二百一十九、Container 裡 Log 放哪裡?

傳統 Server 常常會:

/var/log/backend/app.log

但 Container / Kubernetes 通常更適合讓 Application:

寫到 stdout / stderr

例如:

FastAPI Container
     │
     ▼
stdout

Kubernetes 就能看到:

kubectl logs <pod-name>

然後另外由:

Fluent Bit
Promtail
其他 Log Agent

把每個 Pod 的 Log 收走。


二百二十、Centralized Logging

Production 架構就可能是:

Pod 1 ─┐
Pod 2 ─┤
Pod 3 ─┼──→ Fluent Bit
Pod 4 ─┘
             │
             ▼
       Elasticsearch

             或

            Loki

然後:

Kibana

或:

Grafana

負責查詢。

因此 Application 不需要知道:

Elasticsearch 在哪裡?

Application 只負責:

產生正確的 Structured Log

Log Collector 再負責收集。


二百二十一、接著是 Metrics

Logging 通常是一個 Event 一筆資料。

例如:

Request A → 200

Request B → 200

Request C → 500

但是如果我要知道:

過去五分鐘
500 Error 比例是多少?

一直掃 Log 並不是最適合的方法。

所以需要 Metrics。

例如:

http_requests_total

http_request_duration_seconds

http_errors_total

二百二十二、FastAPI 加 Prometheus Metrics

Prometheus 官方 Python client 可以建立一個 ASGI metrics application,再掛到 FastAPI 的 /metrics;官方也特別提供 FastAPI + Gunicorn 的做法。

安裝:

pip install prometheus-client

然後:

from fastapi import FastAPI
from prometheus_client import make_asgi_app

app = FastAPI()

metrics_app = make_asgi_app()

app.mount("/metrics", metrics_app)

現在:

GET /metrics

就會回傳 Prometheus 格式的 Metrics。


二百二十三、建立自己的 HTTP Metrics

例如想記錄:

API 被呼叫幾次

可以建立 Counter:

from prometheus_client import Counter

REQUEST_COUNT = Counter(
    "http_requests_total",
    "Total HTTP requests",
    [
        "method",
        "path",
        "status",
    ],
)

Middleware:

@app.middleware("http")
async def metrics_middleware(
    request: Request,
    call_next
):
    response = await call_next(request)

    REQUEST_COUNT.labels(
        method=request.method,
        path=request.url.path,
        status=response.status_code,
    ).inc()

    return response

現在:

GET /products
GET /products
GET /users
POST /orders

Prometheus 就可以累積:

http_requests_total{
  method="GET",
  path="/products",
  status="200"
} 2

二百二十四、Counter 是什麼?

Counter 適合:

只會增加的數值

例如:

Request Count

Login Count

Order Count

Error Count

例如:

0
↓
1
↓
2
↓
3

通常不會:

3
↓
2

二百二十五、Histogram:記錄 Response Time

Production 更重要的通常是:

API 到底有多慢?

可以使用 Histogram。

例如:

from prometheus_client import Histogram

REQUEST_DURATION = Histogram(
    "http_request_duration_seconds",
    "HTTP request duration",
    ["method", "path"],
)

Middleware:

@app.middleware("http")
async def metrics_middleware(
    request: Request,
    call_next
):
    start = time.perf_counter()

    response = await call_next(request)

    duration = time.perf_counter() - start

    REQUEST_DURATION.labels(
        method=request.method,
        path=request.url.path,
    ).observe(duration)

    return response

之後 Prometheus 可以計算:

P50

P95

P99

Response Time。


二百二十六、不要把 User ID 放進 Metrics Label

例如這種設定:

user_id=123

user_id=124

user_id=125
...

如果有一百萬使用者,就可能建立大量不同的 time series。

因此 Metrics Label 應該盡量使用有限種類的值,例如:

method

status

route

service

不要輕易使用:

user_id

email

order_id

request_id

這些高基數資料。

這也是 Logs 和 Metrics 的差異之一。


二百二十七、Prometheus 怎麼拿到 /metrics?

FastAPI:

GET /metrics

只是把資料暴露出來。

真正的 Prometheus Server 會定期:

GET /metrics

也就是:

FastAPI Pod
    ▲
    │ scrape
    │
Prometheus

例如:

每 15 秒

抓一次。

資料最後存在 Prometheus:

Prometheus
 │
 ▼
Time Series Database

Grafana 再去查 Prometheus:

FastAPI
   │
   ▼
/metrics
   ▲
   │
Prometheus
   │
   ▼
Grafana

二百二十八、Kubernetes 多 Pod 時怎麼辦?

假設:

Pod 1
Pod 2
Pod 3

每個 Pod 都有:

/metrics

Prometheus 會分別收集:

Pod 1 /metrics

Pod 2 /metrics

Pod 3 /metrics

然後彙整。

這樣即使 HPA:

3 Pods
↓
10 Pods

Prometheus 仍然可以持續監控新 Pod。


二百二十九、Gunicorn 多 Worker 要注意

如果單一 Container 裡使用:

Gunicorn

而且:

workers > 1

Metrics 就要注意 multiprocessing。

Prometheus Python client 官方的 FastAPI + Gunicorn 文件提供了 CollectorRegistry 搭配 MultiProcessCollector 的方式來收集多 worker metrics。

概念就是:

Container

Gunicorn
├── Worker 1
├── Worker 2
├── Worker 3
└── Worker 4

      │
      ▼
合併 Metrics
      │
      ▼
/metrics

否則可能只看到部分 Worker 的資料。


二百三十、接著是 Tracing

Logging 可以看到:

Payment timeout

Metrics 可以看到:

P95 latency
500ms → 3s

但是還是不一定知道:

這三秒到底花在哪?

這時就需要 Tracing。


二百三十一、OpenTelemetry

現在常見做法是使用:

OpenTelemetry

OpenTelemetry Python 可以產生:

Traces
Metrics
Logs

目前官方列出的 Python 狀態中,Tracing 與 Metrics 為 Stable,而 Logs SDK 仍標示為 Development。

因此一個常見策略是:

Logging
→ Python logging + Log Collector

Metrics
→ Prometheus

Tracing
→ OpenTelemetry

之後也可以逐步往統一的 OpenTelemetry 架構演進。


二百三十二、安裝 OpenTelemetry

基本套件:

pip install \
  opentelemetry-api \
  opentelemetry-sdk

OpenTelemetry 官方也提供不同 Library 的 instrumentation package,例如 HTTPX 可以透過 instrumentation library 自動建立相關 telemetry。

FastAPI 專案常會再安裝:

opentelemetry-instrumentation-fastapi

opentelemetry-instrumentation-httpx

opentelemetry-instrumentation-sqlalchemy

這些分別可以協助追蹤:

FastAPI Request

Outbound HTTP

Database Query

二百三十三、Tracing 最核心的概念:Trace 與 Span

假設:

POST /orders

整個 Request 是一個:

Trace

裡面可能包含很多:

Span

例如:

POST /orders
Trace

├── validate_order
│      20ms
│
├── query_database
│      60ms
│
├── call_payment
│      1500ms
│
└── save_order
       50ms

因此馬上可以看出:

call_payment

花了最多時間。


二百三十四、自動 Instrument FastAPI

概念上可以:

from fastapi import FastAPI

from opentelemetry.instrumentation.fastapi import (
    FastAPIInstrumentor,
)

app = FastAPI()

FastAPIInstrumentor.instrument_app(app)

這樣 HTTP Request 就可以自動建立 Span。

例如:

GET /products

產生:

HTTP GET /products

Span。

Instrumentation library 的用途就是避免我們為每個框架或第三方 library 全部手寫 observability code。


二百三十五、手動建立 Business Span

自動 Instrumentation 可以知道:

POST /orders

花了 2 秒。

但不知道你的 Business Logic 哪一段最慢。

這時可以手動建立 Span。

例如:

from opentelemetry import trace

tracer = trace.get_tracer(__name__)


async def create_order():
    with tracer.start_as_current_span(
        "validate_order"
    ):
        await validate_order()

    with tracer.start_as_current_span(
        "process_payment"
    ):
        await process_payment()

    with tracer.start_as_current_span(
        "save_order"
    ):
        await save_order()

OpenTelemetry 官方 Python SDK 就是透過 TracerProvider、Tracer 與 Span 建立這類 Trace。

最後 Trace:

POST /orders
│
├── validate_order
│
├── process_payment
│
└── save_order

會更加有意義。


二百三十六、Trace Context 要能跨 Service

假設:

Order Service
      │
      ▼
Payment Service

如果兩邊都有 OpenTelemetry instrumentation:

Order API
Trace ABC
   │
   ▼
Payment API
Trace ABC

Trace Context 可以透過 HTTP Header 傳遞。

於是觀察平台就知道:

Order Service

跟

Payment Service

屬於同一次 Request。

這是 Distributed Tracing 最重要的能力之一。


二百三十七、HTTPX 也可以自動追蹤

例如 FastAPI 會:

async with httpx.AsyncClient() as client:
    await client.get(
        "https://payment.example.com"
    )

OpenTelemetry 官方文件說明可以為 HTTPX 安裝 instrumentation library,自動對 HTTP client request 建立 telemetry。

概念就會變成:

FastAPI
 │
 ├── POST /orders
 │
 └── HTTP GET payment.example.com

不需要每一次:

client.get()

旁邊都手動建立 Span。


二百三十八、Tracing 資料最後送去哪?

Application 產生 Span 之後,需要 Export。

Production 常見架構是:

FastAPI
   │
   │ OTLP
   ▼
OpenTelemetry Collector
   │
   ▼
Trace Backend

OpenTelemetry 官方建議 Production 將 telemetry 傳給 OpenTelemetry Collector,再由 Collector Export 到實際 backend;OTLP 可以透過 gRPC 或 HTTP 傳輸。

例如:

FastAPI
   │
   ▼
OTel Collector
   │
   ├──→ Jaeger
   │
   └──→ Tempo

二百三十九、為什麼中間還需要 Collector?

也可以:

FastAPI
→ Jaeger

但是加入 Collector 後:

FastAPI
     │
     ▼
OTel Collector
     │
 ┌───┼────────┐
 ▼   ▼        ▼
Jaeger Tempo Vendor

Application 就不需要知道後面到底用哪一套平台。

未來:

Jaeger
→ Tempo

Application 不一定需要修改。

Collector 還可以負責:

Batch

Filter

Sampling

Enrichment

Routing

所以架構會更有彈性。


二百四十、OTLP

OpenTelemetry 常使用:

OTLP

也就是:

OpenTelemetry Protocol

例如 Collector:

4317
→ OTLP/gRPC

4318
→ OTLP/HTTP

OpenTelemetry 官方文件目前就是以這兩種方式作為主要 OTLP exporter 配置。

Kubernetes 裡可能:

FastAPI Pod
    │
    ▼
otel-collector:4317

二百四十一、Logging、Metrics、Tracing 怎麼串起來?

真正好用的 Observability 並不是三套彼此無關的系統。

假設 Grafana 發現:

P95 latency

300ms
↓
3000ms

第一步:

Metrics

告訴你:

系統變慢了

接著查 Trace:

POST /orders
        │
        ├── DB
        │    50ms
        │
        └── Payment
             2700ms

發現:

Payment

特別慢。

最後再利用:

trace_id
request_id

去查 Logs:

payment timeout

connection pool exhausted

third-party returned 503

最後找到真正原因。

完整流程就是:

Metrics
   │
   │ 發現問題
   ▼
Tracing
   │
   │ 定位問題在哪個 Service
   ▼
Logging
   │
   │ 找詳細 Error
   ▼
Root Cause

二百四十二、Trace ID 也可以放進 Log

如果目前 Span:

trace_id
=
ABC123

Log 最好也包含:

{
  "level": "ERROR",
  "message": "payment timeout",
  "request_id": "REQ789",
  "trace_id": "ABC123"
}

這樣在 Trace UI 找到:

ABC123

之後,可以直接搜尋:

trace_id=ABC123

取得相關 Log。

這就是 Observability 資料互相串接的關鍵。


二百四十三、Production FastAPI Observability 架構

完整架構可以變成:

                         User
                          │
                          ▼
                       Ingress
                          │
                          ▼
                       Service
                          │
                          ▼
                   FastAPI Pods
                          │
           ┌──────────────┼──────────────┐
           │              │              │
           ▼              ▼              ▼
         Logs          Metrics         Traces
           │              │              │
           ▼              ▼              ▼
      Fluent Bit      Prometheus    OTel Collector
           │              │              │
           ▼              │              ▼
      Loki / ES           │         Jaeger / Tempo
           │              │              │
           └──────────────┼──────────────┘
                          ▼
                        Grafana

二百四十四、FastAPI 本身應該負責什麼?

Application 不應該負責:

保存所有 Log

保存所有 Metrics

保存所有 Traces

FastAPI 應該負責:

產生 Log

暴露 /metrics

產生 Trace

而 Infrastructure:

Log Collector

Prometheus

OpenTelemetry Collector

負責收集。

Backend 就可以保持比較單純:

                    FastAPI

         ┌────────────┼────────────┐
         │            │            │
         ▼            ▼            ▼
       stdout       /metrics      OTLP
         │            │            │
         ▼            ▼            ▼
       Logs         Metrics      Traces

二百四十五、先做到哪個程度就夠?

如果是一個剛開始進 Production 的 FastAPI 專案,不一定一開始就要全部做完。

可以分階段。

第一階段:

Structured Logging

Request ID

HTTP Status

Response Time

第二階段:

Prometheus

/metrics

Request Count

Error Rate

Latency

第三階段:

OpenTelemetry

FastAPI Tracing

HTTP Client Tracing

Database Tracing

第四階段:

Logs
Metrics
Traces

透過

request_id / trace_id

彼此關聯

這樣會比一次導入所有 Observability 工具更容易掌握。


二百四十六、最後整理

Logging:

FastAPI
   │
   ▼
Python Logging
   │
   ▼
stdout
   │
   ▼
Fluent Bit / Promtail
   │
   ▼
Loki / Elasticsearch

Metrics:

FastAPI
   │
   ▼
/metrics
   ▲
   │
Prometheus
   │
   ▼
Grafana

Tracing:

FastAPI
   │
   ▼
OpenTelemetry
   │
   │ OTLP
   ▼
OTel Collector
   │
   ▼
Jaeger / Tempo

三者的功能:

Logging
→ 這次到底發生什麼事情?

Metrics
→ 整個系統現在健康嗎?

Tracing
→ 這一次 Request 的時間花在哪裡?

真正的 Production 排障流程則可以是:

Alert
 │
 ▼
Metrics
 │
 │ 發現 Error / Latency 異常
 ▼
Tracing
 │
 │ 找到慢在哪個 Service
 ▼
Logging
 │
 │ 找到 Exception / Error Detail
 ▼
確認 Root Cause
 │
 ▼
修正
 │
 ▼
CI/CD Deploy
 │
 ▼
Metrics 驗證是否恢復

這時 Observability 就不只是:

「有裝 Grafana」

而是形成一套真正可以用來:

發現問題
定位問題
理解問題
驗證修復

的 Production 維運能力。


上一篇
上線到prod版本會遇到的問題[Day11]
下一篇
從零開始實作部署[Day13]
系列文
探討k8s部署方式 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言