iT邦幫忙

2026 iThome 鐵人賽

DAY 23
0
Kubernetes

凌晨四點,女友帶著 GPU 來我家學習 Kubernetes:打造 K8s AI Infra 的 30 夜系列 第 23 篇

【Day 23】Envoy AI Gateway:擋在 LLM 前的守衛,進來先檢查 token 額度

  • 分享至 

  • xImage
  •  

前幾天我們看了 Pod 怎麼擴縮、節點從哪裡來、prefill 和 decode 怎麼分到兩台機器上。這三件事的重心都放在負載來了之後。

今天講的是 AI gateway,請求進來之前的那一層:誰進得來、拿到哪個模型、能用多少 token。這一層不會讓機器變多,超過額度的請求會在這裡被擋掉。

今天用的是 Envoy AI Gateway,它在 2026 年 9 月改名叫 Agent Router,下面用新名字。


Kubernetes 上的 gateway

Gateway API 是規格,不是實作

叢集外面的流量要進來,得先經過一個入口。Kubernetes 上這件事原本由 Ingress 負責,而 Gateway API 是它的後繼者。

它定義你能寫哪些物件:GatewayClass 說這一類 gateway 交給哪個 controller 處理、Gateway 定義一個入口(開哪個 port、收什麼協定)、HTTPRoute 說哪個請求送到哪個後端。

真的處理流量的是實作,不是這套 API 本身。Envoy Gateway、kgateway、GKE Gateway 都是同一套 API 的不同實作。

Envoy Gateway

Envoy 是一個 proxy,請求實際上是從它身上經過的。它的設定不是讀檔案,而是透過一個叫 xDS 的介面從外面餵進來。

Envoy Gateway 則是一個 controller。它讀上面那些物件,翻譯成 xDS 餵給 Envoy。所以真的轉發流量的是 Envoy,Envoy Gateway 只是在旁邊告訴它該怎麼轉。

Gateway API 定義的物件只到路由那一層。而限流和指定後端端點這兩件不在那套標準裡,所以它們屬於 Envoy Gateway 自己的 API group gateway.envoyproxy.io。


AI 流量到來,一般的 gateway 少了哪兩件事

model 寫在 body 裡

Kubernetes Gateway API 的 HTTPRoute,在 matches 底下能比對四件事:path、headers、method、queryParams。再加上 HTTPRoute 自己的 hostnames。五個一次用滿是這樣:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: example }
spec:
  hostnames: ["api.example.com"]
  parentRefs: [{ name: aigw }]
  rules:
    - matches:
        - path: { type: PathPrefix, value: /v1 }
          headers: [{ name: x-tenant-id, type: Exact, value: team-a }]
          method: POST
          queryParams: [{ name: debug, type: Exact, value: "true" }]
      backendRefs: [{ name: some-backend, port: 8000 }]

而 OpenAI 相容的請求裡,要用哪個模型寫在 body 的 model 欄位:

{"model": "model-a", "messages": [{"role": "user", "content": "hi"}]}

那四個欄位沒有一個碰得到 body。所以照 model 路由這件事,用 Gateway API 原本的比對規則做不到。

請求數不是可用的單位

一般的限流單位是請求數,Envoy Gateway 的 BackendTrafficPolicy 那個欄位就叫 limit.requests。

但一筆請求的成本差很多。使用者打一句「hi」,輸入是幾個 token;貼一份長文件進來,輸入可能是幾千個。這兩種請求占用的 GPU 時間差很遠,而限流算起來都是一筆。

所以 Agent Router 把額度的單位換成 token。而 token 數要等回應產生完才知道,這件事後面會看到它的後果。

這兩件事有個共同點:解 body 取出 model、讀回應裡的 usage 記下 token 用量,都發生在請求路徑上,而且都要看得到請求和回應的內容。


Agent Router

Agent Router 在一堆模型前面擺一個 OpenAI 相容的入口:照請求裡的 model 決定送去哪個後端、幫你把憑證補上、按 token 計量和限流。

它原本叫 Envoy AI Gateway,2026 年 9 月改名並加入 Agentic AI Foundation。

Agent Router 蓋在 Envoy Gateway 上

Agent Router 沒有取代 Envoy Gateway,它接在旁邊,而且接在三個地方。

第一個是 controller。Agent Router 自己也是一個 controller,管那幾個 AI 專用的資源。而 Envoy Gateway 把物件翻成 xDS 之後,會把結果交給 Agent Router 再加工一次。這個交接點是 Envoy Gateway 提供的擴充機制,叫 extension server。

第二個是請求路徑。路徑上多了一個元件,負責從 body 取出 model,以及從回應裡讀 token 用量。

第三個是限流。扣額度的是 Rate Limit Service,那是 Envoy Gateway 本來就有的服務,而計數要存在外面,所以還需要一個 Redis。


實驗環境

今天在沒有 GPU 的 EC2 用 kind 開一個單節點叢集,裝 Envoy Gateway 加 Agent Router 加一個 Redis,後端是自己寫的假後端。

demo 看什麼
一 怎麼接在 Envoy Gateway 上 兩個 controller、一個原生 sidecar、AIGatewayRoute 生出一份 HTTPRoute
二 一個入口路由到兩個模型 同一個網址,只差 body 裡的 model
三 憑證綁在後端不綁在路由 同一份路由,兩個後端的 Authorization 不同
四 限流扣的是 token 不是請求數 同樣的額度,一筆大的和多筆小的,被擋的發數不同

kind

# kind.yaml
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: aigw-demo
nodes:
  - role: control-plane
kind create cluster --config kind.yaml
Set kubectl context to "kind-aigw-demo"

假後端

一個假後端要同時滿足三件事:

能力 給哪個 demo
兩個模型名稱,兩個 Service demo 二的路由
回傳它收到的 authorization demo 三驗憑證真的被注入
token 數可由請求指定,預設值三個互不相等 demo 四控制第幾發被擋、分辨扣的是哪個計數器

最後一條的預設值故意設成 input 1、output 100、total 300,三個不一樣。換 cost 型別的時候才分得出扣到哪一個。

# 01-fake-backend.yaml
apiVersion: v1
kind: ConfigMap
metadata: { name: fake-src, namespace: default }
data:
  server.py: |
    import json, os, time
    from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

    NAME = os.environ.get("BACKEND_NAME", "unknown")
    DEFAULTS = {"input": 1, "output": 100, "total": 300}

    def pick(hdrs, body, key):
        h = hdrs.get("x-fake-" + key)
        if h is not None:
            return int(h)
        v = body.get("fake_" + key)
        if v is not None:
            return int(v)
        return DEFAULTS[key]

    class H(BaseHTTPRequestHandler):
        protocol_version = "HTTP/1.1"

        def _json(self, code, obj):
            raw = json.dumps(obj).encode()
            self.send_response(code)
            self.send_header("Content-Type", "application/json")
            self.send_header("Content-Length", str(len(raw)))
            self.end_headers()
            self.wfile.write(raw)

        def do_GET(self):
            self._json(200, {"backend": NAME, "path": self.path})

        def do_POST(self):
            n = int(self.headers.get("Content-Length") or 0)
            raw = self.rfile.read(n) if n else b"{}"
            try:
                body = json.loads(raw or b"{}")
            except Exception:
                body = {}
            pin = pick(self.headers, body, "input")
            pout = pick(self.headers, body, "output")
            ptot = pick(self.headers, body, "total")
            auth = self.headers.get("authorization", "<none>")
            model = body.get("model", "<none>")
            self._json(200, {
                "id": "chatcmpl-fake-" + NAME,
                "object": "chat.completion",
                "created": int(time.time()),
                "model": model,
                "choices": [{
                    "index": 0,
                    "finish_reason": "stop",
                    "message": {
                        "role": "assistant",
                        "content": "served_by=%s model=%s authorization=%s" % (NAME, model, auth),
                    },
                }],
                "usage": {
                    "prompt_tokens": pin,
                    "completion_tokens": pout,
                    "total_tokens": ptot,
                },
            })

    ThreadingHTTPServer(("0.0.0.0", 8000), H).serve_forever()
---
apiVersion: apps/v1
kind: Deployment
metadata: { name: fake-a, namespace: default }
spec:
  replicas: 1
  selector: { matchLabels: { app: fake-a } }
  template:
    metadata: { labels: { app: fake-a } }
    spec:
      containers:
        - name: s
          image: python:3.12-slim
          command: ["python3", "/src/server.py"]
          env: [{ name: BACKEND_NAME, value: fake-a }]
          ports: [{ containerPort: 8000, name: http }]
          volumeMounts: [{ name: src, mountPath: /src }]
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            periodSeconds: 5
      volumes:
        - name: src
          configMap: { name: fake-src }
---
apiVersion: v1
kind: Service
metadata: { name: fake-a, namespace: default }
spec:
  selector: { app: fake-a }
  ports: [{ port: 8000, targetPort: 8000 }]
---
# fake-b 跟上面完全一樣,只有名稱、selector、labels、BACKEND_NAME 把 a 換成 b

token 數有兩種指定方式:x-fake-input 這類標頭,或者 body 裡的 fake_input 欄位,標頭優先。

寫兩種是因為請求的 body 會被路徑上那個元件解開再組回去,自己加的欄位不一定留得住,而標頭不會被動到。

kubectl apply -f 01-fake-backend.yaml
kubectl rollout status deploy/fake-a --timeout=180s
kubectl rollout status deploy/fake-b --timeout=180s
deployment "fake-a" successfully rolled out
deployment "fake-b" successfully rolled out

先不經 gateway 直接打一次,確認假後端自己是對的:

kubectl port-forward svc/fake-a 18000:8000 >/tmp/pf-a.log 2>&1 &
echo $! > /tmp/pf-a.pid
sleep 3
curl -s -X POST localhost:18000/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"model-a","messages":[{"role":"user","content":"hi"}]}'
{"id":"chatcmpl-fake-fake-a","object":"chat.completion","model":"model-a",
 "choices":[{"index":0,"finish_reason":"stop",
   "message":{"role":"assistant",
     "content":"served_by=fake-a model=model-a authorization=<none>"}}],
 "usage":{"prompt_tokens":1,"completion_tokens":100,"total_tokens":300}}

再確認標頭的優先序和 authorization 回得出來:

curl -s -X POST localhost:18000/v1/chat/completions -H 'Content-Type: application/json' \
  -H 'x-fake-input: 150' -H 'Authorization: Bearer secret-123' \
  -d '{"model":"model-a","fake_input":9999}' \
| python3 -c 'import sys,json;d=json.load(sys.stdin);print(d["usage"]);print(d["choices"][0]["message"]["content"])'
{'prompt_tokens': 150, 'completion_tokens': 100, 'total_tokens': 300}
served_by=fake-a model=model-a authorization=Bearer secret-123

標頭的 150 蓋過 body 的 9999,收到的 authorization 也回傳得出來。

kill $(cat /tmp/pf-a.pid)

裝 Agent Router

B=https://raw.githubusercontent.com/envoyproxy/ai-gateway/v1.1.0
curl -sLo envoy-gateway-values.yaml           $B/manifests/envoy-gateway-values.yaml
curl -sLo envoy-gateway-values-ratelimit.yaml $B/examples/token_ratelimit/envoy-gateway-values-addon.yaml
kubectl apply -f $B/examples/token_ratelimit/redis.yaml
namespace/redis-system created
service/redis created
deployment.apps/redis created
helm upgrade -i aieg-crd oci://docker.io/envoyproxy/ai-gateway-crds-helm \
  --version v1.1.0 -n envoy-ai-gateway-system --create-namespace
NAME: aieg-crd   STATUS: deployed   REVISION: 1
helm upgrade -i eg oci://docker.io/envoyproxy/gateway-helm \
  --version v1.9.1 -n envoy-gateway-system --create-namespace \
  -f envoy-gateway-values.yaml -f envoy-gateway-values-ratelimit.yaml
NAME: eg   STATUS: deployed   REVISION: 1
kubectl -n envoy-gateway-system rollout status deploy/envoy-gateway --timeout=300s
deployment "envoy-gateway" successfully rolled out
helm upgrade -i aieg oci://docker.io/envoyproxy/ai-gateway-helm \
  --version v1.1.0 -n envoy-ai-gateway-system
NAME: aieg   STATUS: deployed   REVISION: 1

values.yaml 改了什麼

cat envoy-gateway-values.yaml
config:
  envoyGateway:
    extensionApis:
      enableEnvoyPatchPolicy: true
      enableBackend: true
    extensionManager:
      hooks:
        xdsTranslator:
          translation:
            listener: { includeAll: true }
            route:    { includeAll: true }
            cluster:  { includeAll: true }
            secret:   { includeAll: true }
          post: [Translation, Cluster, Route]
      service:
        fqdn:
          hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
          port: 1063

enableBackend: true 打開的是 Envoy Gateway 的 Backend API,原始檔案裡標著 Required。這決定了後面寫路由的時候,backendRef 要指向一個 Backend 物件,不能直接指 Service。

extensionManager 那一段就是接點:service.fqdn 指向 ai-gateway-controller,而 listener、route、cluster、secret 四種都設 includeAll: true,表示 Envoy Gateway 要把整份 xDS 都交出去給它加工。

限流那份 addon 還動了一個地方:

kubectl -n envoy-gateway-system get deploy envoy-ratelimit \
  -o jsonpath='{.spec.template.spec.containers[*].image}'
docker.io/envoyproxy/ratelimit:60d8e81b

Envoy Gateway 自己的預設是 ratelimit:8fe6ea42,被 addon 用 StrategicMerge 蓋掉了。所以 token 限流不只是設定不同,連限流器的映像都要換成另一個版本。

Gateway 與路由

# 02-gateway-route.yaml
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata: { name: aigw }
spec:
  controllerName: gateway.envoyproxy.io/gatewayclass-controller
---
# kind 上要清掉 Envoy 預設的 cpu 和 memory requests,否則排不進去
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata: { name: aigw, namespace: default }
spec:
  provider:
    type: Kubernetes
    kubernetes:
      envoyDeployment:
        container:
          resources: {}
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: aigw, namespace: default }
spec:
  gatewayClassName: aigw
  listeners:
    - { name: http, protocol: HTTP, port: 80 }
  infrastructure:
    parametersRef:
      group: gateway.envoyproxy.io
      kind: EnvoyProxy
      name: aigw
---
# backendRef 不能直接指 Service,必須是 Envoy Gateway 的 Backend
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata: { name: be-a, namespace: default }
spec:
  endpoints:
    - fqdn: { hostname: fake-a.default.svc.cluster.local, port: 8000 }
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata: { name: be-b, namespace: default }
spec:
  endpoints:
    - fqdn: { hostname: fake-b.default.svc.cluster.local, port: 8000 }
---
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata: { name: svc-a, namespace: default }
spec:
  schema: { name: OpenAI }
  backendRef: { name: be-a, kind: Backend, group: gateway.envoyproxy.io }
---
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata: { name: svc-b, namespace: default }
spec:
  schema: { name: OpenAI }
  backendRef: { name: be-b, kind: Backend, group: gateway.envoyproxy.io }
---
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIGatewayRoute
metadata: { name: demo, namespace: default }
spec:
  parentRefs:
    - { name: aigw, kind: Gateway, group: gateway.networking.k8s.io }
  rules:
    # 比對的是 x-ai-eg-model 這個標頭,它是從 body 的 model 合成出來的
    - matches: [{ headers: [{ type: Exact, name: x-ai-eg-model, value: model-a }] }]
      backendRefs: [{ name: svc-a }]
    - matches: [{ headers: [{ type: Exact, name: x-ai-eg-model, value: model-b }] }]
      backendRefs: [{ name: svc-b }]
    # 第三條只為了確認帶斜線的真實模型名 Exact 比對過不過
    - matches: [{ headers: [{ type: Exact, name: x-ai-eg-model, value: "Qwen/Qwen2.5-7B-Instruct" }] }]
      backendRefs: [{ name: svc-a }]
  llmRequestCosts:
    - { metadataKey: llm_input_token,  type: InputToken }
    - { metadataKey: llm_output_token, type: OutputToken }
    - { metadataKey: llm_total_token,  type: TotalToken }
kubectl apply -f 02-gateway-route.yaml
gatewayclass.gateway.networking.k8s.io/aigw created
envoyproxy.gateway.envoyproxy.io/aigw created
gateway.gateway.networking.k8s.io/aigw created
backend.gateway.envoyproxy.io/be-a created
backend.gateway.envoyproxy.io/be-b created
aiservicebackend.aigateway.envoyproxy.io/svc-a created
aiservicebackend.aigateway.envoyproxy.io/svc-b created
aigatewayroute.aigateway.envoyproxy.io/demo created

llmRequestCosts 的作用是宣告要把哪些數字記下來,不是設額度。額度寫在另一份 BackendTrafficPolicy 裡,而那份只對 input 和 output 兩個設了上限。所以 llm_total_token 有記,但這天沒有拿它來擋請求。


一、怎麼接在 Envoy Gateway 上

兩個 controller

kubectl get pods -A
envoy-ai-gateway-system   ai-gateway-controller-59899b6864-mcxm5   1/1 Running
envoy-gateway-system      envoy-gateway-6fbfccc98d-cp9hd           1/1 Running
envoy-gateway-system      envoy-ratelimit-857f488cdf-n7hkz         1/1 Running
redis-system              redis-6f6dbbd547-4vj4j                   1/1 Running
default                   fake-a-f7c68f445-ql5md                   1/1 Running
default                   fake-b-5857b7d5c5-9b7rw                  1/1 Running

兩個 controller 在兩個 namespace,加一個 Rate Limit Service 和它的 Redis。

Envoy Gateway 本身沒有任何 AI 的概念,它只知道要把 xDS 交給設定裡指定的那個服務。那個機制是通用的,指誰就交給誰,而這次指的是 Agent Router 的 controller。

六個 CRD,這天用三個

kubectl get crd | grep aigateway.envoyproxy.io
aigatewayroutes.aigateway.envoyproxy.io
aiservicebackends.aigateway.envoyproxy.io
backendsecuritypolicies.aigateway.envoyproxy.io
gatewayconfigs.aigateway.envoyproxy.io
mcproutes.aigateway.envoyproxy.io
quotapolicies.aigateway.envoyproxy.io

這天用到前三個。mcproutes 是 MCP 的路由,quotapolicies 是另一套額度機制,都沒碰。

Envoy 只有一份,但那個 Pod 裡有三個容器

kubectl get gateway aigw -o wide
NAME   CLASS   ADDRESS   PROGRAMMED   AGE
aigw   aigw              False        31s

kind 沒有 LoadBalancer,所以停在 PROGRAMMED=False。這不影響測試,port-forward 照樣通。

kubectl get pods -n envoy-gateway-system -l gateway.envoyproxy.io/owning-gateway-name=aigw
envoy-default-aigw-7b45c098-5d965d6f5d-qnh6t   3/3   Running

資料面只有一份,名稱是 envoy-<namespace>-<gateway>-<hash>。Agent Router 沒有自己的資料面。但那個 Pod 顯示 3/3,所以裡面不只 Envoy:

P=$(kubectl -n envoy-gateway-system get pods \
     -l gateway.envoyproxy.io/owning-gateway-name=aigw -o jsonpath='{.items[0].metadata.name}')
kubectl -n envoy-gateway-system get pod $P \
  -o jsonpath='{range .spec.containers[*]}{.name}{"\t"}{.image}{"\n"}{end}'
envoy              docker.io/envoyproxy/envoy:distroless-v1.39.1
shutdown-manager   docker.io/envoyproxy/gateway:v1.9.1
kubectl -n envoy-gateway-system get pod $P \
  -o jsonpath='{range .spec.initContainers[*]}{.name}{" restartPolicy="}{.restartPolicy}{"\n"}{end}'
ai-gateway-extproc restartPolicy=Always

ai-gateway-extproc 是 initContainer 加 restartPolicy: Always,那是 Kubernetes 的原生 sidecar。它跟 Envoy 在同一個 Pod 裡,不是另一個 Service,所以請求不會跨 Pod 多跳一次。

所以「加上去的一層」實際上在三個地方各加了一點:

加在哪 加了什麼
控制面 一個 controller,掛在 Envoy Gateway 的 xDS post hook
資料面 一個原生 sidecar,塞進 Envoy 自己的 Pod
限流 換掉 ratelimit 的映像,外加一個 Redis

AIGatewayRoute 生出一份 HTTPRoute

kubectl get httproute -A
NAMESPACE   NAME   HOSTNAMES   AGE
default     demo               42s
kubectl get httproute demo -o yaml
metadata:
  ownerReferences: [{ kind: AIGatewayRoute, name: demo }]
spec:
  parentRefs: [{ group: gateway.networking.k8s.io, kind: Gateway, name: aigw }]
  rules:
  - backendRefs: [{ group: gateway.envoyproxy.io, kind: Backend, name: be-a, weight: 1 }]
    filters:
    - type: ExtensionRef
      extensionRef: { group: gateway.envoyproxy.io, kind: HTTPRouteFilter,
                      name: ai-eg-host-rewrite-demo }
    matches:
    - headers: [{ name: x-ai-eg-model, type: Exact, value: model-a }]
      path: { type: PathPrefix, value: / }
    timeouts: { request: 60s }
  # model-b 和 Qwen/Qwen2.5-7B-Instruct 兩條同形狀
  - name: route-not-found
    filters:
    - type: ExtensionRef
      extensionRef: { group: gateway.envoyproxy.io, kind: HTTPRouteFilter,
                      name: ai-eg-route-not-found-response-demo }
    matches: [{ path: { type: PathPrefix, value: / } }]
kubectl get httproutefilter
ai-eg-host-rewrite-demo               owner: AIGatewayRoute
ai-eg-route-not-found-response-demo   owner: AIGatewayRoute

AIGatewayRoute 不是新的路由機制,它是 HTTPRoute 的產生器。從頭到尾的順序是這樣:

AIGatewayRoute demo
  → HTTPRoute demo(同名,ownerReferences 指回去)
      每條 rule 補上 path PathPrefix /
      每條 rule 補上 timeouts.request 60s
      每條 rule 掛上 HTTPRouteFilter ai-eg-host-rewrite-demo
      最後自動補一條什麼都接的規則:route-not-found
  → 兩個 HTTPRouteFilter,也由 AIGatewayRoute 擁有
  → Envoy Gateway 翻成 xDS
  → 經 post hook 交給 ai-gateway-controller 加工
  → extproc sidecar 在請求路徑上執行

而它生出來的 HTTPRoute 比對的是 x-ai-eg-model 標頭,不是 body。body 到標頭那一步是 extproc 做的。


二、一個入口路由到兩個模型

SVC=$(kubectl -n envoy-gateway-system get svc -o name | grep envoy-default-aigw)
kubectl -n envoy-gateway-system port-forward "$SVC" 8080:80 >/tmp/pf-gw.log 2>&1 &
echo $! > /tmp/pf-gw.pid
sleep 6
for m in model-a model-b Qwen/Qwen2.5-7B-Instruct nope; do
  printf '%-28s ' "$m"
  curl -s -o /tmp/r.json -w '%{http_code} ' -X POST localhost:8080/v1/chat/completions \
    -H 'Content-Type: application/json' \
    -d "{\"model\":\"$m\",\"messages\":[{\"role\":\"user\",\"content\":\"hi\"}]}"
  python3 -c 'import json;print(json.load(open("/tmp/r.json"))["choices"][0]["message"]["content"])' \
    2>/dev/null || head -c 120 /tmp/r.json
  echo
done
model-a                   200 served_by=fake-a model=model-a authorization=<none>
model-b                   200 served_by=fake-b model=model-b authorization=<none>
Qwen/Qwen2.5-7B-Instruct  200 served_by=fake-a model=Qwen/Qwen2.5-7B-Instruct authorization=<none>
nope                      404 No matching route found. It is likely because the model
                              specified in your request is not configured in the Gateway.

同一個網址、同一種格式、同一個方法,只差 body 裡的 model 一個字。帶斜線的真實模型名 Exact 也比對得到。

那個 404 不是 Envoy 的通用錯誤,也不是假後端回的。前一節看 HTTPRoute 的時候,最後有一條叫 route-not-found 的規則,它的 path 是 /,所以前面三條比對 model 的規則都沒中時,請求就落到它身上,由它掛的 filter 回這段訊息。那條規則是 AIGatewayRoute 自動補的。

Envoy 的 access log 也看得到同一件事:那筆 404 的 route_name 結尾是 rule/3,而 AIGatewayRoute 只寫了三條規則,所以 rule 3 就是多出來的那條。它的 response_code_details 是 direct_response,upstream_host 是 null,表示訊息由 filter 直接回,沒有往任何後端轉。

所以整條路徑是這樣接起來的:

客戶端    把 model 寫在 body 裡
extproc   讀出來,塞成 x-ai-eg-model 標頭
Envoy     用它原本就會的標頭比對決定去哪

解法不是教 Envoy 讀 body,是在前面加一步翻譯。


三、憑證綁在後端,不綁在路由

BackendSecurityPolicy 掛在哪裡

kubectl get crd backendsecuritypolicies.aigateway.envoyproxy.io \
  -o jsonpath='{.spec.versions[?(@.name=="v1beta1")].schema.openAPIV3Schema.properties.spec.properties.type.enum}'
["APIKey","AWSCredentials","AzureAPIKey","AzureCredentials","GCPCredentials","AnthropicAPIKey"]
kubectl explain backendsecuritypolicy.spec.targetRefs
targetRefs is the names of the AIServiceBackend or InferencePool resources
this BackendSecurityPolicy is being attached to. Attaching multiple
BackendSecurityPolicies to the same resource is invalid.

掛的對象是 AIServiceBackend,不是 AIGatewayRoute。另外 apiKey.secretRef 那個 Secret 的鍵名必須是 apiKey。

只給 svc-a 一把 key

# 04-backend-security.yaml
apiVersion: v1
kind: Secret
metadata: { name: key-a, namespace: default }
type: Opaque
stringData:
  apiKey: secret-for-fake-a
---
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: BackendSecurityPolicy
metadata: { name: auth-a, namespace: default }
spec:
  type: APIKey
  apiKey:
    secretRef: { name: key-a }
  targetRefs:
    - { group: aigateway.envoyproxy.io, kind: AIServiceBackend, name: svc-a }
kubectl apply -f 04-backend-security.yaml
secret/key-a created
backendsecuritypolicy.aigateway.envoyproxy.io/auth-a created
kubectl get backendsecuritypolicy auth-a \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}{"\n"}{end}'
Accepted=True ReconciliationSucceeded

結果

for m in model-a model-b Qwen/Qwen2.5-7B-Instruct; do
  printf '%-26s ' "$m"
  curl -s -X POST localhost:8080/v1/chat/completions \
    -H 'Content-Type: application/json' \
    -d "{\"model\":\"$m\",\"messages\":[{\"role\":\"user\",\"content\":\"hi\"}]}" \
  | python3 -c 'import sys,json;print(json.load(sys.stdin)["choices"][0]["message"]["content"])'
done
model-a                    served_by=fake-a model=model-a authorization=Bearer secret-for-fake-a
model-b                    served_by=fake-b model=model-b authorization=<none>
Qwen/Qwen2.5-7B-Instruct   served_by=fake-a model=Qwen/Qwen2.5-7B-Instruct authorization=Bearer secret-for-fake-a

svc-a 收到 key,svc-b 什麼都沒有。同一份 AIGatewayRoute,兩個後端的憑證狀態不同。

第三行是真正的證據。Qwen/Qwen2.5-7B-Instruct 是另一條路由規則,但它也指向 svc-a,所以它也拿到了那把 key,而那條規則我一個字都沒有碰。

注入的形式是 Authorization: Bearer <key>。CRD 的說明只講會注入到 Authorization 標頭,沒有提到會補上 Bearer 這個前綴,那是實際跑起來才看到的。


四、限流扣的是 token 不是請求數

額度

# 03-ratelimit.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata: { name: quota, namespace: default }
spec:
  targetRefs:
    - { name: aigw, kind: Gateway, group: gateway.networking.k8s.io }
  rateLimit:
    type: Global
    global:
      rules:
        # 額度一:輸入 token,1000 一小時,每個 x-tenant-id 各自一份
        - clientSelectors: [{ headers: [{ name: x-tenant-id, type: Distinct }] }]
          limit: { requests: 1000, unit: Hour }
          cost:
            request: { from: Number, number: 0 }
            response:
              from: Metadata
              metadata: { namespace: io.envoy.ai_gateway, key: llm_input_token }
        # 額度二:輸出 token,數字刻意設成一樣
        - clientSelectors: [{ headers: [{ name: x-tenant-id, type: Distinct }] }]
          limit: { requests: 1000, unit: Hour }
          cost:
            request: { from: Number, number: 0 }
            response:
              from: Metadata
              metadata: { namespace: io.envoy.ai_gateway, key: llm_output_token }
kubectl apply -f 03-ratelimit.yaml
backendtrafficpolicy.gateway.envoyproxy.io/quota created
kubectl get backendtrafficpolicy quota \
  -o jsonpath='{range .status.ancestors[*]}{range .conditions[*]}{.type}={.status} {.reason}{"\n"}{end}{end}'
Accepted=True Accepted

這份 BackendTrafficPolicy 掛在 aigw 這個 Gateway 上,做的事是給每一個 x-tenant-id 各自一份額度:輸入 token 一小時 1000 個,輸出 token 一小時 1000 個,兩邊分開算。超過的請求回 429 Too Many Requests。

兩個額度的數字刻意設成一樣,這樣第幾發被擋的差別只能來自讀的是哪個計數器。

cost.request.number 設 0 的意思是請求進來的時候只檢查額度,不從額度裡扣,扣多少完全由回應那邊的 token 數決定。

欄位名是 requests,但這裡的 1000 是 token 數。一般的限流一筆請求扣 1,所以 requests: 1000 就是 1000 筆;而 cost.response 叫它改去 io.envoy.ai_gateway 這個 metadata namespace 拿 token 數來扣,於是每一筆扣的是自己用掉的 token。欄位名沒有跟著改。

一筆大的和多筆小的

要證明扣的是 token 不是請求數,做法是讓兩組流量的請求數不一樣、token 總量差不多,然後看各自在第幾發被擋。所以兩個 tenant 共用同一個 1000 的額度,一個每次送 1200 個輸入 token,另一個每次送 150 個。

下面這個函式指定 tenant 和 token 數,然後只印 HTTP 狀態碼:

shot() {   # shot <tenant> <input> <output>
  curl -s -o /dev/null -w '%{http_code}' -X POST localhost:8080/v1/chat/completions \
    -H 'Content-Type: application/json' -H "x-tenant-id: $1" \
    -H "x-fake-input: $2" -H "x-fake-output: $3" \
    -d '{"model":"model-a","messages":[{"role":"user","content":"hi"}]}'
}
T=$(date +%s)

用時間戳當 tenant 名稱的一部分,每次跑都是一個沒用過的 tenant,計數器從零開始,不用等一小時。

for i in 1 2 3; do printf "第 %d 發  %s\n" $i "$(shot a-$T 1200 1)"; sleep 1; done
第 1 發  200
第 2 發  429
第 3 發  429
for i in $(seq 1 9); do printf "第 %d 發  %s\n" $i "$(shot b-$T 150 1)"; sleep 1; done
第 1 發  200
第 2 發  200
第 3 發  200
第 4 發  200
第 5 發  200
第 6 發  200
第 7 發  200
第 8 發  429
第 9 發  429

同一個額度 1000,一筆 1200 的在第 2 發就被擋,七筆 150 的撐到第 8 發。如果限的是請求數,兩組被擋的發數會一樣。它們不一樣,所以限的是 token。

第 7 發值得單獨看。那一發送出去之前,累計已經是 6 乘 150 等於 900,還沒超過 1000,所以它過了;過完之後累計變成 1050,已經超額。第 8 發才被擋。

原因是額度是在回應之後才扣的。系統知道你超額的時候,你已經超了。這不是實作瑕疵,是按 token 算額度的必然,因為用了多少要等生成結束才知道。所以額度要留餘裕,剛好設成上限的話一定會超過一發。

從 gateway 這一側看那個 429

上面那些 429 都是客戶端看到的狀態碼。Envoy 的 access log 可以看到它自己怎麼記這件事:

P=$(kubectl -n envoy-gateway-system get pods \
     -l gateway.envoyproxy.io/owning-gateway-name=aigw -o jsonpath='{.items[0].metadata.name}')
kubectl -n envoy-gateway-system logs $P -c envoy \
| python3 -c "
import json, sys
for line in sys.stdin:
    if not line.strip().startswith('{'):
        continue
    d = json.loads(line)
    if d.get('response_code') == 429:
        print(json.dumps(d, indent=2, sort_keys=True))
        break
"
{
  ":authority": "localhost:8080",
  "bytes_received": 63,
  "bytes_sent": 0,
  "duration": 4,
  "method": "POST",
  "response_code": 429,
  "response_code_details": "request_rate_limited",
  "response_flags": "RL",
  "route_name": "httproute/default/demo/rule/0/match/0/*",
  "upstream_cluster": "httproute/default/demo/rule/0",
  "upstream_host": null,
  "x-envoy-origin-path": "/v1/chat/completions",
  "x-request-id": "de3d62d9-0be9-44ba-b515-8768d65c4c52"
}

response_flags 是 RL、response_code_details 是 request_rate_limited,所以這個 429 是 Envoy 自己因為限流發出來的。不是後端回的,也不是連線失敗。

而 upstream_host 是 null、bytes_sent 是 0,表示這筆請求沒有送到任何後端。被擋下來的請求不會花到後端的運算。

換計數器,同一種流量結果完全不同

for i in 1 2; do printf "第 %d 發  %s\n" $i "$(shot c-$T 1 1200)"; sleep 1; done
第 1 發  200
第 2 發  429

輸入只用掉 1 個 token,輸出額度卻爆了。如果只對輸入設上限,這種流量會整批溜過去。

Redis 裡的計數器

額度的計數存在 Redis 裡,所以可以直接把它讀出來,跟上面送出去的 token 數核對:

kubectl -n redis-system exec deploy/redis -- redis-cli --scan > /tmp/keys.txt
sort /tmp/keys.txt | while IFS= read -r k; do
  v=$(kubectl -n redis-system exec deploy/redis -- redis-cli get "$k")
  printf "%-40s %s\n" "$(echo "$k" | sed 's|.*/\*_||')" "$v"
done
rule-0-match-0_a-...1791342000   1200
rule-0-match-0_b-...1791342000   1050
rule-0-match-0_c-...1791342000      1
rule-1-match-0_a-...1791342000      1
rule-1-match-0_b-...1791342000      7
rule-1-match-0_c-...1791342000   1200

key 的組成是路由位置、第幾條 cost 規則、tenant、這是哪一個小時。六個 key 等於三個 tenant 乘兩條 cost 規則。

tenant 送了什麼 InputToken 計數器 OutputToken 計數器
a 1 筆成功,input 1200、output 1 1200 1
b 7 筆成功,各 input 150、output 1 1050 7
c 1 筆成功,input 1、output 1200 1 1200

輸入和輸出是兩個獨立的計數器。a 的輸入用掉 1200 已經超過,而輸出只用了 1;c 剛好相反。

key 裡帶著 tenant 名稱,那是 type: Distinct 的效果,所以額度是每個 tenant 各一份。

被擋的請求不算進額度。a 一共打了三發,計數器是 1200 而不是 3600,只有成功回應的那一發被扣,這就是 cost.request.number: 0 的效果。

kill $(cat /tmp/pf-gw.pid)

其他方案

以下幾個今天都沒有測,只是記錄它們的定位。

名稱 定位
LiteLLM 應用層的 proxy,Python 寫的,不走 Gateway API,設定是一份 config.yaml 加一個 Postgres。有官方 Helm chart,放進 Kubernetes 很正常。供應商覆蓋 100 多家,是這幾個裡面使用者最多的
agentgateway Rust 寫的資料面,不是建在 Envoy 上,但有 Gateway API。主打 MCP 和 agent 之間的協定
kgateway Envoy 上成熟的 Gateway API 實作,但 v2.2.0 把 AI 功能全部移出去了,現在 AI 的部分要看 agentgateway
Gloo Gateway Solo.io 的商業版。它的開源核心捐給 CNCF 之後改名 kgateway

這幾個擺在一起之後,選擇其實是兩個決定疊在一起。

第一個決定是 AI 流量要不要走 Gateway API 這條路。走的話設定是叢集資源,kubectl 看得到、RBAC 管得動、GitOps 的 diff 看得懂。不走的話設定是一份檔案加一個資料庫,而那個 proxy 自己是你要擴縮的應用。

第二個決定才是在那條路上選誰,而這一題的差別在資料面:Agent Router 讓 AI 流量跟其他 HTTP 流量共用同一個 Envoy,agentgateway 給 AI 流量一條專門寫的 Rust 資料面。


小結

AI gateway 不是另一種 gateway。它是一般的 gateway 加兩件事:路由的依據從 header 和 path 換成 request body 裡的 model,額度的單位從請求數換成 token。

這兩件都需要一個站在請求路徑上、而且看得到請求和回應內容的元件。但底層沒有被換掉,資料面還是 Envoy,路由最後還是落在 HTTPRoute 上。所以 AI 流量跟叢集裡其他 HTTP 流量共用同一套 TLS、同一套 RBAC、同一套 GitOps。

用 token 當額度的單位有一個躲不掉的後果:用了多少要等回應產生完才知道,所以系統發現超額的時候,額度已經被超過了。這不是實作上的瑕疵,是這個單位本身帶來的,所以額度要留餘裕,不能剛好設成上限。

而這一層要在有多個使用者、多個模型、或者有預算要管的時候才用得上。自己跑一個模型自己用,前面幾天的東西就夠了。


參考資料

Agent Router joins the Agentic AI Foundation
Envoy AI Gateway v1.0 release announcement
System Architecture Overview, Envoy AI Gateway
Usage-based rate limiting, Envoy AI Gateway
Connecting to AI Providers, Envoy AI Gateway
kgateway v2.2.x release notes
LiteLLM: Docker, Deployment
HTTPRoute, Gateway API
Gateway API 專案說明


上一篇
【Day 22】Kubernetes 上的 PD 分離與 KV 傳輸
下一篇
【Day 24】Gateway API Inference Extension:突破 Kubernetes 傳統路由限制的推理負載平衡
系列文
凌晨四點,女友帶著 GPU 來我家學習 Kubernetes:打造 K8s AI Infra 的 30 夜 共 24 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言