L1 規則引擎(Day 15–16)、微調後的 L2(Day 20)都有了。今天做三件事:用一個框架把它們串成 pipeline、掛進 LiteLLM 閘道、讓整套東西在沒有網路的環境跑起來。
自己寫 check_l1(); check_l2() 兩行很簡單,但真實部署會遇到:scanner 要能獨立開關、每個 scanner 要有自己的 timeout、判定結果要有統一結構、輸出端與輸入端的 scanner 要分開管理。這些東西 LLM Guard(Protect AI 開源)已經做好了,它的設計正好是「一串 input scanners + 一串 output scanners」。
LLM Guard
├── input_scanners: [L1RuleScanner, ShieldGemmaScanner, TokenLimit]
└── output_scanners: [L1RuleScanner(out), ShieldGemmaScanner(out), URLAllowlist]
LLM Guard 內建的 PromptInjection scanner 用的是它自己的模型;本系列不用內建的,而是把 Day 20 的微調版包成自訂 scanner——這樣 Day 28 的對照才是「我們的 L2」而不是「別人的 L2」。
# guards/scanners.py
from llm_guard.input_scanners.base import Scanner
class L1RuleScanner(Scanner):
def __init__(self, engine, direction):
self.engine, self.direction = engine, direction
def scan(self, prompt: str, context=None) -> tuple[str, bool, float]:
v = self.engine.check(prompt, self.direction, context or {})
if v.action == "block": return prompt, False, 1.0
if v.action == "mask": return v.masked_text, True, 0.5
if v.action == "allow": return prompt, True, 0.0
return prompt, True, 0.3 # escalate / tag:放行但標記,交給 L2
class ShieldGemmaScanner(Scanner):
def __init__(self, client, t_low, t_high):
self.client, self.t_low, self.t_high = client, t_low, t_high
def scan(self, prompt, context=None):
p = self.client.score(prompt, POLICY_INJECTION)
if p >= self.t_high: return prompt, False, p # block
if p >= self.t_low: return prompt, True, p # 灰色地帶,標記給 L3(Day 22)
return prompt, True, p
框架預設會跑完所有 scanner。L1 判 allow 的流量不該再進 L2——這是 Day 15 量測的「多少流量在 L1 結束」能否兌現成延遲節省的關鍵。做法是在 L1 scanner 判 allow 時設一個旗標,L2 scanner 看到旗標直接回傳:
class ShieldGemmaScanner(Scanner):
def scan(self, prompt, context=None):
if context and context.get("l1_allow"):
return prompt, True, 0.0
...
LiteLLM 有 async_pre_call_hook 與 async_post_call_success_hook,正好是 Day 3 的兩個攔截點。
# gateway/guard_hook.py
from litellm.integrations.custom_logger import CustomLogger
class GuardHook(CustomLogger):
async def async_pre_call_hook(self, user_api_key_dict, cache, data, call_type):
text = extract_user_content(data["messages"]) # Day 9:最後三輪
for chunk in data.get("rag_chunks", []): # RAG 段落單獨掃
_, ok, _ = input_pipeline.scan(chunk)
if not ok: raise GuardBlocked("rag_chunk")
sanitized, ok, scores = input_pipeline.scan(text)
audit(direction="in", scores=scores, mode=current_mode())
if not ok: raise GuardBlocked("input")
return data
async def async_post_call_success_hook(self, data, user_api_key_dict, response):
text = response.choices[0].message.content
sanitized, ok, scores = output_pipeline.scan(text)
audit(direction="out", scores=scores, mode=current_mode())
if not ok: raise GuardBlocked("output")
response.choices[0].message.content = sanitized # 遮罩後的版本
return response
# litellm_config.yaml
litellm_settings:
callbacks: ["gateway.guard_hook.GuardHook"]
串流輸出用 Day 12 決定的「掃過才放行」:post-call hook 在串流模式下要改用分段緩衝,實作放在 repo,量 TTFT 影響:
【此處貼串流模式 TTFT 對照截圖】
Day 4 的階梯,在這裡落地:
# gateway/health.py
class GuardCircuit:
def __init__(self, fail_threshold=3, cool_down=30):
self.failures, self.open_until = 0, 0
def call(self, fn, *a):
if time.time() < self.open_until:
raise CircuitOpen
try:
r = fn(*a); self.failures = 0; return r
except Exception:
self.failures += 1
if self.failures >= self.fail_threshold:
self.open_until = time.time() + self.cool_down
raise
L2 的 circuit open 時,pipeline 進入 degraded 模式:只跑 L1、輸入端依 Day 4 設定(本系列:fail-closed 降級到 L1)、所有稽核記錄打上 mode=degraded。恢復前跑 smoke test(eval 集抽 20 筆)。
金融客戶的 air-gap 環境沒有網路。要確認的事:
HF_HUB_OFFLINE=1,權重從 /models 讀,不能有任何 from_pretrained("google/...") 觸發下載。docker save / docker load。/config 讀,Day 29 的離線更新包只換這個目錄。# 驗證離線:拔網路後啟動
docker network disconnect bridge guard-l2
docker compose up -d
curl localhost:4000/health # 閘道 + 兩層護欄都應為 healthy
使用者 ─▶ LiteLLM ─▶ [L1 YAML 規則] ─▶ [ShieldGemma 微調版] ─▶ 主模型
↓ 短路 allow ↓ 灰色地帶 → L3(Day 22)
稽核日誌 ◀────────────────┘
| Day | 做了什麼 |
|---|---|
| 15 | L1 規則引擎,YAML 規則,短路 |
| 16 | 台灣 PII checksum,合成資料 |
| 17 | 選型:授權、產地、政策三關後才看 benchmark |
| 18 | ARM 部署,延遲曲線 |
| 19 | 訓練資料,eval 隔離 |
| 20 | LoRA 微調,三方對照 |
| 21 | LLM Guard 組裝,LiteLLM hook,熔斷降級,離線 |
Week 4 開始。Day 22:L3 LLM-as-judge——灰色地帶送給誰判、judge 本身被注入怎麼辦、輸出為什麼要限制成結構化格式,以及雲端(Gemini)與地端(本機大模型)兩種 judge 的成本。
Instagram @aid3fend。
更多 AI 資安筆記:aid3fend.com