iT邦幫忙

2026 iThome 鐵人賽

DAY 21
0
AI Security

《30 天打造 AI Guardrails》系列 第 21 篇

Day 21|LLM Guard 框架層:離線自架的組裝

  • 分享至 

  • xImage
  •  

Week 3 收尾:把零件裝進閘道

L1 規則引擎(Day 15–16)、微調後的 L2(Day 20)都有了。今天做三件事:用一個框架把它們串成 pipeline、掛進 LiteLLM 閘道、讓整套東西在沒有網路的環境跑起來。

為什麼用框架而不是自己寫 pipeline

自己寫 check_l1(); check_l2() 兩行很簡單,但真實部署會遇到:scanner 要能獨立開關、每個 scanner 要有自己的 timeout、判定結果要有統一結構、輸出端與輸入端的 scanner 要分開管理。這些東西 LLM Guard(Protect AI 開源)已經做好了,它的設計正好是「一串 input scanners + 一串 output scanners」。

LLM Guard
├── input_scanners:  [L1RuleScanner, ShieldGemmaScanner, TokenLimit]
└── output_scanners: [L1RuleScanner(out), ShieldGemmaScanner(out), URLAllowlist]

LLM Guard 內建的 PromptInjection scanner 用的是它自己的模型;本系列不用內建的,而是把 Day 20 的微調版包成自訂 scanner——這樣 Day 28 的對照才是「我們的 L2」而不是「別人的 L2」。

自訂 scanner

# guards/scanners.py
from llm_guard.input_scanners.base import Scanner

class L1RuleScanner(Scanner):
    def __init__(self, engine, direction):
        self.engine, self.direction = engine, direction

    def scan(self, prompt: str, context=None) -> tuple[str, bool, float]:
        v = self.engine.check(prompt, self.direction, context or {})
        if v.action == "block":   return prompt, False, 1.0
        if v.action == "mask":    return v.masked_text, True, 0.5
        if v.action == "allow":   return prompt, True, 0.0
        return prompt, True, 0.3          # escalate / tag:放行但標記,交給 L2

class ShieldGemmaScanner(Scanner):
    def __init__(self, client, t_low, t_high):
        self.client, self.t_low, self.t_high = client, t_low, t_high

    def scan(self, prompt, context=None):
        p = self.client.score(prompt, POLICY_INJECTION)
        if p >= self.t_high:  return prompt, False, p        # block
        if p >= self.t_low:   return prompt, True, p         # 灰色地帶,標記給 L3(Day 22)
        return prompt, True, p

L1 的 allow 要真的短路

框架預設會跑完所有 scanner。L1 判 allow 的流量不該再進 L2——這是 Day 15 量測的「多少流量在 L1 結束」能否兌現成延遲節省的關鍵。做法是在 L1 scanner 判 allow 時設一個旗標,L2 scanner 看到旗標直接回傳:

class ShieldGemmaScanner(Scanner):
    def scan(self, prompt, context=None):
        if context and context.get("l1_allow"):
            return prompt, True, 0.0
        ...

掛進 LiteLLM

LiteLLM 有 async_pre_call_hook 與 async_post_call_success_hook,正好是 Day 3 的兩個攔截點。

# gateway/guard_hook.py
from litellm.integrations.custom_logger import CustomLogger

class GuardHook(CustomLogger):
    async def async_pre_call_hook(self, user_api_key_dict, cache, data, call_type):
        text = extract_user_content(data["messages"])        # Day 9:最後三輪
        for chunk in data.get("rag_chunks", []):              # RAG 段落單獨掃
            _, ok, _ = input_pipeline.scan(chunk)
            if not ok: raise GuardBlocked("rag_chunk")
        sanitized, ok, scores = input_pipeline.scan(text)
        audit(direction="in", scores=scores, mode=current_mode())
        if not ok: raise GuardBlocked("input")
        return data

    async def async_post_call_success_hook(self, data, user_api_key_dict, response):
        text = response.choices[0].message.content
        sanitized, ok, scores = output_pipeline.scan(text)
        audit(direction="out", scores=scores, mode=current_mode())
        if not ok: raise GuardBlocked("output")
        response.choices[0].message.content = sanitized     # 遮罩後的版本
        return response
# litellm_config.yaml
litellm_settings:
  callbacks: ["gateway.guard_hook.GuardHook"]

串流輸出用 Day 12 決定的「掃過才放行」:post-call hook 在串流模式下要改用分段緩衝,實作放在 repo,量 TTFT 影響:

【此處貼串流模式 TTFT 對照截圖】

健康檢查、熔斷、降級

Day 4 的階梯,在這裡落地:

# gateway/health.py
class GuardCircuit:
    def __init__(self, fail_threshold=3, cool_down=30):
        self.failures, self.open_until = 0, 0

    def call(self, fn, *a):
        if time.time() < self.open_until:
            raise CircuitOpen
        try:
            r = fn(*a); self.failures = 0; return r
        except Exception:
            self.failures += 1
            if self.failures >= self.fail_threshold:
                self.open_until = time.time() + self.cool_down
            raise

L2 的 circuit open 時,pipeline 進入 degraded 模式:只跑 L1、輸入端依 Day 4 設定(本系列:fail-closed 降級到 L1)、所有稽核記錄打上 mode=degraded。恢復前跑 smoke test(eval 集抽 20 筆)。

完全離線

金融客戶的 air-gap 環境沒有網路。要確認的事:

  1. 模型權重本機載入:HF_HUB_OFFLINE=1,權重從 /models 讀,不能有任何 from_pretrained("google/...") 觸發下載。
  2. LLM Guard 內建 scanner 的模型:如果有用到內建 scanner(例如 TokenLimit 的 tokenizer),它們的模型檔也要預先放好。
  3. 容器映像離線匯入:docker save / docker load。
  4. 規則檔與閾值設定:從 /config 讀,Day 29 的離線更新包只換這個目錄。
# 驗證離線:拔網路後啟動
docker network disconnect bridge guard-l2
docker compose up -d
curl localhost:4000/health   # 閘道 + 兩層護欄都應為 healthy

Week 3 回顧與地端線目前的樣子

使用者 ─▶ LiteLLM ─▶ [L1 YAML 規則] ─▶ [ShieldGemma 微調版] ─▶ 主模型
                         ↓ 短路 allow        ↓ 灰色地帶 → L3(Day 22)
                    稽核日誌 ◀────────────────┘

 

Day 做了什麼
15 L1 規則引擎,YAML 規則,短路
16 台灣 PII checksum,合成資料
17 選型:授權、產地、政策三關後才看 benchmark
18 ARM 部署,延遲曲線
19 訓練資料,eval 隔離
20 LoRA 微調,三方對照
21 LLM Guard 組裝,LiteLLM hook,熔斷降級,離線

明天預告

Week 4 開始。Day 22:L3 LLM-as-judge——灰色地帶送給誰判、judge 本身被注入怎麼辦、輸出為什麼要限制成結構化格式,以及雲端(Gemini)與地端(本機大模型)兩種 judge 的成本。


追蹤 AId3fend

Instagram @aid3fend。
更多 AI 資安筆記:aid3fend.com


上一篇
Day 20|微調實作與原版對照評估
系列文
《30 天打造 AI Guardrails》 共 21 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言