iT邦幫忙

2026 iThome 鐵人賽

DAY 23
0
AI Security

《30 天打造 AI Guardrails》系列 第 23 篇

Day 23|Agent 工具呼叫護欄:tool input / output 的攔截點

  • 分享至 

  • xImage
  •  

從「說錯話」到「做錯事」

Day 1 講 Agentic Top 10 時說過:agent 的風險從說錯話變成做錯事。Day 3 的架構圖上,跟工具有關的箭頭有三條——工具呼叫參數(出)、工具回傳(入)、記憶寫入(出)。今天把三條都攔起來。

工具呼叫參數:輸出端的特殊型態

模型決定呼叫工具時,產生的是一段結構化資料(function name + arguments)。它是模型的輸出,但會變成外部系統的輸入。攔截的政策跟一般文字不同,是依工具寫的:

# rules/tools.yaml
- tool: transfer_funds
  policy:
    amount: {max: 100000}
    to_account: {allowlist_ref: "verified_accounts"}
    require_l3: true                 # 無條件送 judge(Day 22)
    require_human: {amount_gt: 50000}
- tool: search_documents
  policy:
    query: {max_len: 500, l1_scan: true}
- tool: send_email
  policy:
    to: {domain_allowlist: ["company.com.tw"]}
    body: {output_scan: true}        # 內文過 output rails(PII、URL)

 

三個層次:參數範圍檢查(確定性,L1)、語意檢查(參數內容過 L2/L3)、人工確認(超過門檻停下來等人)。第三個不是護欄的一部分,是 Day 1 說的「護欄只是配角」——Excessive Agency 的根治靠權限設計與人在迴路,護欄只負責在呼叫的那一刻多看一眼。

工具回傳:間接注入的主要入口

工具回傳的內容(搜尋結果、資料庫查詢、網頁內容、其他 agent 的訊息)在 Day 6 是 I-tool 類語料。攔截原則:視同外部輸入,過 input rails,然後才拼進 context。

一個容易漏的細節:工具回傳通常是 JSON,注入句藏在某個字串欄位裡。要逐欄位掃,不是把整個 JSON 序列化成一段掃——後者會被大量結構字元稀釋(跟 Day 9 講 RAG 整包送的問題一樣)。

def scan_tool_result(result: dict) -> dict:
    for path, value in walk_strings(result):        # 遞迴走訪所有字串值
        _, ok, score = input_pipeline.scan(value)
        if not ok:
            raise GuardBlocked(f"tool_result:{path}")
        audit(direction="in", source="tool", path=path, score=score)
    return result

 

記憶寫入

Agent 把「使用者偏好」「先前結論」寫進長期記憶時,寫入的內容過 output rails;下次讀出來拼進 context 時,視情況再過一次 input rails。後者的成本高,本系列的設定:寫入必掃、讀出只在記憶來源不可信(多使用者共用記憶)時掃。

雲端線:Google ADK 的 callback

Agent Development Kit 有 before_tool_callback 與 after_tool_callback,正好是參數與回傳的兩個攔截點。Model Armor 在這裡以「呼叫 sanitize API」的方式接:

# agent/guarded_agent.py
from google.adk.agents import Agent

def before_tool(tool, args, tool_context):
    policy = TOOL_POLICIES.get(tool.name)
    if policy and not policy.check_args(args):           # L1 範圍檢查
        return {"error": "blocked_by_policy"}            # 回傳 dict 即跳過工具執行
    for k, v in args.items():
        if isinstance(v, str) and ma.sanitize_prompt(v).blocked:
            return {"error": f"blocked_arg:{k}"}
    return None                                          # None = 放行

def after_tool(tool, args, tool_context, tool_response):
    for path, value in walk_strings(tool_response):
        if ma.sanitize_prompt(value).blocked:            # 工具回傳視同輸入
            return {"error": f"blocked_result:{path}"}
    return None

agent = Agent(model="gemini-2.5-flash", tools=[transfer_funds, search_documents],
              before_tool_callback=before_tool, after_tool_callback=after_tool)

 

【作者確認】ADK callback 的簽章與回傳語意以你安裝的版本為準。 【此處貼 agent 嘗試呼叫超額轉帳被 before_tool 擋下的執行記錄截圖】 【此處貼工具回傳含注入句、被 after_tool 擋下的截圖】

地端線:LiteLLM hook 的擴充

Day 21 的 GuardHook 加兩段:post-call hook 檢查回應裡的 tool_calls,pre-call hook 檢查 messages 中 role: tool 的訊息。

async def async_post_call_success_hook(self, data, user_api_key_dict, response):
    msg = response.choices[0].message
    for tc in msg.tool_calls or []:
        policy = TOOL_POLICIES.get(tc.function.name)
        args = json.loads(tc.function.arguments)
        if not policy.check_args(args): raise GuardBlocked(f"tool_args:{tc.function.name}")
        if policy.require_l3 and judge(json.dumps(args), task)["verdict"] == "block":
            raise GuardBlocked("tool_args_l3")
    ...

 

明天預告

Day 24:MCP server 場景。當工具不是你自己寫的、而是接一個第三方 MCP server 時——tool description 本身可能有注入、回傳格式不受控、供應鏈不可信。Model Armor 與地端線各自能做到哪裡。


追蹤 AId3fend

Instagram @aid3fend。

更多 AI 資安筆記:aid3fend.com


上一篇
Day 22|L3 LLM-as-judge:何時開、成本多少、判斷如何可解釋
系列文
《30 天打造 AI Guardrails》 共 23 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言