Day 1 講 Agentic Top 10 時說過:agent 的風險從說錯話變成做錯事。Day 3 的架構圖上,跟工具有關的箭頭有三條——工具呼叫參數(出)、工具回傳(入)、記憶寫入(出)。今天把三條都攔起來。
模型決定呼叫工具時,產生的是一段結構化資料(function name + arguments)。它是模型的輸出,但會變成外部系統的輸入。攔截的政策跟一般文字不同,是依工具寫的:
# rules/tools.yaml
- tool: transfer_funds
policy:
amount: {max: 100000}
to_account: {allowlist_ref: "verified_accounts"}
require_l3: true # 無條件送 judge(Day 22)
require_human: {amount_gt: 50000}
- tool: search_documents
policy:
query: {max_len: 500, l1_scan: true}
- tool: send_email
policy:
to: {domain_allowlist: ["company.com.tw"]}
body: {output_scan: true} # 內文過 output rails(PII、URL)
三個層次:參數範圍檢查(確定性,L1)、語意檢查(參數內容過 L2/L3)、人工確認(超過門檻停下來等人)。第三個不是護欄的一部分,是 Day 1 說的「護欄只是配角」——Excessive Agency 的根治靠權限設計與人在迴路,護欄只負責在呼叫的那一刻多看一眼。
工具回傳的內容(搜尋結果、資料庫查詢、網頁內容、其他 agent 的訊息)在 Day 6 是 I-tool 類語料。攔截原則:視同外部輸入,過 input rails,然後才拼進 context。
一個容易漏的細節:工具回傳通常是 JSON,注入句藏在某個字串欄位裡。要逐欄位掃,不是把整個 JSON 序列化成一段掃——後者會被大量結構字元稀釋(跟 Day 9 講 RAG 整包送的問題一樣)。
def scan_tool_result(result: dict) -> dict:
for path, value in walk_strings(result): # 遞迴走訪所有字串值
_, ok, score = input_pipeline.scan(value)
if not ok:
raise GuardBlocked(f"tool_result:{path}")
audit(direction="in", source="tool", path=path, score=score)
return result
Agent 把「使用者偏好」「先前結論」寫進長期記憶時,寫入的內容過 output rails;下次讀出來拼進 context 時,視情況再過一次 input rails。後者的成本高,本系列的設定:寫入必掃、讀出只在記憶來源不可信(多使用者共用記憶)時掃。
Agent Development Kit 有 before_tool_callback 與 after_tool_callback,正好是參數與回傳的兩個攔截點。Model Armor 在這裡以「呼叫 sanitize API」的方式接:
# agent/guarded_agent.py
from google.adk.agents import Agent
def before_tool(tool, args, tool_context):
policy = TOOL_POLICIES.get(tool.name)
if policy and not policy.check_args(args): # L1 範圍檢查
return {"error": "blocked_by_policy"} # 回傳 dict 即跳過工具執行
for k, v in args.items():
if isinstance(v, str) and ma.sanitize_prompt(v).blocked:
return {"error": f"blocked_arg:{k}"}
return None # None = 放行
def after_tool(tool, args, tool_context, tool_response):
for path, value in walk_strings(tool_response):
if ma.sanitize_prompt(value).blocked: # 工具回傳視同輸入
return {"error": f"blocked_result:{path}"}
return None
agent = Agent(model="gemini-2.5-flash", tools=[transfer_funds, search_documents],
before_tool_callback=before_tool, after_tool_callback=after_tool)
【作者確認】ADK callback 的簽章與回傳語意以你安裝的版本為準。 【此處貼 agent 嘗試呼叫超額轉帳被 before_tool 擋下的執行記錄截圖】 【此處貼工具回傳含注入句、被 after_tool 擋下的截圖】
Day 21 的 GuardHook 加兩段:post-call hook 檢查回應裡的 tool_calls,pre-call hook 檢查 messages 中 role: tool 的訊息。
async def async_post_call_success_hook(self, data, user_api_key_dict, response):
msg = response.choices[0].message
for tc in msg.tool_calls or []:
policy = TOOL_POLICIES.get(tc.function.name)
args = json.loads(tc.function.arguments)
if not policy.check_args(args): raise GuardBlocked(f"tool_args:{tc.function.name}")
if policy.require_l3 and judge(json.dumps(args), task)["verdict"] == "block":
raise GuardBlocked("tool_args_l3")
...
Day 24:MCP server 場景。當工具不是你自己寫的、而是接一個第三方 MCP server 時——tool description 本身可能有注入、回傳格式不受控、供應鏈不可信。Model Armor 與地端線各自能做到哪裡。
Instagram @aid3fend。
更多 AI 資安筆記:aid3fend.com