前面幾天,我把 AI Security Lab 從單純的 LLM 一路擴充到 RAG。
目前已經經歷:
Prompt Injection
↓
Security Gateway
↓
Security Event
↓
Wazuh Monitoring
↓
RAG
↓
Indirect Prompt Injection
↓
RAG Defense
做到 Day 23,主要處理的問題都還圍繞在:
LLM 看到了什麼?
LLM 回答了什麼?
就算模型真的受到 Prompt Injection 影響,大部分情況下,最後產生的仍然只是:
Text Response
但是從今天開始,事情會開始變得不太一樣。
因為我要讓 LLM 不只是:
回答問題
而是可以:
決定使用 Tool
↓
呼叫 Tool
↓
對系統真的產生 Action
也就是正式進入:
AI Agent Security
如果是一般 LLM:
User
↓
LLM
↓
Response
例如:
User:
幫我看看 notes.txt 裡面寫什麼。
LLM:
我沒辦法直接讀取你的檔案。
因為 LLM 本身沒有:
File System Access
但是如果替模型提供:
Tools
架構就會變成:
User
↓
LLM
↓
Tool Decision
↓
Tool Execution
↓
Real Resource
例如:
User:
幫我讀取 notes.txt
↓ Agent
Tool:
read_file
↓ Python
open notes.txt
↓ Result
檔案內容
這時候 LLM 已經不只是:
Language Model
而開始具備:
Agency
也就是可以透過工具影響外部環境。
今天先不攻擊 Agent。
跟前面的流程一樣,我想先建立:
Agent Baseline
今天希望做到:
User
↓
qwen3:4b
↓
Tool Decision
↓
Python Tool
↓
Workspace
先提供兩個 Tool:
list_files
read_file
讓 Agent 可以:
列出 Workspace 裡的檔案
讀取 Workspace 裡的檔案
今天暫時不加入:
delete_file
execute_command
PowerShell
send_email
download_file
因為目前只是建立一個安全範圍比較小的 Agent Lab。
做 Agent 有很多現成 Framework。
例如可以直接讓 Framework 幫忙處理:
Tool Calling
Agent Loop
Tool Execution
但這個系列的目的本來就是:
AI Security
如果一開始就把整個 Agent 行為藏在 Framework 裡,
後面 Day 25 要研究:
Tool Abuse
Excessive Agency
Unauthorized Action
反而不容易看清楚到底是哪一層出了問題。
所以今天我還是選擇:
自己寫最基本的 Agent Loop。
流程很單純:
User
↓
LLM
↓
JSON Decision
↓
Python Parse
↓
Tool Execution
這樣每一步我都看得到。
首先回到專案:
cd C:\Users\user\Desktop\AI-Security-Lab
建立:
mkdir agent
mkdir agent\workspace
目前專案多了一個:
AI-Security-Lab/
│
├─ agent/
│ ├─ workspace/
│ │ ├─ notes.txt
│ │ └─ project.txt
│ │
│ ├─ tools.py
│ └─ agent.py
│
├─ app/
├─ attacks/
├─ defense/
├─ logs/
├─ rag/
└─ tests/
這個:
workspace/
就是今天 Agent 可以接觸的環境。
我不希望測試 Agent 一開始就能直接碰:
C:\
或:
C:\Users\user\
因為那裡可能有真正的:
文件
設定
程式
個人資料
而我們現在只是在做 Security Lab。
所以我另外建立:
agent/workspace/
讓 Agent 所有檔案操作都限制在這裡。
目前放:
notes.txt
project.txt
這也是一個很重要的概念:
就算只是測試 Agent,也不要直接給它整台電腦的權限。
第一個:
agent/workspace/notes.txt
內容:
AI Security Lab Agent Workspace
目前正在進行 AI Agent Security 測試。
Day 24 的目標是建立一個可以使用 Tool 的 AI Agent。
目前 Agent 可以列出檔案與讀取 Workspace 中的文字檔案。
第二個:
agent/workspace/project.txt
內容:
Project: AI Security Lab
Components:
- Security Gateway
- RAG Security
- Wazuh Monitoring
- AI Agent
接下來就讓 Agent 嘗試讀取這些資料。
新增:
agent/tools.py
今天只有:
list_files()
read_file()
程式:
from pathlib import Path
BASE_DIR = Path(__file__).resolve().parent
WORKSPACE_DIR = BASE_DIR / "workspace"
def list_files():
files = []
for path in WORKSPACE_DIR.iterdir():
if path.is_file():
files.append(path.name)
return {
"success": True,
"files": files,
}
def read_file(filename):
file_path = (
WORKSPACE_DIR / filename
).resolve()
workspace_path = (
WORKSPACE_DIR.resolve()
)
try:
file_path.relative_to(
workspace_path
)
except ValueError:
return {
"success": False,
"error": (
"Access outside workspace "
"is not allowed."
),
}
if not file_path.exists():
return {
"success": False,
"error": "File not found.",
}
if not file_path.is_file():
return {
"success": False,
"error": "Not a file.",
}
content = file_path.read_text(
encoding="utf-8"
)
return {
"success": True,
"filename": filename,
"content": content,
}
這個 Tool 很簡單:
Agent
↓
list_files()
↓
agent/workspace/
↓
取得檔案名稱
例如:
notes.txt
project.txt
它不需要任何 Argument。
所以未來 LLM 要使用它,只需要產生:
{
"action": "tool",
"tool": "list_files",
"arguments": {}
}
第二個:
read_file()
需要:
filename
例如:
{
"action": "tool",
"tool": "read_file",
"arguments": {
"filename": "notes.txt"
}
}
Python 收到之後才真正執行:
read_file("notes.txt")
也就是:
LLM 決定
↓
Application 執行
這兩層其實要分清楚。
雖然 Day 24 還沒有正式進入 Agent Defense,
但我還是沒有讓:
read_file()
直接接受任何 Path。
程式會檢查:
file_path.relative_to(
workspace_path
)
如果有人嘗試:
../../something.txt
離開:
agent/workspace/
就回傳:
Access outside workspace is not allowed.
這不是 Day 25 要做的完整權限控制。
只是最基本的:
Sandbox Boundary
不然為了測 Agent,直接讓它存取真正的系統檔案沒有必要。
跟 Day 21 做 RAG 時一樣,
我沒有一開始就把全部東西接在一起。
先測:
python
然後:
from agent.tools import list_files, read_file
測試:
print(list_files())
應該可以取得:
notes.txt
project.txt
再測:
print(
read_file("notes.txt")
)
確認可以讀到:
AI Security Lab Agent Workspace
這代表:
Tool Layer
已經可以正常運作。
接下來才讓 LLM 決定什麼時候使用它。
新增:
agent/agent.py
首先定義 Agent 可以使用哪些 Tool:
AGENT_SYSTEM_PROMPT = """
You are an AI Security Lab agent.
You can use the following tools:
1. list_files
Lists files in the agent workspace.
2. read_file
Reads a file from the agent workspace.
When a tool is required, respond ONLY with valid JSON.
For list_files:
{
"action": "tool",
"tool": "list_files",
"arguments": {}
}
For read_file:
{
"action": "tool",
"tool": "read_file",
"arguments": {
"filename": "notes.txt"
}
}
If no tool is required, respond:
{
"action": "answer",
"content": "your answer"
}
"""
也就是要求模型只能做兩種 Decision。
第一種:
ANSWER
第二種:
TOOL
例如使用者問:
什麼是 AI Security?
模型可以回答:
{
"action": "answer",
"content": "AI Security 是..."
}
但如果使用者問:
幫我讀取 notes.txt
模型應該回:
{
"action": "tool",
"tool": "read_file",
"arguments": {
"filename": "notes.txt"
}
}
注意:
這時候 LLM 還沒有真的讀檔案。
它只是產生了一個:
Tool Decision
模型繼續使用前面一路使用的:
qwen3:4b
Ollama API:
OLLAMA_URL = (
"http://localhost:11434/api/chat"
)
MODEL_NAME = "qwen3:4b"
建立:
def ask_agent(user_input):
payload = {
"model": MODEL_NAME,
"messages": [
{
"role": "system",
"content": (
AGENT_SYSTEM_PROMPT
),
},
{
"role": "user",
"content": user_input,
},
],
"stream": False,
"think": False,
}
response = requests.post(
OLLAMA_URL,
json=payload,
timeout=300,
)
response.raise_for_status()
model_output = response.json()[
"message"
]["content"]
return model_output
這裡繼續保留:
timeout = 300
也沒有另外限制 num_predict。
接著建立:
def execute_tool(
tool_name,
arguments
):
if tool_name == "list_files":
return list_files()
if tool_name == "read_file":
return read_file(
arguments.get(
"filename",
""
)
)
return {
"success": False,
"error": "Unknown tool.",
}
這個 Function 就是今天很重要的一個邊界:
LLM
↓
Tool Name + Arguments
↓
execute_tool()
↓
Python Function
↓
Real Resource
以前模型輸出:
我想讀 notes.txt
什麼事情都不會發生。
現在如果模型輸出:
{
"tool": "read_file",
"arguments": {
"filename": "notes.txt"
}
}
Application 就真的可能:
Open File
最後建立:
def run_agent(user_input):
model_output = ask_agent(
user_input
)
try:
decision = json.loads(
model_output
)
except json.JSONDecodeError:
return {
"success": False,
"error": (
"Agent returned "
"invalid JSON."
),
"raw_output": model_output,
}
if decision.get(
"action"
) == "answer":
return {
"success": True,
"action": "answer",
"response": decision.get(
"content",
""
),
}
if decision.get(
"action"
) == "tool":
tool_name = decision.get(
"tool"
)
arguments = decision.get(
"arguments",
{}
)
tool_result = execute_tool(
tool_name,
arguments
)
return {
"success": True,
"action": "tool",
"tool": tool_name,
"arguments": arguments,
"tool_result": tool_result,
}
return {
"success": False,
"error": (
"Unknown agent action."
),
"decision": decision,
}
現在整個 Agent Loop 就變成:
User Input
↓
ask_agent()
↓
qwen3:4b
↓
JSON Decision
↓
json.loads()
↓
Action?
│
├── answer
│ ↓
│ Response
│
└── tool
↓
execute_tool()
↓
Python Function
↓
Tool Result
第一個測試:
from agent.agent import run_agent
result = run_agent(
"幫我列出 workspace 裡有哪些檔案"
)
print(result)
理想上 Agent 會先決定:
{
"action": "tool",
"tool": "list_files",
"arguments": {}
}
接著 Application 執行:
list_files()
最後取得:
notes.txt
project.txt
整個流程:
User
↓
LLM
↓
list_files
↓
Tool Execution
↓
Workspace
↓
File List
這也是這個 Lab 第一次真的讓:
LLM Decision
觸發:
Application Action
第二個:
result = run_agent(
"幫我讀取 notes.txt"
)
print(result)
模型需要自己決定:
需要使用 Tool
然後選:
read_file
Argument:
notes.txt
最後 Python 才真正執行:
read_file("notes.txt")
取得檔案內容。
這裡其實就是 Agent 跟一般 Chatbot 最大的差別之一。
一般 Chatbot:
LLM
↓
Text
Agent:
LLM
↓
Decision
↓
Action
第三個我測:
result = run_agent(
"什麼是 AI Security?"
)
print(result)
這個問題不需要:
list_files
也不需要:
read_file
所以理想情況下 Agent 應該選:
{
"action": "answer",
"content": "..."
}
這個測試是在確認:
Agent 不是什麼問題都直接使用 Tool。
Tool 應該是在:
有需要
的時候才呼叫。
做到這裡之後,我開始把 Agent Tool Call 拆成三層來看。
第一層:
Tool Selection
模型選了:
list_files
還是:
read_file
第二層:
Tool Arguments
例如:
{
"filename": "notes.txt"
}
第三層:
Tool Execution
Application 真的執行:
read_file("notes.txt")
這三件事情不能混在一起。
因為後面做 Agent Security 時,每一層都可能出問題。
以前:
Prompt Injection
↓
LLM
↓
Bad Response
風險主要是:
錯誤回答
敏感資訊洩漏
System Prompt Leakage
但是 Agent:
Prompt Injection
↓
LLM
↓
Bad Tool Decision
↓
Tool Execution
↓
Real Action
風險開始變成:
讀取不該讀的資料
修改不該修改的資料
刪除檔案
執行 Command
呼叫外部 API
寄送資料
所以:
Agent Security 的問題,不只是模型會不會被騙,而是模型被騙之後擁有多少權限。
今天故意只提供:
list_files
read_file
沒有:
write_file
delete_file
execute_command
run_powershell
send_email
因為 Day 24 的目的只是:
建立 Agent Baseline,證明 LLM 可以選擇 Tool 並觸發實際 Application Action。
不是要真的讓 Agent 擁有高風險權限。
尤其:
execute_command
這種 Tool 如果沒有權限限制,
就可能讓:
LLM Output
直接變成:
Operating System Command
攻擊面會突然大很多。
雖然今天還不是 Agent Defense,
但 read_file() 本身已經有:
Workspace Boundary
這件事情其實很重要。
假設只在 System Prompt 寫:
Only read files inside the workspace.
但是 Python Function 寫成:
open(filename)
那真正的安全機制其實只有:
希望 LLM 聽話
但今天是:
LLM
↓
read_file()
↓
Path Validation
↓
Workspace Only
就算模型產生錯誤 Path,
Application Layer 還是可以拒絕。
這跟前面做 Security Gateway 的概念其實一樣:
不要把安全責任全部交給 LLM。
做到 Agent 之後,也開始碰到一個新的概念:
Excessive Agency
簡單來說,就是:
AI Agent 被賦予超過完成任務所需要的能力或權限。
例如今天只是:
幫我整理文件內容
其實可能只需要:
read_file
如果我卻給 Agent:
read_file
write_file
delete_file
execute_command
send_email
database_admin
那就代表它擁有大量根本不需要的權限。
如果 Agent 後來被 Prompt Injection 控制,
影響就會從:
回答錯誤
變成:
真的執行錯誤 Action
這也是 Day 25 要開始研究的問題。
做到 Day 24,我覺得 Agent Security 最值得問的問題不是:
LLM 會不會犯錯?
因為答案幾乎一定是:
會
真正要問的是:
如果 LLM 犯錯,
Application 允許它做到什麼程度?
例如:
LLM 選錯 Tool
↓
能不能執行?
LLM 給錯 Argument
↓
會不會被拒絕?
LLM 想讀 Workspace 外的檔案
↓
Application 會不會阻擋?
這才開始進入:
Agent Security
現在整個 AI Security Lab 又多了一條新的 Trust Boundary。
RAG:
External Document
↓
Retrieved Context
↓
LLM
Agent:
LLM
↓
Tool Decision
↓
Application
↓
Real Resource
所以目前至少已經有:
User Input
↓
Security Boundary
Retrieved Context
↓
Security Boundary
LLM Output
↓
Security Boundary
Tool Call
↓
新的 Security Boundary
今天只是先把最後這條建立起來。
Day 25 才正式攻擊。
目前可以畫成:
User
↓
Security Gateway
↓
LLM
↓
Agent Decision
↓
┌────────────────────┐
│ │
Answer Tool
│ │
↓ ↓
Response Tool Selection
↓
Tool Arguments
↓
execute_tool()
↓
┌──────────────────┐
│ │
list_files read_file
│ │
└────────┬─────────┘
↓
Workspace
這是整個系列第一次讓:
LLM
真正連到:
Real Resource
今天完成:
建立 agent/ 目錄
建立 Agent Workspace
建立 notes.txt
建立 project.txt
建立 list_files Tool
建立 read_file Tool
建立基本 Workspace Boundary
測試 Tool Function
建立 Agent System Prompt
建立 JSON Tool Decision Format
連接 Ollama qwen3:4b
建立 ask_agent()
建立 execute_tool()
建立 run_agent()
測試 list_files Tool Selection
測試 read_file Tool Selection
測試 Tool Arguments
測試 Tool Execution
測試不需要 Tool 的一般問題
完成第一版 AI Agent Baseline
如果用一句話總結 Day 24:
LLM 一旦擁有 Tool,模型的輸出就不再只是文字,而可能變成真正的系統操作。
以前:
LLM Output
↓
User 看
現在:
LLM Output
↓
Application 解析
↓
Tool Execution
↓
Real Action
也因此 AI Security 的問題開始從:
模型說了什麼?
往:
模型能做什麼?
移動。
最近四天剛好看到 AI Application Attack Surface 一步一步增加。
Day 21
RAG
↓
讓 LLM 可以讀外部知識
Day 22
Indirect Prompt Injection
↓
外部知識也可能是攻擊來源
Day 23
RAG Defense
↓
Retrieved Context 也需要 Security Check
Day 24
AI Agent
↓
讓 LLM 可以開始執行 Action
從:
Read
正式走向:
Act
Day 25|Agent Attack:當 AI 擁有太多權限——Excessive Agency
今天的 Agent 還很單純:
list_files
read_file
下一篇要開始故意測它的 Security Boundary。
例如:
Agent 能不能被誘導使用不該使用的 Tool?
Agent 能不能讀取不該讀的 File?
如果 User Prompt 嘗試操控 Tool Arguments 會怎樣?
以及最重要的:
如果我們給 Agent 太多權限,
Prompt Injection 的影響會被放大到什麼程度?
Day 25 就會正式進入:
Prompt Injection
↓
Agent Manipulation
↓
Tool Abuse
↓
Unauthorized Action
也就是 Agent 系統很重要的一個安全問題:
Excessive Agency:真正危險的可能不是 LLM 被騙,而是被騙之後,它剛好有權限真的照做。