iT邦幫忙

2026 iThome 鐵人賽

DAY 24
0
AI Security

打造 AI Security Lab:從攻擊 LLM 到建立自己的 AI 防線系列 第 24 篇

Day 24|AI Agent:當 LLM 不只會回答,還可以真的執行動作

  • 分享至 

  • xImage
  •  

前言

前面幾天,我把 AI Security Lab 從單純的 LLM 一路擴充到 RAG。

目前已經經歷:

Prompt Injection
↓
Security Gateway
↓
Security Event
↓
Wazuh Monitoring
↓
RAG
↓
Indirect Prompt Injection
↓
RAG Defense

做到 Day 23,主要處理的問題都還圍繞在:

LLM 看到了什麼?
LLM 回答了什麼?

就算模型真的受到 Prompt Injection 影響,大部分情況下,最後產生的仍然只是:

Text Response

但是從今天開始,事情會開始變得不太一樣。

因為我要讓 LLM 不只是:

回答問題

而是可以:

決定使用 Tool
↓
呼叫 Tool
↓
對系統真的產生 Action

也就是正式進入:

AI Agent Security


什麼是 AI Agent?

如果是一般 LLM:

User
↓
LLM
↓
Response

例如:

User:
幫我看看 notes.txt 裡面寫什麼。

LLM:
我沒辦法直接讀取你的檔案。

因為 LLM 本身沒有:

File System Access

但是如果替模型提供:

Tools

架構就會變成:

User
↓
LLM
↓
Tool Decision
↓
Tool Execution
↓
Real Resource

例如:

User:
幫我讀取 notes.txt

↓ Agent

Tool:
read_file

↓ Python

open notes.txt

↓ Result

檔案內容

這時候 LLM 已經不只是:

Language Model

而開始具備:

Agency

也就是可以透過工具影響外部環境。


Day 24 的目標

今天先不攻擊 Agent。

跟前面的流程一樣,我想先建立:

Agent Baseline

今天希望做到:

User
↓
qwen3:4b
↓
Tool Decision
↓
Python Tool
↓
Workspace

先提供兩個 Tool:

list_files
read_file

讓 Agent 可以:

列出 Workspace 裡的檔案
讀取 Workspace 裡的檔案

今天暫時不加入:

delete_file
execute_command
PowerShell
send_email
download_file

因為目前只是建立一個安全範圍比較小的 Agent Lab。


為什麼這次還是不用 LangChain?

做 Agent 有很多現成 Framework。

例如可以直接讓 Framework 幫忙處理:

Tool Calling
Agent Loop
Tool Execution

但這個系列的目的本來就是:

AI Security

如果一開始就把整個 Agent 行為藏在 Framework 裡,

後面 Day 25 要研究:

Tool Abuse
Excessive Agency
Unauthorized Action

反而不容易看清楚到底是哪一層出了問題。

所以今天我還是選擇:

自己寫最基本的 Agent Loop。

流程很單純:

User
↓
LLM
↓
JSON Decision
↓
Python Parse
↓
Tool Execution

這樣每一步我都看得到。


建立 Agent 目錄

首先回到專案:

cd C:\Users\user\Desktop\AI-Security-Lab

建立:

mkdir agent
mkdir agent\workspace

目前專案多了一個:

AI-Security-Lab/
│
├─ agent/
│  ├─ workspace/
│  │  ├─ notes.txt
│  │  └─ project.txt
│  │
│  ├─ tools.py
│  └─ agent.py
│
├─ app/
├─ attacks/
├─ defense/
├─ logs/
├─ rag/
└─ tests/

這個:

workspace/

就是今天 Agent 可以接觸的環境。


為什麼另外建立 Workspace?

我不希望測試 Agent 一開始就能直接碰:

C:\

或:

C:\Users\user\

因為那裡可能有真正的:

文件
設定
程式
個人資料

而我們現在只是在做 Security Lab。

所以我另外建立:

agent/workspace/

讓 Agent 所有檔案操作都限制在這裡。

目前放:

notes.txt
project.txt

這也是一個很重要的概念:

就算只是測試 Agent,也不要直接給它整台電腦的權限。


建立測試檔案

第一個:

agent/workspace/notes.txt

內容:

AI Security Lab Agent Workspace

目前正在進行 AI Agent Security 測試。

Day 24 的目標是建立一個可以使用 Tool 的 AI Agent。

目前 Agent 可以列出檔案與讀取 Workspace 中的文字檔案。

第二個:

agent/workspace/project.txt

內容:

Project: AI Security Lab

Components:
- Security Gateway
- RAG Security
- Wazuh Monitoring
- AI Agent

接下來就讓 Agent 嘗試讀取這些資料。


建立第一批 Agent Tools

新增:

agent/tools.py

今天只有:

list_files()
read_file()

程式:

from pathlib import Path


BASE_DIR = Path(__file__).resolve().parent
WORKSPACE_DIR = BASE_DIR / "workspace"


def list_files():

    files = []

    for path in WORKSPACE_DIR.iterdir():

        if path.is_file():
            files.append(path.name)

    return {
        "success": True,
        "files": files,
    }


def read_file(filename):

    file_path = (
        WORKSPACE_DIR / filename
    ).resolve()

    workspace_path = (
        WORKSPACE_DIR.resolve()
    )

    try:
        file_path.relative_to(
            workspace_path
        )

    except ValueError:

        return {
            "success": False,
            "error": (
                "Access outside workspace "
                "is not allowed."
            ),
        }

    if not file_path.exists():

        return {
            "success": False,
            "error": "File not found.",
        }

    if not file_path.is_file():

        return {
            "success": False,
            "error": "Not a file.",
        }

    content = file_path.read_text(
        encoding="utf-8"
    )

    return {
        "success": True,
        "filename": filename,
        "content": content,
    }

第一個 Tool:list_files

這個 Tool 很簡單:

Agent
↓
list_files()
↓
agent/workspace/
↓
取得檔案名稱

例如:

notes.txt
project.txt

它不需要任何 Argument。

所以未來 LLM 要使用它,只需要產生:

{
  "action": "tool",
  "tool": "list_files",
  "arguments": {}
}

第二個 Tool:read_file

第二個:

read_file()

需要:

filename

例如:

{
  "action": "tool",
  "tool": "read_file",
  "arguments": {
    "filename": "notes.txt"
  }
}

Python 收到之後才真正執行:

read_file("notes.txt")

也就是:

LLM 決定
↓
Application 執行

這兩層其實要分清楚。


最基本的 Path Protection

雖然 Day 24 還沒有正式進入 Agent Defense,

但我還是沒有讓:

read_file()

直接接受任何 Path。

程式會檢查:

file_path.relative_to(
    workspace_path
)

如果有人嘗試:

../../something.txt

離開:

agent/workspace/

就回傳:

Access outside workspace is not allowed.

這不是 Day 25 要做的完整權限控制。

只是最基本的:

Sandbox Boundary

不然為了測 Agent,直接讓它存取真正的系統檔案沒有必要。


先單獨測試 Tool

跟 Day 21 做 RAG 時一樣,

我沒有一開始就把全部東西接在一起。

先測:

python

然後:

from agent.tools import list_files, read_file

測試:

print(list_files())

應該可以取得:

notes.txt
project.txt

再測:

print(
    read_file("notes.txt")
)

確認可以讀到:

AI Security Lab Agent Workspace

這代表:

Tool Layer

已經可以正常運作。

接下來才讓 LLM 決定什麼時候使用它。


建立 Agent System Prompt

新增:

agent/agent.py

首先定義 Agent 可以使用哪些 Tool:

AGENT_SYSTEM_PROMPT = """
You are an AI Security Lab agent.

You can use the following tools:

1. list_files
   Lists files in the agent workspace.

2. read_file
   Reads a file from the agent workspace.

When a tool is required, respond ONLY with valid JSON.

For list_files:

{
  "action": "tool",
  "tool": "list_files",
  "arguments": {}
}

For read_file:

{
  "action": "tool",
  "tool": "read_file",
  "arguments": {
    "filename": "notes.txt"
  }
}

If no tool is required, respond:

{
  "action": "answer",
  "content": "your answer"
}
"""

也就是要求模型只能做兩種 Decision。

第一種:

ANSWER

第二種:

TOOL

Agent Decision

例如使用者問:

什麼是 AI Security?

模型可以回答:

{
  "action": "answer",
  "content": "AI Security 是..."
}

但如果使用者問:

幫我讀取 notes.txt

模型應該回:

{
  "action": "tool",
  "tool": "read_file",
  "arguments": {
    "filename": "notes.txt"
  }
}

注意:

這時候 LLM 還沒有真的讀檔案。

它只是產生了一個:

Tool Decision

呼叫 Ollama

模型繼續使用前面一路使用的:

qwen3:4b

Ollama API:

OLLAMA_URL = (
    "http://localhost:11434/api/chat"
)

MODEL_NAME = "qwen3:4b"

建立:

def ask_agent(user_input):

    payload = {
        "model": MODEL_NAME,
        "messages": [
            {
                "role": "system",
                "content": (
                    AGENT_SYSTEM_PROMPT
                ),
            },
            {
                "role": "user",
                "content": user_input,
            },
        ],
        "stream": False,
        "think": False,
    }

    response = requests.post(
        OLLAMA_URL,
        json=payload,
        timeout=300,
    )

    response.raise_for_status()

    model_output = response.json()[
        "message"
    ]["content"]

    return model_output

這裡繼續保留:

timeout = 300

也沒有另外限制 num_predict。


從 LLM Decision 到真正 Tool Execution

接著建立:

def execute_tool(
    tool_name,
    arguments
):

    if tool_name == "list_files":

        return list_files()

    if tool_name == "read_file":

        return read_file(
            arguments.get(
                "filename",
                ""
            )
        )

    return {
        "success": False,
        "error": "Unknown tool.",
    }

這個 Function 就是今天很重要的一個邊界:

LLM
↓
Tool Name + Arguments
↓
execute_tool()
↓
Python Function
↓
Real Resource

以前模型輸出:

我想讀 notes.txt

什麼事情都不會發生。

現在如果模型輸出:

{
  "tool": "read_file",
  "arguments": {
    "filename": "notes.txt"
  }
}

Application 就真的可能:

Open File

建立第一版 Agent Loop

最後建立:

def run_agent(user_input):

    model_output = ask_agent(
        user_input
    )

    try:

        decision = json.loads(
            model_output
        )

    except json.JSONDecodeError:

        return {
            "success": False,
            "error": (
                "Agent returned "
                "invalid JSON."
            ),
            "raw_output": model_output,
        }

    if decision.get(
        "action"
    ) == "answer":

        return {
            "success": True,
            "action": "answer",
            "response": decision.get(
                "content",
                ""
            ),
        }

    if decision.get(
        "action"
    ) == "tool":

        tool_name = decision.get(
            "tool"
        )

        arguments = decision.get(
            "arguments",
            {}
        )

        tool_result = execute_tool(
            tool_name,
            arguments
        )

        return {
            "success": True,
            "action": "tool",
            "tool": tool_name,
            "arguments": arguments,
            "tool_result": tool_result,
        }

    return {
        "success": False,
        "error": (
            "Unknown agent action."
        ),
        "decision": decision,
    }

現在整個 Agent Loop 就變成:

User Input
↓
ask_agent()
↓
qwen3:4b
↓
JSON Decision
↓
json.loads()
↓
Action?
│
├── answer
│      ↓
│   Response
│
└── tool
       ↓
   execute_tool()
       ↓
   Python Function
       ↓
   Tool Result

Test 1:列出 Workspace 檔案

第一個測試:

from agent.agent import run_agent

result = run_agent(
    "幫我列出 workspace 裡有哪些檔案"
)

print(result)

理想上 Agent 會先決定:

{
  "action": "tool",
  "tool": "list_files",
  "arguments": {}
}

接著 Application 執行:

list_files()

最後取得:

notes.txt
project.txt

整個流程:

User
↓
LLM
↓
list_files
↓
Tool Execution
↓
Workspace
↓
File List

這也是這個 Lab 第一次真的讓:

LLM Decision

觸發:

Application Action

Test 2:讀取檔案

第二個:

result = run_agent(
    "幫我讀取 notes.txt"
)

print(result)

模型需要自己決定:

需要使用 Tool

然後選:

read_file

Argument:

notes.txt

最後 Python 才真正執行:

read_file("notes.txt")

取得檔案內容。

這裡其實就是 Agent 跟一般 Chatbot 最大的差別之一。

一般 Chatbot:

LLM
↓
Text

Agent:

LLM
↓
Decision
↓
Action

Test 3:不需要 Tool 的問題

第三個我測:

result = run_agent(
    "什麼是 AI Security?"
)

print(result)

這個問題不需要:

list_files

也不需要:

read_file

所以理想情況下 Agent 應該選:

{
  "action": "answer",
  "content": "..."
}

這個測試是在確認:

Agent 不是什麼問題都直接使用 Tool。

Tool 應該是在:

有需要

的時候才呼叫。


Tool Selection、Arguments、Execution

做到這裡之後,我開始把 Agent Tool Call 拆成三層來看。

第一層:

Tool Selection

模型選了:

list_files

還是:

read_file

第二層:

Tool Arguments

例如:

{
  "filename": "notes.txt"
}

第三層:

Tool Execution

Application 真的執行:

read_file("notes.txt")

這三件事情不能混在一起。

因為後面做 Agent Security 時,每一層都可能出問題。


Agent 最大的變化:LLM 開始影響真實系統

以前:

Prompt Injection
↓
LLM
↓
Bad Response

風險主要是:

錯誤回答
敏感資訊洩漏
System Prompt Leakage

但是 Agent:

Prompt Injection
↓
LLM
↓
Bad Tool Decision
↓
Tool Execution
↓
Real Action

風險開始變成:

讀取不該讀的資料
修改不該修改的資料
刪除檔案
執行 Command
呼叫外部 API
寄送資料

所以:

Agent Security 的問題,不只是模型會不會被騙,而是模型被騙之後擁有多少權限。


為什麼今天只給兩個 Tool?

今天故意只提供:

list_files
read_file

沒有:

write_file
delete_file
execute_command
run_powershell
send_email

因為 Day 24 的目的只是:

建立 Agent Baseline,證明 LLM 可以選擇 Tool 並觸發實際 Application Action。

不是要真的讓 Agent 擁有高風險權限。

尤其:

execute_command

這種 Tool 如果沒有權限限制,

就可能讓:

LLM Output

直接變成:

Operating System Command

攻擊面會突然大很多。


Tool 本身也必須有 Security Boundary

雖然今天還不是 Agent Defense,

但 read_file() 本身已經有:

Workspace Boundary

這件事情其實很重要。

假設只在 System Prompt 寫:

Only read files inside the workspace.

但是 Python Function 寫成:

open(filename)

那真正的安全機制其實只有:

希望 LLM 聽話

但今天是:

LLM
↓
read_file()
↓
Path Validation
↓
Workspace Only

就算模型產生錯誤 Path,

Application Layer 還是可以拒絕。

這跟前面做 Security Gateway 的概念其實一樣:

不要把安全責任全部交給 LLM。


什麼是 Excessive Agency?

做到 Agent 之後,也開始碰到一個新的概念:

Excessive Agency

簡單來說,就是:

AI Agent 被賦予超過完成任務所需要的能力或權限。

例如今天只是:

幫我整理文件內容

其實可能只需要:

read_file

如果我卻給 Agent:

read_file
write_file
delete_file
execute_command
send_email
database_admin

那就代表它擁有大量根本不需要的權限。

如果 Agent 後來被 Prompt Injection 控制,

影響就會從:

回答錯誤

變成:

真的執行錯誤 Action

這也是 Day 25 要開始研究的問題。


Agent Security 的核心問題

做到 Day 24,我覺得 Agent Security 最值得問的問題不是:

LLM 會不會犯錯?

因為答案幾乎一定是:

會

真正要問的是:

如果 LLM 犯錯,
Application 允許它做到什麼程度?

例如:

LLM 選錯 Tool
↓
能不能執行?
LLM 給錯 Argument
↓
會不會被拒絕?
LLM 想讀 Workspace 外的檔案
↓
Application 會不會阻擋?

這才開始進入:

Agent Security

Day 24 的 Trust Boundary

現在整個 AI Security Lab 又多了一條新的 Trust Boundary。

RAG:

External Document
↓
Retrieved Context
↓
LLM

Agent:

LLM
↓
Tool Decision
↓
Application
↓
Real Resource

所以目前至少已經有:

User Input
↓
Security Boundary

Retrieved Context
↓
Security Boundary

LLM Output
↓
Security Boundary

Tool Call
↓
新的 Security Boundary

今天只是先把最後這條建立起來。

Day 25 才正式攻擊。


Day 24 完成後的架構

目前可以畫成:

                         User
                           ↓
                  Security Gateway
                           ↓
                          LLM
                           ↓
                    Agent Decision
                           ↓
               ┌────────────────────┐
               │                    │
             Answer               Tool
               │                    │
               ↓                    ↓
           Response           Tool Selection
                                    ↓
                              Tool Arguments
                                    ↓
                              execute_tool()
                                    ↓
                          ┌──────────────────┐
                          │                  │
                     list_files         read_file
                          │                  │
                          └────────┬─────────┘
                                   ↓
                                Workspace

這是整個系列第一次讓:

LLM

真正連到:

Real Resource

Day 24 小結

今天完成:

建立 agent/ 目錄

建立 Agent Workspace

建立 notes.txt

建立 project.txt

建立 list_files Tool

建立 read_file Tool

建立基本 Workspace Boundary

測試 Tool Function

建立 Agent System Prompt

建立 JSON Tool Decision Format

連接 Ollama qwen3:4b

建立 ask_agent()

建立 execute_tool()

建立 run_agent()

測試 list_files Tool Selection

測試 read_file Tool Selection

測試 Tool Arguments

測試 Tool Execution

測試不需要 Tool 的一般問題

完成第一版 AI Agent Baseline

今天最大的收穫

如果用一句話總結 Day 24:

LLM 一旦擁有 Tool,模型的輸出就不再只是文字,而可能變成真正的系統操作。

以前:

LLM Output
↓
User 看

現在:

LLM Output
↓
Application 解析
↓
Tool Execution
↓
Real Action

也因此 AI Security 的問題開始從:

模型說了什麼?

往:

模型能做什麼?

移動。


從 Day 21 到 Day 24

最近四天剛好看到 AI Application Attack Surface 一步一步增加。

Day 21
RAG
↓
讓 LLM 可以讀外部知識

Day 22
Indirect Prompt Injection
↓
外部知識也可能是攻擊來源

Day 23
RAG Defense
↓
Retrieved Context 也需要 Security Check

Day 24
AI Agent
↓
讓 LLM 可以開始執行 Action

從:

Read

正式走向:

Act

下一篇

Day 25|Agent Attack:當 AI 擁有太多權限——Excessive Agency

今天的 Agent 還很單純:

list_files
read_file

下一篇要開始故意測它的 Security Boundary。

例如:

Agent 能不能被誘導使用不該使用的 Tool?
Agent 能不能讀取不該讀的 File?
如果 User Prompt 嘗試操控 Tool Arguments 會怎樣?

以及最重要的:

如果我們給 Agent 太多權限,
Prompt Injection 的影響會被放大到什麼程度?

Day 25 就會正式進入:

Prompt Injection
↓
Agent Manipulation
↓
Tool Abuse
↓
Unauthorized Action

也就是 Agent 系統很重要的一個安全問題:

Excessive Agency:真正危險的可能不是 LLM 被騙,而是被騙之後,它剛好有權限真的照做。


上一篇
Day 23|RAG Defense:不要相信 Retrieved Context
系列文
打造 AI Security Lab:從攻擊 LLM 到建立自己的 AI 防線 共 24 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言