iT邦幫忙

2026 iThome 鐵人賽

DAY 27
0
Software Development

30 天從 Full-Stack Engineer 進化到 System Design:從 0 設計可支撐百萬使用者的系統系列 第 27 篇

# Day 27|AI Agent Architecture:Tool Calling、Memory、Planning、Workflow 到底是什麼?

  • 分享至 

  • xImage
  •  

Day 26 我們完成了 RAG System。

今天開始進入另一個很重要的 AI System Design 主題:

AI Agent Architecture

這一篇不會把之前已經詳細解釋過的名詞重新講一次,例如:

LLM
Token
Context Window
Inference
RAG
Embedding
Vector Database
API
Message Queue
Cache
Authentication
Authorization
Retry
Timeout
Idempotency
Observability

今天的重點是新的 Agent Concepts。

只要是第一次出現的新縮寫,我都會像 ACID 一樣拆開:

HITL

H = Human = 人
I = In = 在……之中
T = The = 這個
L = Loop = 迴圈

Human-in-the-Loop
= 人類參與迴圈

而第一次出現的新概念會按照:

是什麼
↓
中文意思
↓
為什麼需要
↓
簡單例子
↓
System Design 怎麼使用

來理解。


1. 先從一個普通 LLM 開始

假設 User 問:

What is the weather in Boston today?

如果 LLM 沒有即時 Weather Data,它可能只能:

根據 Training Knowledge 猜

但 User 真正想要的是:

Current Weather API
↓
Current Temperature
↓
Current Rain Probability

所以我們希望 LLM 不只是:

Generate Text

而是能夠:

Understand Goal
↓
Choose Tool
↓
Call Tool
↓
Read Result
↓
Continue
↓
Answer User

這就是 AI Agent 很重要的核心概念。


2. AI Agent

AI Agent(AI 代理 / 智慧代理)

Agent 原本的英文意思可以理解成:

代理人 / 代表某人執行事情的角色

在 AI System 中:

AI Agent 是一個可以根據 Goal、目前 State 與外部資訊,決定下一個
Action,並透過 Tools 執行工作的 AI System。

普通 LLM:

Question
↓
LLM
↓
Text Answer

AI Agent:

Goal
↓
LLM / Decision Logic
↓
Choose Action
↓
Use Tool
↓
Observe Result
↓
Choose Next Action
↓
...
↓
Final Answer

所以:

Agent 不只是回答問題,而是可以採取 Action。


3. Agentic AI

Agentic AI

Agentic
→ 具有自主採取行動能力的 / 代理式的

AI
→ Artificial Intelligence
→ 人工智慧

中文常見:

代理式 AI

它描述的是:

AI 不只是被動生成一次 Output,而是可以為了完成 Goal,持續做
Decision、使用 Tools、處理 Results,再決定下一步。

例如:

User:
Help me find tomorrow's meetings and draft preparation notes.

Agent 可能:

Read Calendar
↓
Find Meetings
↓
Read Related Documents
↓
Summarize
↓
Generate Preparation Notes

4. Goal

Goal(目標)

Agent 最終想完成的事情。

例如:

Book a restaurant
Find cheapest flight
Prepare meeting brief
Analyze customer issue
Generate weekly report

Agent Architecture 很重要的一件事就是:

Agent 必須知道「現在到底想完成什麼」。


5. Task

Task(任務)

為完成 Goal 所需要執行的一項工作。

例如 Goal:

Prepare tomorrow's meeting

可能拆成 Tasks:

1. Find tomorrow's meeting
2. Identify attendees
3. Read previous notes
4. Summarize open issues
5. Create briefing

Goal 通常比較高層。

Task 通常比較具體。


6. Tool

Tool(工具)

Agent 可以呼叫的外部能力。

例如:

Weather API
Calculator
Database
Search Engine
Calendar
Email
Code Executor
RAG Retriever
CRM

LLM 本身可能不知道:

今天氣溫
User 的 Calendar
Company Database
最新 Stock Price

所以 Agent 需要 Tool。


7. Tool Calling

Tool Calling(工具呼叫)

Agent 決定使用某個 Tool,並產生該 Tool 所需要的 Input。

例如:

User:
What's the weather in Boston?

Agent 不直接猜:

Boston is probably...

而是產生:

Call Weather Tool
location = Boston

Tool Result:

18°C, raining

Agent 再回答:

Boston is currently 18°C and raining.

8. Function

Function(函式)

在 Programming 中:

一段可以接收 Input、執行特定 Logic,並可能回傳 Output
的可重複使用程式。

例如:

def get_weather(city):
    ...

Function 可以被 Agent 當成 Tool 的一種介面。


9. Function Calling

Function Calling(函式呼叫)

讓 Model 選擇某個預先定義的 Function,並產生符合 Function Interface 的
Arguments。

例如 System 定義:

get_weather(city)

User:

What's the weather in Boston?

Model 可能輸出:

Function:
get_weather

Arguments:
city = Boston

重要:

Model 通常不是自己直接執行真正的 Backend Function。

常見流程:

LLM chooses function
↓
Application validates request
↓
Application executes function
↓
Function result returned to LLM

10. Argument

Argument(引數)

呼叫 Function 時傳入的實際 Value。

例如:

get_weather("Boston")

這裡:

"Boston"

就是 Argument。


11. Parameter

Parameter(參數)

Function 定義時,用來接收 Input 的變數。

例如:

def get_weather(city):

city 是 Parameter。

呼叫:

get_weather("Boston")

"Boston" 是 Argument。

簡單記:

Parameter
→ Function 定義裡的名字

Argument
→ 呼叫時真正傳進去的值

12. Schema

Schema(結構規格 / 結構描述)

定義 Data 應該有哪些 Fields、Types 與 Rules。

例如 Weather Tool:

{
  "city": "string",
  "unit": "string"
}

Tool Schema 告訴 Model:

Tool 叫什麼
Tool 做什麼
需要哪些 Parameters
Parameter 是什麼 Type
哪些是 Required

13. Structured Output

Structured Output(結構化輸出)

Output 不是自由文字,而是符合固定 Structure。

例如自由文字:

I think you should call weather with Boston.

Structured Output:

{
  "tool": "get_weather",
  "city": "Boston"
}

Agent System 通常更喜歡 Structured Output,因為 Application
比較容易可靠地 Parse 和 Validate。


14. Validation

Validation(驗證)

檢查 Input / Output 是否符合預期 Rules。

例如 Tool 只接受:

unit = celsius
unit = fahrenheit

如果 Model 產生:

unit = banana

Application 應該 Reject,而不是直接執行。


15. Action

Action(行動)

Agent 決定執行的下一個操作。

例如:

Call Search Tool
Send Email
Query Database
Ask User
Stop

Tool Call 是 Action 的一種。


16. Observation

Observation(觀察結果)

Agent 執行 Action 後,從 Environment 得到的 Result。

例如:

Action:
Search "Boston weather"

Observation:
18°C, rain

Agent 再根據 Observation 決定下一步。


17. Environment

Environment(環境)

Agent 可以互動的外部世界。

例如:

Web
Database
File System
Calendar
Email
Application APIs
User

Agent:

Action
↓
Environment
↓
Observation

18. Agent Loop

Agent Loop(Agent 迴圈)

Agent 重複執行「判斷 → Action → Observation → 再判斷」的 Process。

Goal
↓
Decide next action
↓
Action
↓
Observation
↓
Update state
↓
Decide next action
↓
...
↓
Finish

為什麼叫 Loop?

因為:

不是只做一次

而是可能重複很多次。


19. Reasoning Loop

Reasoning Loop(推理迴圈)

Agent 根據目前已知資訊與 Tool
Results,反覆判斷下一步應該做什麼的高層概念。

在實際 Product 中:

不需要、也不應要求 System 暴露 Model 的私人內部推理內容。

System 真正需要保存的是可操作資訊,例如:

Current Goal
Tool Calls
Tool Results
State
Errors
Approval Status
Final Outcome

20. State

State(狀態)

System 在某個時間點,需要記住的目前資訊。

例如 Agent 正在處理 Travel Task:

Destination = Tokyo
Departure = Boston
Date = Dec 20
Budget = $1,500
Flight found = true
Hotel found = false

這些就是 State。


21. State Machine

State Machine(狀態機)

用有限的 States 與 State Transitions 描述 System
如何從一個狀態移動到另一個狀態。

例如:

START
↓
SEARCHING
↓
WAITING_FOR_APPROVAL
↓
BOOKING
↓
COMPLETED

如果失敗:

BOOKING
↓
FAILED

22. Transition

Transition(狀態轉換)

System 從一個 State 移到另一個 State。

例如:

WAITING_FOR_APPROVAL
↓ User approves
BOOKING

User Approval 就是造成 Transition 的 Event。


23. Planning

Planning(規劃)

Agent 在執行前或執行過程中,決定完成 Goal 可能需要哪些 Steps。

例如 Goal:

Plan a trip to New York

Plan:

1. Determine dates
2. Search transportation
3. Search hotel
4. Compare prices
5. Build itinerary

24. Plan

Plan(計畫)

一組預計執行的 Steps。

注意:

Plan 不代表永遠不能改。

例如:

Hotel unavailable

Agent 可能需要:

Re-plan

25. Re-planning

Re-planning(重新規劃)

當 Observation、Failure 或 Requirement 改變時,重新調整 Plan。

例如:

Original:
Book Hotel A

Observation:
Hotel A sold out

Re-plan:
Search Hotel B / C

26. Planner

Planner(規劃器)

負責決定「要做哪些 Steps」的 Component。

Goal
↓
Planner
↓
Plan

例如:

Goal:
Prepare customer meeting

Planner:
1. Read CRM
2. Find previous emails
3. Summarize open issues
4. Create briefing

27. Executor

Executor(執行器)

負責真正執行 Plan 中 Actions 的 Component。

Plan
↓
Executor
↓
Tool Calls

簡單記:

Planner
→ 想要做什麼

Executor
→ 真正去做

28. Planner-Executor Architecture

Planner-Executor Architecture(規劃器-執行器架構)

User Goal
↓
Planner
↓
Plan
↓
Executor
↓
Tools
↓
Results
↓
Planner / Controller
↓
Next Step

適合比較複雜的 Multi-Step Task。

但簡單 Task 不一定需要這麼複雜。


29. Workflow

Workflow(工作流程)

預先定義好的 Steps 與執行順序。

例如:

Receive Invoice
↓
Extract Data
↓
Validate
↓
Ask Manager Approval
↓
Save to Database

這是一個 Workflow。


30. Agent vs Workflow

這是非常重要的面試觀念。

Workflow

Steps mostly known in advance

例如:

A → B → C → D

Agent

Next step may be dynamically decided

例如:

A
↓
LLM decides
├── B
├── C
└── Ask User

所以:

不是所有 AI Application 都需要 Agent。

如果流程固定:

Workflow

通常更簡單、更可預測。


31. Deterministic

Deterministic(確定性的)

在相同 Input / State 下,System 行為可以按照明確 Rules 得到預期結果。

傳統 Workflow 通常比較 Deterministic。

LLM Agent 通常具有較高的:

Variability

32. Variability

Variability(變異性 / 不固定性)

相似 Input 不一定每次都產生完全相同 Decision 或 Output。

這代表 Agent Design 要特別重視:

Validation
Guardrails
Evaluation
Observability

33. Orchestration

Orchestration(協調 / 編排)

原本可以想成:

指揮很多不同 Components,讓它們按照正確順序合作。

Agent System 中:

LLM
Tools
Memory
Workflow
Approval
Retries
State

都可能需要被協調。

這就是:

Agent Orchestration


34. Orchestrator

Orchestrator(協調器 / 編排器)

負責管理 Agent Workflow、State、Tool Execution 與 Transitions 的
Component。

例如:

Orchestrator
├── Call LLM
├── Execute Tool
├── Save State
├── Check Approval
├── Retry Failure
└── Stop when finished

35. Control Flow

Control Flow(控制流程)

Program / Workflow 決定「下一步執行哪裡」的 Logic。

例如:

if approved:
    send_email
else:
    wait

Agent Architecture 也需要 Control Flow,只是部分 Decision 可能交給 LLM。


36. Branch

Branch(分支)

根據 Condition 選擇不同 Execution Path。

Payment success?
├── Yes → Ship
└── No  → Retry

Agent Workflow 也可能:

Need more information?
├── Yes → Ask User
└── No  → Continue

37. Conditional Routing

Conditional Routing(條件式路由)

根據某個 Condition 把 Task 導向不同 Processing Path。

例如:

Question type
├── HR → HR RAG
├── Engineering → Engineering RAG
└── Finance → Finance RAG

38. Router

Router(路由器 / 路由元件)

根據 Input 決定應該走哪條 Path 或使用哪個 Tool / Agent。

例如:

User Request
↓
Router
├── Search Agent
├── Coding Agent
└── Calendar Agent

39. Memory

Memory(記憶)

在 Agent System 中:

保存過去資訊,讓未來 Decision 可以使用。

例如 User 說:

My preferred airport is BOS.

如果 System 需要之後使用:

BOS

就需要某種 Memory / State Storage。

但 Memory 不是單一技術。


40. Short-Term Memory

Short-Term Memory

Short-Term → 短期
Memory → 記憶

中文:短期記憶

主要保存目前 Task / Conversation 需要的暫時資訊。

例如:

Current conversation
Current plan
Recent tool results
Current task state

Task 結束後可能不需要永久保存。


41. Long-Term Memory

Long-Term Memory

Long-Term → 長期
Memory → 記憶

中文:長期記憶

保存跨 Session 或長時間仍有價值的資訊。

例如:

Stable user preference
Previous project decision
Long-lived profile setting

但 Long-Term Memory 必須考慮:

Privacy
User Control
Freshness
Deletion
Incorrect Memory

42. Session

Session(工作階段 / 會話)

User 與 System 在一段連續 Interaction 中的上下文範圍。

例如:

Open Chat
↓
Ask several questions
↓
Close / expire session

這可以是一個 Session。


43. Conversation Memory

Conversation Memory(對話記憶)

保存 Conversation 中的重要資訊,讓後續 Messages 可以延續 Context。

例如:

User:
I want to visit Tokyo.

Later:
Find hotels under $200.

System 要知道:

Tokyo

來自前文。


44. Working Memory

Working Memory(工作記憶)

Agent 正在完成某個 Task 時,暫時保存的 Active Information。

例如:

Goal
Current Step
Tool Results
Pending Questions
Temporary Variables

它和 Short-Term Memory 很接近,實際產品命名可能不同。


45. Memory Store

Memory Store(記憶儲存層)

實際保存 Memory Data 的 Storage。

可能使用:

Database
Key-Value Store
Vector Store
Document Store

選擇取決於 Memory Type。


46. Memory Retrieval

Memory Retrieval(記憶檢索)

從過去保存的 Memory 中找出目前 Task Relevant 的部分。

如果 Long-Term Memory 有 10,000 筆:

不能每次全部塞進 Prompt

所以可能:

Current Task
↓
Retrieve relevant memories
↓
Add to Context

概念和 RAG 很像。


47. Memory Write

Memory Write(寫入記憶)

決定哪些新資訊應該被保存。

不是 User 說的每一句話都應該永久保存。

要考慮:

Is it useful later?
Is it sensitive?
Is it temporary?
Is it accurate?
Does user expect it to persist?

48. Memory Update

Memory Update(記憶更新)

假設舊 Memory:

Preferred city = Boston

User 後來明確說:

I moved to New York.

System 不能永遠保留舊 Value 當成 Current Truth。

因此 Memory System 需要處理:

Update
Conflict
Freshness
Deletion

49. Stale Memory

Stale Memory(過時記憶)

曾經正確,但現在已經過期的 Memory。

例如:

Old job
Old address
Old preference
Expired project deadline

Memory 越多,不代表 Agent 一定越聰明。

錯誤 Memory 反而會造成錯誤 Decision。


50. Checkpoint

Checkpoint(檢查點 / 狀態保存點)

在 Workflow 執行到某個階段時,把目前 State 保存起來。

例如:

Step 1 complete
↓
Save checkpoint
↓
Step 2
↓
System crashes

Restart 後:

Load checkpoint
↓
Continue from Step 2

不用全部重做。


51. Resume

Resume(恢復 / 繼續執行)

從之前保存的 State 繼續 Task。

Checkpoint 的重要用途之一就是:

Resume after failure

52. Durable Execution

Durable Execution

Durable → 持久、能在 Failure 後保留
Execution → 執行

中文可理解為:

可持久化的工作執行

意思:

Long-running Workflow 即使 Process Crash、Server
Restart,也可以根據保存的 State 繼續。

這對 Agent 很重要,因為有些 Task 可能:

Wait 3 hours
Wait for approval
Call many services
Run for days

53. Long-Running Task

Long-Running Task(長時間執行任務)

不會在幾秒內立刻完成,可能持續很久的 Task。

例如:

Wait for customer reply
Monitor job completion
Wait for manager approval
Process thousands of files

這類 Task 通常不能只靠:

One HTTP Request

一直開著等。


54. Human-in-the-Loop

HITL = Human-in-the-Loop

逐字:

H = Human = 人類
I = In = 在……之中
T = The = 這個
L = Loop = 迴圈

完整:

Human-in-the-Loop
= 人類參與迴圈

意思:

AI Workflow 中某些重要 Step 必須讓 Human Review、Approve、Correct 或
Decide。

例如:

Agent drafts email
↓
Human approves
↓
Agent sends email

55. Approval

Approval(批准 / 核准)

Human 明確允許某個 Action 執行。

例如:

Send $10,000 payment?
↓
Require Approval

而不是 Agent 自己直接執行。


56. Approval Gate

Gate(閘門 / 關卡)

必須通過某個 Condition 才能繼續。

Approval Gate(核准關卡)

Agent prepares action
↓
WAITING_FOR_APPROVAL
↓
Human approves?
├── Yes → Execute
└── No  → Cancel

57. High-Risk Action

High-Risk Action(高風險行動)

如果執行錯誤,可能造成重大 Financial、Security、Legal、Privacy 或
Operational Impact 的 Action。

例如:

Send money
Delete production database
Send external email
Change permissions
Publish public content

這類 Action 常需要:

HITL
Approval
Strict Authorization
Audit Log

58. Reversible Action

Reversible(可逆的)

執行後容易恢復。

例如:

Create draft

通常比:

Send final email

更容易 Undo。

Agent Design 可以優先:

Draft first
↓
Approve
↓
Commit

59. Irreversible Action

Irreversible(不可逆的)

執行後很難或不能恢復。

例如:

Permanently delete data
Transfer money
Publish confidential data

越 Irreversible 的 Action,通常越需要嚴格 Guardrails。


60. Side Effect

Side Effect 在之前已經提過,所以這裡只套用:

Tool Call 如果會改變外部世界,就是有 Side Effect 的 Action。

例如:

Read Calendar
→ mostly read-only

Create Calendar Event
→ side effect

Search Email
→ read-only

Send Email
→ side effect

Agent 必須區分:

Read
vs
Write

61. Read-Only Tool

Read-Only

Read → 讀取
Only → 僅

中文:唯讀

Tool 只能讀 Data,不修改外部 State。

例如:

Search documents
Read calendar
Get weather

通常 Risk 較低。


62. Write Tool

Write Tool(寫入工具)

可以修改外部 State 的 Tool。

例如:

Send Email
Delete File
Create Event
Update Database
Place Order

通常需要更嚴格的:

Authorization
Validation
Confirmation
Audit

63. Least Privilege

Least Privilege

Least → 最少
Privilege → 權限

中文:

最小權限原則

意思:

Agent / Tool 只取得完成 Task 真正需要的最小 Permissions。

例如:

如果 Agent 只需要:

Read Calendar

不要給它:

Delete Calendar

權限。


64. Tool Permission

Tool Permission(工具權限)

Agent 是否被允許使用某個 Tool,以及可以做哪些 Operations。

例如:

Calendar:
READ = allowed
CREATE = allowed
DELETE = denied

65. Tool Allowlist

Allowlist(允許清單)

明確列出可以使用的項目,其他預設不允許。

Tool Allowlist(工具允許清單)

例如:

Allowed:
search_docs
read_calendar
create_draft

Not allowed:
delete_database
send_payment

66. Guardrail

Guardrail(護欄 / 安全限制)

原本 Guardrail 是道路旁防止車子衝出去的護欄。

AI System 中:

限制 Agent 不要執行不允許、危險或不符合 Policy 的 Behavior。

例如:

Validate Tool Arguments
Restrict Permissions
Require Approval
Limit Number of Steps
Block Dangerous Actions

67. Agent Guardrail

Agent Guardrail(Agent 安全護欄)

專門限制 Agent Decision / Tool Use 的 Safety Rules。

例如:

Never send email without approval
Never transfer more than $100
Never query another tenant's data
Never delete production resources

68. Policy

Policy(政策 / 規則)

System 定義什麼 Behavior 被允許、禁止或需要額外條件。

例如:

Payments above $500 require human approval.

就是 Policy。


69. Policy Engine

Policy Engine(政策判斷引擎)

根據 Rules 判斷某個 Action 是否允許執行的 Component。

Proposed Action
↓
Policy Engine
↓
ALLOW / DENY / REQUIRE_APPROVAL

這比完全依賴 LLM 自己記住安全規則更可靠。


70. Tool Error

Tool Error(工具錯誤)

Agent 呼叫 Tool 時,Tool Execution 發生 Failure。

例如:

API timeout
Invalid arguments
Permission denied
Service unavailable
Resource not found

Agent 必須知道:

Should retry?
Should choose another tool?
Should ask user?
Should stop?

71. Error Classification

Classification(分類)

Error Classification(錯誤分類)

根據 Error Type 決定不同 Recovery Strategy。

例如:

Timeout
→ maybe retry

Permission denied
→ don't blindly retry

Invalid argument
→ fix input

Resource not found
→ ask user / choose alternative

72. Recoverable Error

Recoverable Error(可恢復錯誤)

有合理方法可以再次嘗試或修正的 Error。

例如:

Temporary timeout
Temporary service unavailable

73. Non-Recoverable Error

Non-Recoverable Error(不可恢復錯誤)

在目前條件下 Retry 通常沒有意義。

例如:

User has no permission
Invalid account ID
Resource permanently deleted

這時應:

Stop
Ask User
Escalate

而不是 Infinite Retry。


74. Max Steps

Max = Maximum = 最大值

Max Steps(最大步數)

Agent 一次 Task 最多允許執行多少 Steps。

例如:

max_steps = 10

為什麼需要?

因為 Agent 可能:

Search
↓
Search again
↓
Search again
↓
Search again
↓
...

永遠停不下來。


75. Infinite Loop

Infinite Loop

Infinite → 無限的
Loop → 迴圈

中文:無限迴圈

System 一直重複某個 Process,無法正常結束。

Agent Example:

Search → no result
↓
Search same query
↓
no result
↓
Search same query
↓
...

Max Steps 可以防止這種情況。


76. Termination Condition

Termination

Terminate → 終止
Termination → 終止

Condition(條件)

Termination Condition(終止條件)

Agent 什麼時候應該停止 Loop。

例如:

Goal completed
User input required
Max steps reached
Critical error
Approval denied
No valid action available

77. Completion Condition

Completion Condition(完成條件)

用來判斷 Goal 是否真的完成的 Rule。

例如:

Goal:

Create a meeting

不是:

I generated meeting details.

就算完成。

真正 Completion 可能是:

Calendar API returned event_id

這才代表 Event 已建立。


78. Tool Result

Tool Result(工具結果)

Tool Execution 真正回傳的 Output。

Agent 不應只相信:

"I think the email was sent."

應看 Tool Result:

status = sent
message_id = 123

79. Grounding Agent Decisions

Grounding 之前已學過。

Agent 中可以延伸:

Decision 應盡可能依據 Tool Results / System State,而不是只靠 Model
猜測外部世界。

例如:

不要:

"The payment probably succeeded."

而是:

Payment API:
status = SUCCESS
transaction_id = ...

80. Tool Selection

Tool Selection(工具選擇)

Agent 根據 Task 選擇最適合的 Tool。

例如:

Calculate 999 * 837

可能:

Calculator Tool

而不是:

Web Search

81. Tool Description

Tool Description(工具描述)

告訴 Model 這個 Tool 是做什麼的,以及什麼時候應該使用。

例如:

search_company_docs:
Search internal company documents.
Use this when the user asks about internal policies.

Tool Description 寫得不好可能導致:

Wrong Tool Selection

82. Tool Contract

Contract(契約 / 約定)

Software 中:

Component 對 Input、Output 與 Behavior 的明確約定。

Tool Contract(工具契約)

包括:

Input Schema
Output Schema
Possible Errors
Permissions
Side Effects

83. Precondition

Precondition

Pre → 之前
Condition → 條件

中文:前置條件

Action 執行前必須成立的 Condition。

例如:

Before sending payment:
User authenticated
Account exists
Balance sufficient
Approval received

84. Postcondition

Postcondition

Post → 之後
Condition → 條件

中文:後置條件

Action 成功執行後應該成立的 Condition。

例如:

Create calendar event

Postcondition:

Event exists
Event ID returned

85. Agent Architecture:最簡單版本

User
↓
Agent
↓
LLM
↓
Tool Selection
↓
Tool
↓
Observation
↓
LLM
↓
Final Answer

這適合:

One or two tool calls
Simple task
Low risk

86. Stateful Agent Architecture

當 Task 變複雜:

User
↓
Agent Orchestrator
├── LLM
├── State Store
├── Memory
├── Tool Registry
├── Guardrails
└── Checkpoint Store
      ↓
    Tools

每個 Step:

Read State
↓
Decide
↓
Validate
↓
Execute
↓
Observe
↓
Update State
↓
Checkpoint

87. Tool Registry

Registry(登錄表 / 註冊中心)

Tool Registry(工具登錄表)

保存 Agent 可以使用哪些 Tools 以及它們的 Definitions。

例如:

Tool Registry

1. search_docs
2. read_calendar
3. create_calendar_event
4. draft_email
5. send_email

每個 Tool 包含:

Name
Description
Schema
Permission
Risk Level

88. Risk Level

Risk Level(風險等級)

對 Action 潛在 Impact 做分類。

例如:

LOW
→ Search public web

MEDIUM
→ Create draft

HIGH
→ Send external email

CRITICAL
→ Transfer money

Risk Level 可以決定:

Need approval?
Need stronger authentication?
Need audit?

89. Audit Log

Audit Log 之前 Security 已接觸過,這裡套用 Agent:

Who requested action?
Which agent?
Which tool?
What arguments?
When?
Was approval given?
What was result?

Agent 能採取 Action 時,Auditability 特別重要。


90. Multi-Agent System

Multi-Agent System

Multi → 多個
Agent → 代理
System → 系統

中文:

多 Agent 系統

意思:

一個 Application 中有多個 Agents,各自負責不同 Role。

例如:

Coordinator Agent
├── Research Agent
├── Coding Agent
└── Review Agent

91. Role

Role(角色)

某個 Agent 被分配的 Responsibility。

例如:

Research Agent
→ Find information

Review Agent
→ Check output

Execution Agent
→ Use tools

92. Coordinator Agent

Coordinator(協調者)

Coordinator Agent(協調 Agent)

負責把 Task 分配給其他 Agents,並整合 Results。

User Goal
↓
Coordinator
├── Agent A
├── Agent B
└── Agent C
↓
Combine Results

93. Handoff

Handoff(交接)

一個 Agent 把 Task / Context 交給另一個 Agent。

例如:

General Agent
↓
Detect coding problem
↓
Handoff
↓
Coding Agent

Handoff 時要清楚傳:

Goal
Relevant Context
Current State
Constraints

94. Single-Agent vs Multi-Agent

不要看到 Agent 就一定做 Multi-Agent。

Single-Agent

優點:

Simpler
Lower latency
Lower cost
Easier debugging

Multi-Agent

可能適合:

Clearly separated roles
Different permissions
Different tools
Parallel specialized work

缺點:

More coordination
More model calls
More failure points
Higher cost
Harder debugging

所以:

Multi-Agent 不是自動比較高級或比較好。


95. Parallel Execution

Parallel Execution(平行執行)

多個互不依賴的 Tasks 同時執行。

例如:

Search flights ──────┐
Search hotels ───────┼→ Combine
Search attractions ──┘

不用:

Flight 完成
↓
Hotel
↓
Attractions

可以降低 Total Latency。


96. Dependency

Dependency(依賴關係)

Task B 必須依賴 Task A 的 Result 才能執行。

例如:

Find customer ID
↓
Then query customer orders

Order Query 依賴 Customer ID。

這兩個不能隨便 Parallel。


97. DAG

DAG = Directed Acyclic Graph

逐字:

D = Directed = 有方向的
A = Acyclic = 無環的
G = Graph = 圖

中文:

有向無環圖

先拆兩個概念。

Directed

Connection 有方向:

A → B

代表 A 之後才能 B。

Acyclic

Acyclic
= 沒有 Cycle
= 不會繞一圈回到原點

所以:

A → B → C

可以。

但:

A → B → C → A

有 Cycle,就不是 DAG。


98. DAG 為什麼跟 Workflow 有關?

假設:

Task A: Read customer
Task B: Read orders
Task C: Read support tickets
Task D: Generate summary

Dependency:

        A
       / \
      B   C
       \ /
        D

B 和 C 可以 Parallel。

D 要等:

B + C

完成。

這種 Workflow 可以用 DAG 表示。


99. Agent 不一定是 DAG

因為 Agent 可能:

A
↓
B
↓
Observe failure
↓
Go back to A-like search

Agent Loop 可以有 Dynamic Cycles。

所以:

Fixed Workflow
→ often modeled as DAG

Agent Loop
→ may dynamically revisit steps

100. Supervisor

Supervisor(監督者)

負責監控或協調其他 Components / Agents 的角色。

Multi-Agent Architecture 可能:

Supervisor Agent
├── Research Agent
├── Analysis Agent
└── Writer Agent

Supervisor 決定:

Who should work next?
Is result sufficient?
Should retry?
Is task complete?

101. Escalation

Escalation(升級處理)

Agent 無法安全或正確完成 Task 時,把問題交給更高權限的 Human /
System。

例如:

Agent confidence too low
↓
Escalate to human support

或:

Payment exceeds limit
↓
Escalate to manager approval

102. Confidence

Confidence(信心程度)

System 對某個 Prediction / Decision 的確定程度。

注意:

LLM 自己說:

"I'm 95% confident"

不一定等於真正校準過的 95% Probability。

Production System 不應只依賴 Model 自報 Confidence 做高風險決策。


103. Threshold

Threshold(門檻值)

超過或低於某個 Value 時觸發不同 Behavior。

例如:

Risk Score > threshold
↓
Require Human Approval

104. Budget

Budget(預算 / 限額)

Agent 不只需要 Money Budget。

還可以有:

Token Budget
Time Budget
Tool Call Budget
Step Budget
Cost Budget

例如:

Maximum 10 tool calls
Maximum $0.50 per task
Maximum 30 seconds

105. Cost Guardrail

Cost Guardrail(成本護欄)

防止 Agent 因為 Loop 或過多 Tool Calls 造成無限制 Cost。

例如:

max_steps = 10
max_tool_calls = 20
max_cost = $1

106. Time Budget

Time Budget(時間預算)

Task 最多允許花多少時間。

例如:

30 seconds

超過:

Stop
Return partial result
Ask user to continue

取決於 Product Requirement。


107. Partial Result

Partial Result(部分結果)

Task 沒有完全完成,但已經取得部分有用成果。

例如:

Found 8 of 10 requested records.

有些 System 可以回 Partial Result,而不是整個 Request 全部失敗。


108. Agent Failure Modes

Failure Mode(失敗模式)

System 可能以什麼方式失敗。

Agent 常見:

Wrong tool
Wrong arguments
Tool failure
Infinite loop
Duplicate action
Unauthorized action
Stale memory
Bad plan
Premature termination
Never terminates
High cost
Hallucinated tool result

109. Premature Termination

Premature

Premature → 過早的
Termination → 終止

中文:

過早終止

例如 Goal:

Send meeting invitation

Agent 只產生:

Here's the invitation text.

卻說:

Done.

如果真正 Requirement 是建立 Calendar Event,那就是過早認為 Task 完成。


110. Hallucinated Tool Result

Hallucination 已學過。

Agent 特別危險的一種情況:

Agent did not call payment API

卻回答:

Your payment was successful.

因此 Tool-Using Agent 要盡量以:

Actual Tool Result

確認 External Action 是否成功。


111. Duplicate Action

Duplicate Action(重複行動)

例如:

Agent calls send_email
↓
Response times out
↓
Agent retries
↓
Email actually sent twice

這就是為什麼之前學過的:

Idempotency

在 Agent Tool Design 非常重要。


112. Agent + Idempotency

例如 Tool:

create_payment(
    idempotency_key = task_123_payment_1
)

第一次其實成功,但 Response Lost。

Agent Retry:

same idempotency_key

Backend 可以辨識:

This action was already processed.

避免重複付款。


113. Agent + RAG

Day 26 的 RAG 可以變成 Agent 的 Tool:

Agent
↓
Need company knowledge?
↓
RAG Tool
↓
Retrieve documents
↓
Observation
↓
Agent continues

所以:

RAG
≠ Agent

但:

RAG can be a tool used by an Agent

114. Agent + API

之前學過 API。

Agent Tool 本質上常包裝:

External API
Internal API
Database
Search
Code

Agent 不應該直接取得:

Unlimited infrastructure access

而應透過:

Controlled Tool Interface

限制它能做的事情。


115. Agent + Message Queue

對 Long-Running Agent:

User Request
↓
Create Task
↓
Queue
↓
Agent Worker
↓
Execute
↓
Checkpoint
↓
Continue

Message Queue 可以幫助:

Buffer Tasks
Retry Workers
Scale Processing

116. Agent + Database

Database 可以保存:

Task
State
Tool Calls
Approvals
Checkpoints
Audit Logs
Results

例如:

agent_task
-----------
task_id
user_id
goal
status
current_step
created_at
updated_at

117. Agent Task Status

可以設計:

PENDING
RUNNING
WAITING_FOR_USER
WAITING_FOR_APPROVAL
COMPLETED
FAILED
CANCELLED

這些是 Agent Task 的 States。


118. WAITING_FOR_USER

WAITING_FOR_USER(等待使用者)

Agent 缺少必要 Information,需要 User 回答後才能繼續。

例如:

Which date do you want?

Agent 不應亂猜 High-Impact Information。


119. CANCELLED

Cancelled(已取消)

Task 被 User 或 System 主動停止。

例如:

User:
Cancel the booking task.

System:

RUNNING
↓
CANCELLED

120. Cancellation

Cancellation(取消)

停止還沒有完成的 Task。

Long-Running Agent 必須考慮:

Can user cancel?
What if tool call is already executing?
What if an external side effect already happened?

取消 Workflow 不代表可以自動 Undo 已發生的 External Actions。


121. Compensation

Compensation 在 Sharding / Distributed Transaction
可能提過,這裡只套用。

如果 Agent:

1. Create hotel booking
2. Create flight booking
3. Flight fails

可能需要:

Cancel hotel booking

這種「用另一個 Action 補償之前已完成 Action」就是 Compensation 思路。


122. Rollback vs Compensation

Rollback:

像 Database Transaction
→ 還原 Transaction 內的變更

External Systems 常不能真正 Rollback。

例如:

Email sent

不能「沒寄過」。

所以 Agent Workflow 常需要:

Compensating Action

而不是假設所有 Tool Calls 都可以 Rollback。


123. Human Review

Human Review(人工審查)

Human 查看 Agent Output / Proposed Action,再決定是否接受。

例如:

Agent generates contract summary
↓
Lawyer reviews
↓
Approve / Edit / Reject

124. Override

Override(覆寫 / 人工取代決策)

Human 可以取代 Agent 原本 Decision。

例如:

Agent recommends DENY
Human authorized reviewer overrides → APPROVE

是否允許 Override 要依 Business Policy。


125. Agent Observability

Observability 已詳細學過。

Agent 特別應該觀察:

Task completion rate
Tool call count
Tool error rate
Average steps
Agent latency
Cost per task
Approval rate
Escalation rate
Loop rate
Failure reasons

126. Trace

Trace 之前 Observability 已學過。

Agent Trace 可以包含:

Task Started
↓
LLM Decision
↓
Tool A Called
↓
Tool A Result
↓
State Updated
↓
Approval Requested
↓
Approval Received
↓
Tool B Called
↓
Completed

注意:

Production Trace 應記錄可操作 Events / Metadata,不需要暴露 Model
私人的 Chain-of-Thought。


127. Agent Evaluation

Agent Evaluation(Agent 評估)

測量 Agent 是否真的完成 Task,而且方法安全、正確、有效率。

不能只評:

Final text looks good

還要評:

Did it choose correct tools?
Did it complete the goal?
Did it violate policy?
Did it use too many steps?
Did it create duplicate side effects?
Did it ask for approval when required?

128. Task Success Rate

Task Success Rate(任務成功率)

Agent 成功完成 Goal 的比例。

例如:

100 tasks
85 completed correctly
Task Success Rate = 85%

129. Tool Success Rate

Tool Success Rate(工具成功率)

Tool Calls 成功完成的比例。

例如:

1,000 tool calls
950 successful
Tool Success Rate = 95%

但 Tool 成功不代表 Agent Goal 一定成功。


130. Step Efficiency

Efficiency(效率)

Step Efficiency(步驟效率)

Agent 是否使用合理數量的 Steps 完成 Task。

例如:

Agent A:

3 tool calls

完成。

Agent B:

25 tool calls

才完成相同 Task。

即使兩個都成功,B:

More latency
More cost
More failure opportunities

131. Agent Benchmark

Benchmark(基準測試)

使用固定 Tasks / Dataset 比較 System Performance。

Agent Benchmark 可以包含:

Task
Expected outcome
Allowed tools
Policy constraints
Maximum steps
Evaluation criteria

132. Golden Task

類似 Day 26 Golden Dataset。

Golden Task(標準測試任務)

已知正確 Goal、Expected Behavior 與 Expected Outcome 的測試 Task。

例如:

Task:
Find customer order #123 and summarize status.

Expected:
Use order lookup tool
Do not call payment tool
Return current status

133. Agent System Design Example:Meeting Preparation Agent

Requirement:

User:
Prepare me for tomorrow's customer meeting.

Agent 需要:

1. Read Calendar
2. Identify meeting
3. Read attendees
4. Search CRM
5. Search previous notes
6. Search relevant emails
7. Build summary

134. Meeting Agent Architecture

User
↓
Agent API
↓
Authentication
↓
Agent Orchestrator
├── State Store
├── Memory
├── Tool Registry
├── Policy Engine
├── Checkpoint Store
└── LLM
      ↓
      Tools
      ├── Calendar Tool
      ├── CRM Tool
      ├── Email Search Tool
      └── RAG Tool

135. Meeting Agent Flow

Goal:
Prepare tomorrow's customer meeting
↓
Planner
↓
Plan:
1. Find meeting
2. Identify company
3. Get CRM information
4. Find previous communication
5. Summarize
↓
Executor
↓
Calendar Tool
↓
Observation
↓
CRM Tool
↓
Observation
↓
Email Search Tool
↓
Observation
↓
RAG Tool
↓
Observation
↓
Generate briefing
↓
Completion Check
↓
Final Result

136. Example:Agent 需要 Ask User

Goal:

Book a restaurant for tomorrow.

缺少:

City?
Time?
Party size?

Agent 不應:

Guess everything

而應進入:

WAITING_FOR_USER

並詢問必要 Information。


137. Example:Approval Gate

User:

Find an appropriate meeting time and email everyone.

可以:

Read Calendar
↓
Find time
↓
Draft Email
↓
WAITING_FOR_APPROVAL
↓
User approves
↓
Send Email Tool
↓
Verify Tool Result
↓
COMPLETED

這比 Agent 自動寄信更安全。


138. Example:Tool Failure

Agent
↓
Calendar API
↓
Timeout

Error Classification:

Temporary Timeout
→ Recoverable

可以:

Retry with backoff

如果:

Permission Denied

則:

Non-Recoverable under current permissions
↓
Ask user / escalate

139. Example:避免 Infinite Loop

錯誤:

Search Tool
↓
No result
↓
Search same query
↓
No result
↓
Search same query
↓
...

加入:

max_steps = 8
max_search_calls = 3

以及:

If no progress:
Ask user or stop

140. Progress

Progress(進展)

Agent 是否比上一個 Step 更接近 Goal。

例如:

Step 1:
Need customer ID

Step 2:
Found customer ID

有 Progress。

但:

Search same query 5 times

可能沒有 Progress。


141. No-Progress Detection

No-Progress Detection(無進展偵測)

偵測 Agent 是否持續重複 Actions,卻沒有取得新 Information 或接近
Goal。

可以作為:

Termination / Escalation Guardrail

142. Production Agent 的核心原則

不要設計成:

Give LLM every permission
↓
Tell it:
"Do whatever is needed"

更合理:

Clear Goal
↓
Limited Tools
↓
Explicit Schemas
↓
Least Privilege
↓
Validate Arguments
↓
Policy Check
↓
Approval for risky actions
↓
Execute
↓
Verify Result
↓
Save State
↓
Audit

143. 面試題:LLM 和 Agent 差在哪?

可以回答:

An LLM primarily generates outputs from its input context. An AI agent
wraps a model inside a larger system that can maintain state, choose
actions, call external tools, observe results, and continue taking
steps until a goal is completed or a termination condition is reached.

簡單版:

LLM
→ Think / Generate

Agent
→ Decide + Act + Observe + Continue

144. 面試題:Tool Calling 是什麼?

Tool Calling allows the model to select a predefined external
capability and produce structured arguments for it. The application
should validate the call, execute the actual tool, return the result
to the agent, and let the agent decide what to do next.


145. 面試題:Function Calling 和 Tool Calling 差在哪?

可以說:

Function Calling is commonly a specific form of Tool Calling where the
model selects a function and generates its arguments. Tool Calling is
the broader concept because a tool could represent a function, API,
database operation, search system, code executor, or another
controlled capability.


146. 面試題:Agent 和 Workflow 差在哪?

Workflow
→ Execution path mostly predefined

Agent
→ Some next-step decisions are dynamic

回答:

I would prefer a deterministic workflow when the business process is
known in advance because it is easier to test, control, and debug. I
would introduce agentic decision-making only where the next step
genuinely depends on dynamic context.

這句很重要:

Use the least amount of autonomy necessary.


147. Autonomy

Autonomy(自主性)

System 可以在沒有 Human 每一步指示的情況下,自行做多少 Decision /
Action。

High Autonomy:

Agent decides many steps itself

Low Autonomy:

Workflow controls most steps
Agent only handles one decision

Autonomy 越高:

Flexibility ↑
Predictability ↓
Risk ↑
Testing complexity ↑

通常是 Trade-off。


148. 面試題:Memory 分哪些?

可以回答:

Short-Term Memory
→ Current task / conversation

Long-Term Memory
→ Useful information across sessions

Working Memory
→ Active task state

Memory Store
→ Where memory is persisted

Memory Retrieval
→ Find relevant past information

並補充:

More memory is not automatically better. Memory needs privacy
controls, freshness management, update/delete behavior, and relevance
filtering.


149. 面試題:什麼是 HITL?

逐字:

H = Human = 人類
I = In = 在……之中
T = The = 這個
L = Loop = 迴圈

完整:

HITL
= Human-in-the-Loop
= 人類參與迴圈

回答:

HITL means inserting a human review, approval, or correction step into
an AI workflow. It is especially useful before high-risk or
irreversible actions such as sending money, deleting data, changing
permissions, or sending important external communications.


150. 面試題:如何防止 Agent 無限 Loop?

可以回答:

Max Steps
Tool Call Limits
Time Budget
Cost Budget
No-Progress Detection
Termination Conditions
Error Classification
Escalation

不是只靠:

Prompt:
"Please don't loop."

151. 面試題:Tool Call Timeout 後可以直接 Retry 嗎?

不能一概而論。

如果 Read Tool:

Search

Retry 通常比較安全。

如果 Write Tool:

Send Payment

Timeout 不代表:

Action did not happen

可能:

Payment succeeded
but response was lost

所以需要:

Idempotency Key
Status Check

避免 Duplicate Side Effect。


152. 面試題:為什麼 Agent 需要 Checkpoint?

Agent tasks can be long-running and may involve multiple tools,
approvals, or waits. A checkpoint persists the current workflow state
so the system can resume after a crash or restart instead of repeating
the entire task and potentially duplicating side effects.


153. 面試題:Single-Agent vs Multi-Agent?

先選:

Single-Agent

如果 Requirement 真的需要,再 Multi-Agent。

Multi-Agent 適合:

Different specialized roles
Different permissions
Independent parallel tasks
Clear ownership boundaries

但會增加:

Coordination
Latency
Cost
Failure points
Debugging complexity

154. 面試題:什麼是 DAG?

逐字:

D = Directed = 有方向的
A = Acyclic = 無環的
G = Graph = 圖

中文:

有向無環圖

Workflow Example:

      A
     / \
    B   C
     \ /
      D

代表:

B and C depend on A
D depends on B and C

B、C 可以 Parallel。


155. 面試題:怎麼評估 Agent?

不要只看 Final Answer。

可以評:

Task Success Rate
Tool Selection Accuracy
Tool Success Rate
Policy Violation Rate
Average Steps
Latency
Cost per Task
Approval Compliance
Duplicate Action Rate
Escalation Rate

156. Day 27 Interview Checklist

Agent Basics
□ AI Agent
□ Agentic AI
□ Goal
□ Task
□ Tool
□ Tool Calling
□ Function
□ Function Ca

上一篇
# Day 26|Design RAG System:如何讓 LLM 回答 Private Documents,而且不是只靠自己的記憶?
系列文
30 天從 Full-Stack Engineer 進化到 System Design:從 0 設計可支撐百萬使用者的系統 共 27 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言