在上一篇中,我們用 MessagesState 與 tools_condition 建立了最基本的 Agent Loop。如下圖所示,通報輸入後,模型決策節點根據目前的訊息決定是否要求 Tool 呼叫;有工具呼叫時,tools_condition 便將流程導向工具執行節點。工具執行結果寫回 State,再交由模型決策節點讀取,決定下一步要查什麼。例如,先搜尋日誌,再根據錯誤呼叫 check_service_health 驗證 Redis。當模型認為證據已足夠,就停止要求工具呼叫並產出根因報告,由 tools_condition 將流程導向 END。

這條流程能順利運作,完全建立在「模型每次都能在正確時機停手」的假設上。在 tools_condition 的機制中,控制流程是否繼續的唯一依據是模型最後一則訊息有沒有包含 tool_calls。這意味著迴圈的控制權完全交給了模型:只要模型持續發出工具呼叫,迴圈就不會結束;模型一旦停止呼叫工具,程式就直接認定任務完成。
把這種未加防護的迴圈放到真實環境時,可能會遇到失控的情況:
這兩種失控代表邊界控制的兩個缺口:模型可能該停不停,也可能不該停卻停。程式不能把邊界控制權交給模型的自由意願,而必須在 State 中維護確定性的控制變數,由 Python 程式剛性守住兩個邊界:
completed;缺少證據或超額中斷時則標記為 incomplete。
這篇的範例在 ai-agent-sample/langgraph/langgraph-incident-investigation/ 目錄,核心實作在 workflow.py。
上一篇的State 只用 messages 傳遞對話。但如果只有對話紀錄,程式無法得知模型已經呼叫了幾次工具,也無法確認模型最後輸出的結論是真的有查證過,還是未經查證的隨意猜測。
為了解決這兩個問題,我們在 workflow.py 的 State 中加入次數上限與結案狀態:
import operator
from typing import Annotated, Literal, TypedDict
from langchain_core.messages import AnyMessage
from langgraph.graph.message import add_messages
IncidentStatus = Literal["investigating", "completed", "incomplete"]
class State(TypedDict):
incident_description: str # 原始事故通報
messages: Annotated[list[AnyMessage], add_messages] # 對話與工具回傳訊息
tool_call_count: int # 模型累計要求的工具呼叫數
max_tool_calls: int # 允許執行的工具呼叫上限
status: IncidentStatus # 調查狀態(investigating / completed / incomplete)
final_report: str | None # 驗收通過的最終報告
stop_reason: str | None # 未完成原因
events: Annotated[list[str], operator.add] # 終端機日誌
這份 State 新增的欄位各司其職:
tool_call_count、max_tool_calls):模型每發起一次工具呼叫,程式就把計數往上加。後續的 Router 只要比對這兩個數字,就能在工具真正執行前攔截超額呼叫。status、final_report、stop_reason):調查啟動時 status 為 investigating。模型在訊息裡宣稱完成不算數,必須等流程結束時由結案節點檢查歷史事實——確認查過日誌與健康度才標記為 completed,超額或證據不足則標記為 incomplete。messages、events):使用 Annotated 指定 reducer(add_messages 與 operator.add)。各節點更新 State 時只需回傳異動欄位,例如回傳 {"tool_call_count": 1} 即可覆寫計數,而回傳的訊息或日誌則會自動追加到清單尾端。State 負責完整記錄每次排查的資料與狀態。有了這些變數,下一步就能讓 Router 依據次數上限守住迴圈邊界。
Agent 節點呼叫模型後,第一件事就是更新 State 中的工具呼叫計數。一則 AIMessage 可能同時包含多個平行 tool call,因此必須累計 len(response.tool_calls) 筆數,而不是固定加一:
def investigate(state: State) -> StateUpdate:
# 展開 State 累積的訊息(包含初始通報、先前決策與 ToolMessage),供模型推論下一步。
prompt = [
SystemMessage(
content=SYSTEM_PROMPT.format(
now=MOCK_NOW.isoformat(),
incident_description=state["incident_description"],
)
),
*state["messages"],
]
response = model_with_tools.invoke(prompt)
if not isinstance(response, AIMessage):
raise TypeError("incident model must return an AIMessage")
# 一則 AIMessage 可能同時要求多個平行工具,因此計算實際要求筆數,不能只固定加一。
requested_calls = len(response.tool_calls)
if requested_calls:
tool_names = "、".join(
tool_call["name"] for tool_call in response.tool_calls
)
# 工具尚未真正執行,先將本輪要求的工具數累加回 State,供接下來的 Router 比對上限。
return {
"messages": [response],
"tool_call_count": state["tool_call_count"] + requested_calls,
"status": "investigating",
"events": [f"Agent 要求呼叫:{tool_names}。"],
}
# 沒有 tool call 代表模型要求停止循環,先保存回覆,後續由 Router 導向結案節點驗收。
return {"messages": [response]}
此時工具尚未執行,只是把最新累計量寫入 State。緊接著由條件邊的路由函式 route_after_investigation 讀取 State 中的控制變數進行分流:
def route_after_investigation(state: State) -> Route:
last_message = state["messages"][-1]
if not isinstance(last_message, AIMessage):
raise TypeError("route_after_investigation requires an AIMessage")
if (
last_message.tool_calls
and state["tool_call_count"] <= state["max_tool_calls"]
):
return "tools"
return "finish"
這個 Router 取代了原本的 tools_condition,落實了次數上限檢查:
"tools" 讓工具節點執行。"finish" 脫離迴圈。若模型在最後一輪要求的工具導致總數超過 max_tool_calls,Router 會直接攔截並轉向結案節點。這批超額的工具在執行前就會被擋下,不會送入 tools 節點執行,避免了浪費執行時間與外部系統負擔。
脫離迴圈後,流程進入 finish_investigation 節點。在結帳逾時的情境中,排查原則要求指控相依服務前必須查證其狀態。結案節點不聽信模型的片面宣稱,直接檢查 State 中的客觀執行事實:
def finish_investigation(state: State) -> StateUpdate:
last_message = state["messages"][-1]
if not isinstance(last_message, AIMessage):
raise TypeError("finish_investigation requires an AIMessage")
# 若最後一則訊息仍包含 tool call,代表是因為超過上限被強制中斷,最後一批工具未執行。
if last_message.tool_calls:
reason = (
f"Agent 要求的工具呼叫已超過上限 "
f"{state['max_tool_calls']},未執行最後一批工具。"
)
return {
"status": "incomplete",
"final_report": None,
"stop_reason": reason,
"events": [f"調查未完成:{reason}"],
}
# 從歷史訊息中找出所有實際執行過並回傳結果的 Tool 名稱。
observed_tools = {
message.name
for message in state["messages"]
if isinstance(message, ToolMessage)
}
# 檢查是否同時具備日誌搜尋與相依服務健康度的 ToolMessage。
if not {"search_logs", "check_service_health"} <= observed_tools:
reason = "Agent 未同時取得日誌與健康度檢查結果,不能完成調查。"
return {
"status": "incomplete",
"final_report": None,
"stop_reason": reason,
"events": [f"調查未完成:{reason}"],
}
# 滿足所有必要證據,接受模型最後的分析報告並標記為完成。
return {
"status": "completed",
"final_report": last_message.text,
"stop_reason": None,
"events": ["必要證據齊全,完成調查報告。"],
}
這個節點完成了三種狀態判定:
tool_calls,表示工具在執行前被攔截。State 的 status 設為 incomplete,final_report 設為 None,並在 stop_reason 記錄超額。observed_tools 缺少 search_logs 或 check_service_health 的 ToolMessage。即使模型文字寫得再詳盡,State 依然標記為 incomplete 並記錄缺少必要證據。final_report,並將 status 更新為 completed。這確保了外部呼叫端拿到的 status 是經由程式驗收的業務事實,而不是模型的片面宣稱。
這裡寫死檢查日誌與健康度兩項工具,是針對本範例的情境(顧客結帳逾時,查日誌發現疑似是後端的 Redis 連線失敗)所設定的結案標準。在真實事故中,模型最常見的早退行為是:只在日誌看見一行連線錯誤,就立刻腦補推論下游服務掛了,根本不願進一步查證即時健康度。因此範例以此作為防線。
在更多元的業務情境中,結案檢查不一定要寫死單一工具組合,而能採用更具彈性的驗收方式:
ToolMessage 中,避免模型在未經查驗的情況下憑空捏造數據。最後,在 workflow.py 中,節點與邊的組裝方式如下:
builder = StateGraph(State)
builder.add_node("investigator_agent", investigate)
builder.add_node("tools", ToolNode(INCIDENT_TOOLS))
builder.add_node("finish_investigation", finish_investigation)
builder.add_edge(START, "investigator_agent")
builder.add_conditional_edges(
"investigator_agent",
route_after_investigation,
{
"tools": "tools",
"finish": "finish_investigation",
},
)
builder.add_edge("tools", "investigator_agent")
builder.add_edge("finish_investigation", END)
這段組裝有兩個關鍵結構:
tools -> investigator_agent 形成推論迴圈:工具執行完畢後產生 ToolMessage,這條回邊將新的 Observation 帶回模型節點,驅動模型進行下一輪決策。finish_investigation 是唯一通往 END 的閘門:investigator_agent 不直接連向 END。只要脫離迴圈,無論原因為何,流程都必須通過 finish_investigation。這保證 Graph 結束時,State 內一定具備確定性的 status 與對應的報告或原因。接著,在 demo.py 中執行 Graph 時,傳入了 recursion_limit 設定:
result = graph.invoke(
initial_state,
{"recursion_limit": 20},
)
這兩個上限位在不同的層級,具備不同的保護責任:
max_tool_calls(業務邏輯上限):保存在 State 中,由業務程式精確累計模型請求的工具呼叫數。當達到上限時,流程會受控地轉向 finish_investigation,產出 status="incomplete" 的結構化結果。呼叫端能正常讀取失敗原因,決定是否指派工程師接手或重新調查。recursion_limit(底層執行步數上限):由 LangGraph 計算節點轉換的總步數。一次工具調查循環至少包含 investigator_agent 與 tools 兩個節點,因此 6 次工具呼叫可能累積十幾個步驟。recursion_limit 是防止工程師在 Router 判斷式中寫出邏輯漏洞(例如計數未更新導致死循環)時的最後保險絲;一旦觸發會直接拋出 GraphRecursionError 例外中斷程式。兩者不能互相取代。工具呼叫上限交由 State 與 Router 正常收尾,底層無窮迴圈的例外防護則由 recursion_limit 作為最後一道防線。
最後,在目前的設計中,我們為 Agent 建立了次數限制與結果驗收兩道邊界。但這兩道防線能放心讓流程全自動運轉,有一個重要的前提:流程中註冊的工具(search_logs 與 check_service_health)全部屬於唯讀查詢工具。
唯讀工具即使多查了幾次,頂多耗費額度與時間,不會對線上系統造成副作用。但在真實的事故處理中,排查出故障根因往往只是前半段,下一步必然會涉及修復與處置,例如重啟服務及調整設定。這類寫入與變更操作具備不可逆性與破壞性。如果只靠次數上限就讓模型在迴圈中自主執行處置工具,一旦模型在參數中填錯了服務名稱,或是誤殺了正在進行交易的正常連線,就會在線上引發二次災難。
面對高風險的寫入操作,程式不能直接放行,而必須在工具執行前暫停迴圈、保存目前狀態,並將處置參數交由線上工程師審查確認。下一篇將探討如何引入中斷機制與持久化儲存,實作由真人審批把關的處置流程。