昨天,把圖片多模態給結束了,今天,要開啟一個同樣蠻複雜的項目 - subagent 子代理!
這邊我們會定義好要如何進行,並把子 agent 的呼叫部分大致搞定;而明天,首先來做子 agent 的進度渲染(進行到哪了),然後把整個呼叫與回報流程串接完整。
也叫做子代理,就是讓主要的模型去呼叫底下的子模型,達到分工的效果。
那為什麼要這麼分工呢?
這邊最主要的原因又是上下文佔用,假設現在你讓 agent 幫你檢查整個專案中的渲染問題,那它要先調用工具查看你的專案,定位到有關渲染的檔案,然後才能閱讀此檔案並開始解決問題。
這過程中,來來回回可能需要調用個 3 到 5 次工具來查看文件,但是,這之中一定會查看到一些無關緊要的地方,而它們也全被模型一字不漏地記下來。
那現在有子 agent 可以分擔的話,情況就不一樣了!主 agent 調用了子 agent,給它「在專案中查找出有關於渲染部分的檔案」這項任務,定位到了檔案,主 agent 就省去了在龐大專案中漫無目的試錯與讀取不相干檔案的過程,成功節省了主 agent 的上下文空間。
在這裡,我把調用子 agent 的方式做成了「工具」,讓模型以調用工具的方式來調用子 agent。
而我們針對的是「單任務形式」的子 agent,簡單來說就是只讓它執行單次指派的任務(在背景自主查完資料、產出回答後即關閉),並且不保留對話歷史(所以能節省主 agent 的上下文)。
有一點我在開發時覺得蠻麻煩的,也是需要特別注意的地方,
還記得之前在做工具調用時,有特別設計了「多工具並發執行」嗎?
沒錯,也就是說,主 agent 是可以同時調用多個子 agent「一起來執行」的,而渲染時就需要特別隔開它們每一個的狀態,所以看似簡單的渲染邏輯硬是多了一大截。
這裡,簡單定義了以下兩種角色,並給了規則、限定工具。
限定工具可以避免不必要的工具佔用上下文。
# src/meowgent/tool.py
...
from ...
ROLE_PRESETS = {
"explorer": {
"rule": "你是一個專業的程式碼檢索專員。請在工作目錄中快速查找程式碼與檔案,並提供最簡明扼要的摘要結論。嚴禁修改任何檔案。",
"tools": ["read_file", "list_file", "grep_search", "web_fetch"], # 唯讀工具
},
"reviewer": {
"rule": "你是一個嚴謹的代碼審查專員(Code Reviewer)。請仔細閱讀給定的程式碼檔案,指出架構設計問題、潛在 Bug 或可改進之處。",
"tools": ["read_file"], # 只需要讀檔
}
}
...
首先,先建立工具函數,參數部分,task 是主 agent 要指派的任務,role 就是上面的角色:
# src/meowgent/tool.py
...
@tool_register(True)
def subagent_once(
task: Annotated[str, "要交付的任務說明"],
role: Annotated[
Literal["explorer", "reviewer"],
"子 Agent 的角色:'explorer'(唯讀快速檢索代碼)、'reviewer'(審查代碼與抓 Bug)"
]
):
from agent import Agent
from providers import OllamaProvider
先回去看一下 main.py,我們在呼叫模型前做了什麼事情:
而這三件事情就是這邊我們也要來做的,一個一個來吧!
這邊,我特別把子模型的部分獨立出來設定:
就不特別解釋語法部分了,忘記的話回去看 Day 17。
# src/meowgent/config/config_schema.py
...
class ModelsConfig(...):
...
class SubagentConfig(BaseModel):
""" 子模型相關 """
sub_agent_model: str = Field(
default="qwen3-vl:4b-thinking",
description="子模型"
)
sub_temperature: float = Field(
default=0.1,
ge=0, le=1.5,
description="溫度係數(0~1.5)"
)
class AgentConfig(...):
...
class MeowgentConfig(...):
models: ...
sub_agent: SubagentConfig = Field(default_factory=SubagentConfig)
agent: ...
這邊講一下獨立出來的原因,主 agent 和子 agent 在使用時,常會讓主 agent 使用較強大的模型,子 agent 用較快速的模型,來做到分工。
主 agent 做複雜思考,子 agent 快速給出工具調用結果。
這也就是為什麼我要特別把子模型的設定獨立出一個專屬的區塊。
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
_, config = ConfigManager.load_config()
sub_model_name = config.sub_agent.sub_agent_model
這裡,把剛才在 ROLE_PRESETS 設定的東西取出:
注意看
ROLE_PRESETS的結構:ROLE_PRESETS = { "explorer": { "rule": "你是一個專業的程式碼檢索專員。請在工作目錄中快速查找程式碼與檔案,並提供最簡明扼要的摘要結論。嚴禁修改任何檔案。", "tools": ["read_file", "list_file", "grep_search", "web_fetch"], # 唯讀工具 },...
explorer是ROLE_PRESETS字典的鍵,而值又是一個字典,
裡面的鍵分別為rule、tools,而值為字串及串列。所以,下面的語法是先取
ROLE_PRESETS的鍵(傳入subagent_once的角色參數,例如explorer),然後再取rule、tools兩個鍵得到要的內容。
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
rule = ROLE_PRESETS[role]["rule"]
tool_list = ROLE_PRESETS[role]["tools"]
這部分,要從 get_all_xml() 開始改:
在參數部分加上工具列表 tool_list,然後判斷如果有傳入工具列表,就按照列表上給工具。
新增參數的部分,我們都要加上預設值,讓之前寫好的調用不用重寫一次。
# src/meowgent/prompt/mcp_schema.py
def get_all_xml(tool_list: Optional[list] = None) -> str:
global ...
if tool_list: # 有指定工具
tool_nodes = []
for tool_name, tool_func in TOOL_REGISTRY.items():
if tool_name in tool_list:
tool_nodes.append(_tool_to_mcp_xml(tool_func))
return "<tools>\n" + "\n".join(tool_nodes) + "\n</tools>"
if not _MCP_TOOL_CACHE:
接下來是 get_system_prompt():當函式接收到 tool_list 時,就將它轉交給 get_all_xml() 來取得篩選後的工具列表:
# src/meowgent/prompt/system_prompt.py
...
def get_system_prompt(..., tool_list: Optional[list] = None) -> ...:
...
if enable_tools:
mcp_tool = get_all_xml(tool_list)
Agent() 裡以下部分也要改一下,建立物件時加入 tool_list 參數,並在調用 get_system_prompt() 時將其傳入:
# src/meowgent/agent.py
...
class Agent():
def __init__(..., tool_list: Optional[list] = None):
...
self.tool_list = tool_list
...
def chat(...) -> ...:
...
temp_history_messages = [
{
"role": "system",
"content": get_system_prompt(
...,
tool_list=self.tool_list
)
}
] + ...
system prompt 的部分,剩下的等等用
renew_system_prompt()來改,這邊先處理好建立 Agent 物件要用的工具。
下面這部分,就沒什麼好講的了,跟 main.py 中不同的地方是 tool_approval_mode 因為這邊給子 agent 的工具都是「唯讀」類型的,所以工具一律放行,就不特別做工具同意的審核了。
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
sub_model = Agent(
OllamaProvider(
model_name=sub_model_name,
temperature=config.sub_agent.sub_temperature,
),
max_turns=config.agent.max_turns,
tool_approval_mode="approval_all",
tool_list=tool_list
)
這裡就依照取得的值傳入 renew_system_prompt() 就好:
工作目錄的部分,因為在
main.py或commands.py中選完路徑都有做os.chdir(path)這個動作(進入資料夾),所以用cwd()就可以取得。
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
sub_model.renew_system_prompt(
rule=rule,
model_name=sub_model_name,
path=str(Path.cwd())
)
這裡為了分辨出多個子 agent,所以我們把每個加上一組數字 id 作為辨識:itertools.count() 是一個可以用作計數器的函數,傳入起點(這裡是 1),它會產生一個迭代器物件,然後就可以用 next() 取用。
# src/meowgent/tool.py
from ...
import itertools
_subagent_counter = itertools.count(1)
...
@...
def subagent_once(...) -> ...:
...
subagent_id = f"{role}#{next(_subagent_counter)}"
# 子 agent id,用於辨識身份
因為 chat() 會以串流的方式「流式」輸出文字,所以要在迴圈外部建立一個變數做收集:
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
final_report = "" # 子 agent 串流收集出的回答
然後就可以開始做回答的串流接收了,這邊一樣回去看一下 main.py 怎麼做的:
# src/meowgent/cli/main.py
...
if __name__ == "__main__":
...
...
...
for stream_content in model.chat(
user_input=user_input,
tool_approval=ask_tool_approval,
images=images
# 這邊忽略 -> 因為沒要給子 agent 讀圖片
):
...
tool_approval 要求回傳 Callable[[str, dict], bool] 這樣的函數進去,但為了省事,這邊直接用 lambda 做一個回傳 True 的函數出來:
# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
...
for stream_content in sub_model.chat(
user_input=task,
tool_approval=lambda *args: True
# *args 接收 t.tool_name、t.args,直接回傳 True
):
好啦!今天的部分先到這裡了~明天,繼續完成這個功能!
(小抱怨一下,最近庫存壓力很大啊哈哈!昨天的寫到了晚上 11:50 幾才發,今天也是,都是壓著線完成的啊。
庫存從最一開始的大約 15 篇被我揮霍到完全沒有了,怎麼會這樣啊T_T)