iT邦幫忙

2026 iThome 鐵人賽

DAY 24
0

昨天,把圖片多模態給結束了,今天,要開啟一個同樣蠻複雜的項目 - subagent 子代理!
這邊我們會定義好要如何進行,並把子 agent 的呼叫部分大致搞定;而明天,首先來做子 agent 的進度渲染(進行到哪了),然後把整個呼叫與回報流程串接完整。


什麼是 subagent?

也叫做子代理,就是讓主要的模型去呼叫底下的子模型,達到分工的效果。

那為什麼要這麼分工呢?
這邊最主要的原因又是上下文佔用,假設現在你讓 agent 幫你檢查整個專案中的渲染問題,那它要先調用工具查看你的專案,定位到有關渲染的檔案,然後才能閱讀此檔案並開始解決問題。
這過程中,來來回回可能需要調用個 3 到 5 次工具來查看文件,但是,這之中一定會查看到一些無關緊要的地方,而它們也全被模型一字不漏地記下來。

那現在有子 agent 可以分擔的話,情況就不一樣了!主 agent 調用了子 agent,給它「在專案中查找出有關於渲染部分的檔案」這項任務,定位到了檔案,主 agent 就省去了在龐大專案中漫無目的試錯與讀取不相干檔案的過程,成功節省了主 agent 的上下文空間。


用什麼樣的形式?

在這裡,我把調用子 agent 的方式做成了「工具」,讓模型以調用工具的方式來調用子 agent。

而我們針對的是「單任務形式」的子 agent,簡單來說就是只讓它執行單次指派的任務(在背景自主查完資料、產出回答後即關閉),並且不保留對話歷史(所以能節省主 agent 的上下文)。


多個子 agent

有一點我在開發時覺得蠻麻煩的,也是需要特別注意的地方,
還記得之前在做工具調用時,有特別設計了「多工具並發執行」嗎?
沒錯,也就是說,主 agent 是可以同時調用多個子 agent「一起來執行」的,而渲染時就需要特別隔開它們每一個的狀態,所以看似簡單的渲染邏輯硬是多了一大截。


子 agent 角色

這裡,簡單定義了以下兩種角色,並給了規則、限定工具。

限定工具可以避免不必要的工具佔用上下文。

# src/meowgent/tool.py
...
from ...

ROLE_PRESETS = {
    "explorer": {
        "rule": "你是一個專業的程式碼檢索專員。請在工作目錄中快速查找程式碼與檔案,並提供最簡明扼要的摘要結論。嚴禁修改任何檔案。",
        "tools": ["read_file", "list_file", "grep_search", "web_fetch"],  # 唯讀工具
    },
    "reviewer": {
        "rule": "你是一個嚴謹的代碼審查專員(Code Reviewer)。請仔細閱讀給定的程式碼檔案,指出架構設計問題、潛在 Bug 或可改進之處。",
        "tools": ["read_file"],  # 只需要讀檔
    }
}
...

工具函數的初始化

首先,先建立工具函數,參數部分,task 是主 agent 要指派的任務,role 就是上面的角色:

# src/meowgent/tool.py
...
@tool_register(True)
def subagent_once(
    task: Annotated[str, "要交付的任務說明"],
    role: Annotated[
        Literal["explorer", "reviewer"],
        "子 Agent 的角色:'explorer'(唯讀快速檢索代碼)、'reviewer'(審查代碼與抓 Bug)"
    ]
):
	from agent import Agent
    from providers import OllamaProvider 

先回去看一下 main.py,我們在呼叫模型前做了什麼事情:

  1. 載入設定檔。
  2. 初始化 Agent 物件。
  3. 取得 system prompt。

而這三件事情就是這邊我們也要來做的,一個一個來吧!


設定檔

這邊,我特別把子模型的部分獨立出來設定:

就不特別解釋語法部分了,忘記的話回去看 Day 17。

# src/meowgent/config/config_schema.py
...
class ModelsConfig(...):
	...

class SubagentConfig(BaseModel):
    """ 子模型相關 """

    sub_agent_model: str = Field(
        default="qwen3-vl:4b-thinking",
        description="子模型"
    )

    sub_temperature: float = Field(
        default=0.1,
        ge=0, le=1.5,
        description="溫度係數(0~1.5)"
    )
    
class AgentConfig(...):
	...
    
class MeowgentConfig(...):
	models: ...

    sub_agent: SubagentConfig = Field(default_factory=SubagentConfig)

    agent: ...

這邊講一下獨立出來的原因,主 agent 和子 agent 在使用時,常會讓主 agent 使用較強大的模型,子 agent 用較快速的模型,來做到分工。

主 agent 做複雜思考,子 agent 快速給出工具調用結果。

這也就是為什麼我要特別把子模型的設定獨立出一個專屬的區塊。

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	_, config = ConfigManager.load_config()

    sub_model_name = config.sub_agent.sub_agent_model

取得角色行為及限定的工具

這裡,把剛才在 ROLE_PRESETS 設定的東西取出:

注意看 ROLE_PRESETS 的結構:

ROLE_PRESETS = {
    "explorer": {
        "rule": "你是一個專業的程式碼檢索專員。請在工作目錄中快速查找程式碼與檔案,並提供最簡明扼要的摘要結論。嚴禁修改任何檔案。",
        "tools": ["read_file", "list_file", "grep_search", "web_fetch"],  # 唯讀工具
    },...

explorer 是 ROLE_PRESETS 字典的鍵,而值又是一個字典,
裡面的鍵分別為 rule、tools,而值為字串及串列。

所以,下面的語法是先取 ROLE_PRESETS 的鍵(傳入 subagent_once 的角色參數,例如 explorer),然後再取 rule、tools 兩個鍵得到要的內容。

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	rule = ROLE_PRESETS[role]["rule"]

    tool_list = ROLE_PRESETS[role]["tools"]

建立 Agent 物件

這部分,要從 get_all_xml() 開始改:
在參數部分加上工具列表 tool_list,然後判斷如果有傳入工具列表,就按照列表上給工具。

新增參數的部分,我們都要加上預設值,讓之前寫好的調用不用重寫一次。

# src/meowgent/prompt/mcp_schema.py
def get_all_xml(tool_list: Optional[list] = None) -> str:
	global ...
	
	if tool_list: # 有指定工具

        tool_nodes = []
        for tool_name, tool_func in TOOL_REGISTRY.items():

            if tool_name in tool_list: 
                tool_nodes.append(_tool_to_mcp_xml(tool_func))

        return "<tools>\n" + "\n".join(tool_nodes) + "\n</tools>"
    
    if not _MCP_TOOL_CACHE:

接下來是 get_system_prompt():當函式接收到 tool_list 時,就將它轉交給 get_all_xml() 來取得篩選後的工具列表:

# src/meowgent/prompt/system_prompt.py
...
def get_system_prompt(..., tool_list: Optional[list] = None) -> ...:
	...
	
	if enable_tools:
        mcp_tool = get_all_xml(tool_list)

Agent() 裡以下部分也要改一下,建立物件時加入 tool_list 參數,並在調用 get_system_prompt() 時將其傳入:

# src/meowgent/agent.py
...
class Agent():
	def __init__(..., tool_list: Optional[list] = None):
		...
		self.tool_list = tool_list
		...
	def chat(...) -> ...:
			...
			temp_history_messages = [
                {
                    "role": "system",
                    "content": get_system_prompt(
                        ...,
                        tool_list=self.tool_list
                    )
                }
            ] + ...

system prompt 的部分,剩下的等等用 renew_system_prompt() 來改,這邊先處理好建立 Agent 物件要用的工具。

下面這部分,就沒什麼好講的了,跟 main.py 中不同的地方是 tool_approval_mode 因為這邊給子 agent 的工具都是「唯讀」類型的,所以工具一律放行,就不特別做工具同意的審核了。

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	sub_model = Agent(
        OllamaProvider(
            model_name=sub_model_name,
            temperature=config.sub_agent.sub_temperature,
        ),
        max_turns=config.agent.max_turns,
        tool_approval_mode="approval_all",
        tool_list=tool_list
    )

system prompt

這裡就依照取得的值傳入 renew_system_prompt() 就好:

工作目錄的部分,因為在 main.py 或 commands.py 中選完路徑都有做 os.chdir(path) 這個動作(進入資料夾),所以用 cwd() 就可以取得。

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	sub_model.renew_system_prompt(
        rule=rule,
        model_name=sub_model_name,
        path=str(Path.cwd())
    )

辨識身份

這裡為了分辨出多個子 agent,所以我們把每個加上一組數字 id 作為辨識:
itertools.count() 是一個可以用作計數器的函數,傳入起點(這裡是 1),它會產生一個迭代器物件,然後就可以用 next() 取用。

# src/meowgent/tool.py
from ...
import itertools

_subagent_counter = itertools.count(1)

...

@...
def subagent_once(...) -> ...:
	...
	
	subagent_id = f"{role}#{next(_subagent_counter)}"
	# 子 agent id,用於辨識身份

讓子 agent 開始運作

因為 chat() 會以串流的方式「流式」輸出文字,所以要在迴圈外部建立一個變數做收集:

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	final_report = "" # 子 agent 串流收集出的回答

然後就可以開始做回答的串流接收了,這邊一樣回去看一下 main.py 怎麼做的:

# src/meowgent/cli/main.py
...
if __name__ == "__main__":
	...
		...
			...
			
			for stream_content in model.chat(
				user_input=user_input,
				tool_approval=ask_tool_approval,
				images=images
				# 這邊忽略 -> 因為沒要給子 agent 讀圖片
			):
				...

tool_approval 要求回傳 Callable[[str, dict], bool] 這樣的函數進去,但為了省事,這邊直接用 lambda 做一個回傳 True 的函數出來:

# src/meowgent/tool.py
...
@...
def subagent_once(...) -> ...:
	...
	
	for stream_content in sub_model.chat(
        user_input=task,
        tool_approval=lambda *args: True
        # *args 接收 t.tool_name、t.args,直接回傳 True
    ):

好啦!今天的部分先到這裡了~明天,繼續完成這個功能!
(小抱怨一下,最近庫存壓力很大啊哈哈!昨天的寫到了晚上 11:50 幾才發,今天也是,都是壓著線完成的啊。
庫存從最一開始的大約 15 篇被我揮霍到完全沒有了,怎麼會這樣啊T_T)


上一篇
Day 23 - 圖片讀取的多模態能力 - 下
系列文
手刻 AI Agent!大一新生的 Python 實戰筆記 共 24 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言