iT邦幫忙

2026 iThome 鐵人賽

DAY 20
0
自我挑戰組

AI Agent 從零開始系列 第 20 篇

讓 Agent 自己檢查自己(Reflection)

  • 分享至 

  • xImage
  •  

在 Agent 架構中,Reflection(自我反思/自我檢查) 是讓 Agent 從「瞎蒙硬猜」跨越到「高可靠度」的關鍵機制。

如果說 Planning 是 Agent 執行前的「路線規劃」,那麼 Reflection 就是執行中與執行後的「品質品管(QA)與自我除錯(Self-Debugging)」。它讓 Agent 能在不依賴人類干預的情況下,發現自身的錯誤、評價輸出的質量,並自我迭代修正。


一、 Reflection 的三種核心運作型態

在工程落地時,Reflection 通常分為以下三種常見模式:

1. Input Guard / Self-Consistency  ──► [執行前/生成中檢查] (預防錯誤)
2. Reflexion (Post-Execution)      ──► [執行後反思]      (讀取 Error Log 並修正 Action)
3. Evaluator-Optimizer Pattern     ──► [雙 Agent 審查]   (Generator 生成 -> Critic 批改)

  1. Self-Correction (單體即時反思): Agent 在生成程式碼或回答後,主動發起一次內部 Prompt:「檢查你剛才輸出的解答,有沒有邊界條件漏掉?邏輯有沒有矛盾?」
  2. Reflexion 模式 (對外環境反饋反思): Agent 將 Tool Call 發送出去(如:執行 Python 代碼),捕捉到編譯器回傳的 SyntaxError 或 NullPointerExceptions,將 Error Stream 餵給 Reflection Prompt,生成「失敗原因分析」並更新下一次的行動策略。
  3. Evaluator-Optimizer 模式 (審查者-優化者雙角色):
  • Generator:負責產出初稿(Code、Report、SQL Query)。
  • Evaluator:專職批判,依據既定標準(如:單元測試 Pass 率、語法規範、邏輯完整度)打分並提供具體修改建議(Feedback)。

二、 Python 實戰:建構一個具備 Reflection & Self-Correction 的代碼生成器

這個範例展示一個 Generator + Evaluator (Critic) 的 Reflection 迴圈。當 Generator 生成的 C++ 代碼或 SQL 無法通過檢驗時,Evaluator 會指出漏洞,促使 Generator 自我修正,直到達到品質標準。

import json
from typing import List, Optional
from pydantic import BaseModel, Field
from openai import OpenAI

client = OpenAI()

# ==========================================
# 1. 定義 Evaluator 的評估架構 Schema
# ==========================================
class CodeEvaluation(BaseModel):
    is_passed: bool = Field(description="程式碼是否完全符合需求且無邏輯/記憶體/語法漏洞")
    score: int = Field(description="程式碼品質評分 (0 - 100)")
    critique: str = Field(description="具體的批判與改進建議。若通過則填寫優點,若未通過請列出明確 Bug 或不符之處。")
    suggested_fix: Optional[str] = Field(None, description="針對發現的漏洞,提供建議的修正方向")


# ==========================================
# 2. Generator & Evaluator 模組實作
# ==========================================
class ReflectionAgent:
    def __init__(self, model: str = "gpt-4o-mini"):
        self.model = model

    def _generate_code(self, task: str, feedback_history: List[str] = None) -> str:
        """Generator:負責撰寫/修正程式碼"""
        prompt = f"任務:{task}\n"
        if feedback_history:
            prompt += "\n之前嘗試失敗的反思與批判建議紀錄:\n"
            for idx, fb in enumerate(feedback_history, 1):
                prompt += f"--- 嘗試 {idx} 的反饋 ---\n{fb}\n"
            prompt += "\n請根據上述反思建議,重新寫出修正後更高品質的程式碼。"

        response = client.chat.completions.create(
            model=self.model,
            messages=[
                {"role": "system", "content": "你是一個資深的進階 C++ 軟體工程師,請只輸出高質量的 C++ 程式碼。"},
                {"role": "user", "content": prompt}
            ]
        )
        return response.choices[0].message.content

    def _evaluate_code(self, task: str, code: str) -> CodeEvaluation:
        """Evaluator:扮演嚴格的 Code Reviewer,檢查潛在漏洞"""
        prompt = (
            f"原始任務需求:【{task}】\n\n"
            f"待審查的 C++ 程式碼:\n```cpp\n{code}\n```\n\n"
            f"請扮演極度嚴苛的 Code Reviewer,審查這段程式碼:\n"
            f"1. 邏輯是否正確?\n"
            f"2. 是否存在邊界條件問題 (Edge Cases) 或記憶體洩漏 (Memory Leaks)?\n"
            f"3. 是否符合Modern C++ Best Practices (如 RAII)?\n"
            f"評估後請給出分數與具體審查意見。"
        )
        completion = client.beta.chat.completions.parse(
            model=self.model,
            messages=[{"role": "user", "content": prompt}],
            response_format=CodeEvaluation,
        )
        return completion.choices[0].message.parsed

    def run(self, task: str, max_reflections: int = 3) -> str:
        print(f"🎯 [Task]: {task}\n")
        feedback_history = []
        current_code = ""

        for attempt in range(1, max_reflections + 1):
            print(f"🔄 === 第 {attempt} 次嘗試 (Attempt {attempt}) ===")
            
            # 1. 生成 (Generate)
            current_code = self._generate_code(task, feedback_history)
            print("📝 [Generated Code Draft]:")
            print(current_code)

            # 2. 檢查/反思 (Reflection / Evaluation)
            eval_res = self._evaluate_code(task, current_code)
            print(f"\n🧐 [Evaluator Score]: {eval_res.score} / 100")
            print(f"🔍 [Critique]: {eval_res.critique}")

            # 3. 判斷是否通過
            if eval_res.is_passed or eval_res.score >= 90:
                print("\n✅ [Reflection Passed!] 程式碼通過自我審查品質標準。")
                return current_code

            # 4. 記錄反思紀錄,進入下一輪迭代
            print(f"💡 [Suggested Fix]: {eval_res.suggested_fix}")
            print("⚠️ 未達標,觸發 Self-Correction 進行重新優化...\n")
            
            feedback = (
                f"得分: {eval_res.score}\n"
                f"問題批判: {eval_res.critique}\n"
                f"修復建議: {eval_res.suggested_fix}"
            )
            feedback_history.append(feedback)

        print("⚠️ 已達最大反思輪次上限,輸出當前最佳版本。")
        return current_code


# ==========================================
# 3. 測試執行
# ==========================================
if __name__ == "__main__":
    agent = ReflectionAgent()
    # 給出一個容易包含邊界條件與記憶體管理陷阱的需求
    task_description = (
        "請用純 Modern C++ 寫一個手動實現的單向鏈結串列 (Singly Linked List) 的 reverse() 函數,"
        "並確保沒有記憶體洩漏,支援頭指標為 nullptr 的邊界情況。"
    )
    final_code = agent.run(task_description, max_reflections=3)


三、 執行軌跡:Reflection 如何抓出 Bug

當 Prompt 傳入時,Generator 初稿可能寫出簡單的迭代轉置,但 Evaluator 會發揮 Reflection 的作用,抓出隱藏的瑕疵:

🎯 [Task]: 請用純 Modern C++ 寫一個手動實現的單向鏈結串列 (Singly Linked List) 的 reverse() 函數...

🔄 === 第 1 次嘗試 (Attempt 1) ===
📝 [Generated Code Draft]:
void reverse(Node*& head) {
    Node* prev = nullptr;
    Node* current = head;
    while (current != nullptr) {
        Node* next = current->next;
        current->next = prev;
        prev = current;
        current = next;
    }
    head = prev;
}

🧐 [Evaluator Score]: 75 / 100
🔍 [Critique]: 程式碼完成了基本的鏈結串列反轉,但缺乏完整的結構定義與動態記憶體解構(Destructor)示範。沒有展現 Modern C++ (RAII) 對於鏈結串列節點釋放的處理,且未提供測試用例。
💡 [Suggested Fix]: 提供完整的 struct Node 與 LinkedList 類別,並用 RAII (或智能指標/明確解構子) 確保記憶體安全。

⚠️️ 未達標,觸發 Self-Correction 進行重新優化...

🔄 === 第 2 次嘗試 (Attempt 2) ===
... (Generator 讀取了上一次的 Critique,補齊了完整的類別設計、RAII 解構子與測試用例) ...

🧐 [Evaluator Score]: 95 / 100
🔍 [Critique]: 程式碼結構嚴密,包含完整的 LinkedList 類別、RAII 解構子以防記憶體洩漏,reverse() 函數正確處理了 head == nullptr 與單節點的邊界條件。

✅ [Reflection Passed!] 程式碼通過自我審查品質標準。


四、 現代 Agent 架構中 Reflection 的核心價值

維度 無 Reflection 具備 Reflection
錯誤率 (Error Rate) 高(一次性生成,幻覺或語法錯誤直接暴露給用戶) 顯著降低(內部完成 2~3 輪除錯才輸出)
複雜任務完成度 低(無法應對編譯錯誤或環境動態變化) 高(結合 Code Execution 形成「寫 Code $\rightarrow$ 執行 $\rightarrow$ 看 Error Log $\rightarrow$ 修正」閉環)
成本與延遲 低(單次 LLM 呼叫) 較高(需額外 1~3 次 LLM 審查與修正呼叫)

五、 生產環境中的 Reflection 設計建議

  1. 不要為了 Reflection 而 Reflection:簡單任務(如提取 JSON 欄位)不需要開啟 Reflection,否則浪費 Token 與增加 latency。
  2. 結合 Deterministic Code(硬代碼檢查):讓 Reflection 結合真正的 Python linter (flake8)、C++ 編譯器或 Unit Test。用真實環境的 Error 訊息來引導 Reflection,比讓模型單純「空想反思」精準數倍。
  3. 專職化 Prompts: Generator 與 Evaluator 使用不同的 System Prompt。甚至可以給 Evaluator 設定偏執(Paranoid)的人設,專門刁鑽地尋找漏洞。

上一篇
Planning Agent
下一篇
打造 Coding Agent
系列文
AI Agent 從零開始 共 22 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言