在 Agent 架構中,Reflection(自我反思/自我檢查) 是讓 Agent 從「瞎蒙硬猜」跨越到「高可靠度」的關鍵機制。
如果說 Planning 是 Agent 執行前的「路線規劃」,那麼 Reflection 就是執行中與執行後的「品質品管(QA)與自我除錯(Self-Debugging)」。它讓 Agent 能在不依賴人類干預的情況下,發現自身的錯誤、評價輸出的質量,並自我迭代修正。
在工程落地時,Reflection 通常分為以下三種常見模式:
1. Input Guard / Self-Consistency ──► [執行前/生成中檢查] (預防錯誤)
2. Reflexion (Post-Execution) ──► [執行後反思] (讀取 Error Log 並修正 Action)
3. Evaluator-Optimizer Pattern ──► [雙 Agent 審查] (Generator 生成 -> Critic 批改)
SyntaxError 或 NullPointerExceptions,將 Error Stream 餵給 Reflection Prompt,生成「失敗原因分析」並更新下一次的行動策略。這個範例展示一個 Generator + Evaluator (Critic) 的 Reflection 迴圈。當 Generator 生成的 C++ 代碼或 SQL 無法通過檢驗時,Evaluator 會指出漏洞,促使 Generator 自我修正,直到達到品質標準。
import json
from typing import List, Optional
from pydantic import BaseModel, Field
from openai import OpenAI
client = OpenAI()
# ==========================================
# 1. 定義 Evaluator 的評估架構 Schema
# ==========================================
class CodeEvaluation(BaseModel):
is_passed: bool = Field(description="程式碼是否完全符合需求且無邏輯/記憶體/語法漏洞")
score: int = Field(description="程式碼品質評分 (0 - 100)")
critique: str = Field(description="具體的批判與改進建議。若通過則填寫優點,若未通過請列出明確 Bug 或不符之處。")
suggested_fix: Optional[str] = Field(None, description="針對發現的漏洞,提供建議的修正方向")
# ==========================================
# 2. Generator & Evaluator 模組實作
# ==========================================
class ReflectionAgent:
def __init__(self, model: str = "gpt-4o-mini"):
self.model = model
def _generate_code(self, task: str, feedback_history: List[str] = None) -> str:
"""Generator:負責撰寫/修正程式碼"""
prompt = f"任務:{task}\n"
if feedback_history:
prompt += "\n之前嘗試失敗的反思與批判建議紀錄:\n"
for idx, fb in enumerate(feedback_history, 1):
prompt += f"--- 嘗試 {idx} 的反饋 ---\n{fb}\n"
prompt += "\n請根據上述反思建議,重新寫出修正後更高品質的程式碼。"
response = client.chat.completions.create(
model=self.model,
messages=[
{"role": "system", "content": "你是一個資深的進階 C++ 軟體工程師,請只輸出高質量的 C++ 程式碼。"},
{"role": "user", "content": prompt}
]
)
return response.choices[0].message.content
def _evaluate_code(self, task: str, code: str) -> CodeEvaluation:
"""Evaluator:扮演嚴格的 Code Reviewer,檢查潛在漏洞"""
prompt = (
f"原始任務需求:【{task}】\n\n"
f"待審查的 C++ 程式碼:\n```cpp\n{code}\n```\n\n"
f"請扮演極度嚴苛的 Code Reviewer,審查這段程式碼:\n"
f"1. 邏輯是否正確?\n"
f"2. 是否存在邊界條件問題 (Edge Cases) 或記憶體洩漏 (Memory Leaks)?\n"
f"3. 是否符合Modern C++ Best Practices (如 RAII)?\n"
f"評估後請給出分數與具體審查意見。"
)
completion = client.beta.chat.completions.parse(
model=self.model,
messages=[{"role": "user", "content": prompt}],
response_format=CodeEvaluation,
)
return completion.choices[0].message.parsed
def run(self, task: str, max_reflections: int = 3) -> str:
print(f"🎯 [Task]: {task}\n")
feedback_history = []
current_code = ""
for attempt in range(1, max_reflections + 1):
print(f"🔄 === 第 {attempt} 次嘗試 (Attempt {attempt}) ===")
# 1. 生成 (Generate)
current_code = self._generate_code(task, feedback_history)
print("📝 [Generated Code Draft]:")
print(current_code)
# 2. 檢查/反思 (Reflection / Evaluation)
eval_res = self._evaluate_code(task, current_code)
print(f"\n🧐 [Evaluator Score]: {eval_res.score} / 100")
print(f"🔍 [Critique]: {eval_res.critique}")
# 3. 判斷是否通過
if eval_res.is_passed or eval_res.score >= 90:
print("\n✅ [Reflection Passed!] 程式碼通過自我審查品質標準。")
return current_code
# 4. 記錄反思紀錄,進入下一輪迭代
print(f"💡 [Suggested Fix]: {eval_res.suggested_fix}")
print("⚠️ 未達標,觸發 Self-Correction 進行重新優化...\n")
feedback = (
f"得分: {eval_res.score}\n"
f"問題批判: {eval_res.critique}\n"
f"修復建議: {eval_res.suggested_fix}"
)
feedback_history.append(feedback)
print("⚠️ 已達最大反思輪次上限,輸出當前最佳版本。")
return current_code
# ==========================================
# 3. 測試執行
# ==========================================
if __name__ == "__main__":
agent = ReflectionAgent()
# 給出一個容易包含邊界條件與記憶體管理陷阱的需求
task_description = (
"請用純 Modern C++ 寫一個手動實現的單向鏈結串列 (Singly Linked List) 的 reverse() 函數,"
"並確保沒有記憶體洩漏,支援頭指標為 nullptr 的邊界情況。"
)
final_code = agent.run(task_description, max_reflections=3)
當 Prompt 傳入時,Generator 初稿可能寫出簡單的迭代轉置,但 Evaluator 會發揮 Reflection 的作用,抓出隱藏的瑕疵:
🎯 [Task]: 請用純 Modern C++ 寫一個手動實現的單向鏈結串列 (Singly Linked List) 的 reverse() 函數...
🔄 === 第 1 次嘗試 (Attempt 1) ===
📝 [Generated Code Draft]:
void reverse(Node*& head) {
Node* prev = nullptr;
Node* current = head;
while (current != nullptr) {
Node* next = current->next;
current->next = prev;
prev = current;
current = next;
}
head = prev;
}
🧐 [Evaluator Score]: 75 / 100
🔍 [Critique]: 程式碼完成了基本的鏈結串列反轉,但缺乏完整的結構定義與動態記憶體解構(Destructor)示範。沒有展現 Modern C++ (RAII) 對於鏈結串列節點釋放的處理,且未提供測試用例。
💡 [Suggested Fix]: 提供完整的 struct Node 與 LinkedList 類別,並用 RAII (或智能指標/明確解構子) 確保記憶體安全。
⚠️️ 未達標,觸發 Self-Correction 進行重新優化...
🔄 === 第 2 次嘗試 (Attempt 2) ===
... (Generator 讀取了上一次的 Critique,補齊了完整的類別設計、RAII 解構子與測試用例) ...
🧐 [Evaluator Score]: 95 / 100
🔍 [Critique]: 程式碼結構嚴密,包含完整的 LinkedList 類別、RAII 解構子以防記憶體洩漏,reverse() 函數正確處理了 head == nullptr 與單節點的邊界條件。
✅ [Reflection Passed!] 程式碼通過自我審查品質標準。
| 維度 | 無 Reflection | 具備 Reflection |
|---|---|---|
| 錯誤率 (Error Rate) | 高(一次性生成,幻覺或語法錯誤直接暴露給用戶) | 顯著降低(內部完成 2~3 輪除錯才輸出) |
| 複雜任務完成度 | 低(無法應對編譯錯誤或環境動態變化) | 高(結合 Code Execution 形成「寫 Code $\rightarrow$ 執行 $\rightarrow$ 看 Error Log $\rightarrow$ 修正」閉環) |
| 成本與延遲 | 低(單次 LLM 呼叫) | 較高(需額外 1~3 次 LLM 審查與修正呼叫) |
flake8)、C++ 編譯器或 Unit Test。用真實環境的 Error 訊息來引導 Reflection,比讓模型單純「空想反思」精準數倍。