打造一個專為程式開發設計的 Coding Agent,不能只是簡單地把代碼範例丟給 LLM,而是要建立一個能夠 「閱讀專案脈絡(Context)$\rightarrow$ 規劃變更(Planning)$\rightarrow$ 實作與編輯(Code Editing)$\rightarrow$ 執行與測試(Execution)$\rightarrow$ 自我除錯(Reflection)」 的閉環系統(Closed-Loop System)。
這也是像 Cursor、Aider、Devin 或 Claude Code 這類現代 Coding Agent 能遠超傳統 Prompt 輸出的關鍵所在。
┌─────────────────────────┐
│ User Coding Task │
└────────────┬────────────┘
│
┌───────────────────────────────┼───────────────────────────────┐
▼ ▼ ▼
【1. Context Retrieval】 【2. Code Editing Strategy】 【3. Sandbox Execution】
(AST/Tree-sitter/RAG 檢索) (Aider Style Diff / Search-Replace) (Docker / WASM 隔離執行)
│ │ │
└───────────────────────────────┼───────────────────────────────┘
▼
【4. ReAct & Planning Engine】
(任務拆解、呼叫 Tool、檔名改動)
│
▼
【5. Reflection Loop】
(語法檢查 / Lint / 執行單元測試)
.h / .cpp / .py 檔案內容打包進 Context。make、pytest 或 g++,捕捉真實編譯器與 Runtime 的 stdout / stderr。ls)、搜尋符號(grep/ripgrep)、開啟檔案(read_file)與寫入變更(edit_file)。這份實作使用純 Python + OpenAI API,為 Agent 提供完整的檔案系統操作與 Python 測試沙盒工具,示範如何實現自動 Code Editing 與 Reflection。
import os
import subprocess
import json
from typing import List, Dict, Any
from pydantic import BaseModel, Field
from openai import OpenAI
client = OpenAI()
# ==========================================
# 1. Coding Agent 工具庫 (Tools)
# ==========================================
class FileSystemTools:
@staticmethod
def read_file(filepath: str) -> str:
"""讀取檔案內容"""
if not os.path.exists(filepath):
return f"Error: 檔案 '{filepath}' 不存在。"
with open(filepath, "r", encoding="utf-8") as f:
return f.read()
@staticmethod
def write_file(filepath: str, content: str) -> str:
"""覆寫或新建檔案"""
os.makedirs(os.path.dirname(filepath) or ".", exist_ok=True)
with open(filepath, "w", encoding="utf-8") as f:
f.write(content)
return f"Successfully wrote to '{filepath}'"
@staticmethod
def search_replace_edit(filepath: str, search_block: str, replace_block: str) -> str:
"""Aider 風格的精準代碼區塊替換 (Search & Replace)"""
if not os.path.exists(filepath):
return f"Error: 檔案 '{filepath}' 不存在。"
with open(filepath, "r", encoding="utf-8") as f:
content = f.read()
if search_block not in content:
return f"Error: 找不到待替換的原始程式碼區塊。\nSearched block:\n{search_block}"
new_content = content.replace(search_block, replace_block, 1)
with open(filepath, "w", encoding="utf-8") as f:
f.write(new_content)
return f"Successfully updated '{filepath}'"
@staticmethod
def run_tests(test_command: str = "python3 -m unittest discover") -> str:
"""沙盒執行器:執行單元測試並回傳完整 Log"""
try:
result = subprocess.run(
test_command,
shell=True,
capture_output=True,
text=True,
timeout=30
)
output = f"Return Code: {result.returncode}\n"
output += f"--- STDOUT ---\n{result.stdout}\n"
output += f"--- STDERR ---\n{result.stderr}\n"
return output
except Exception as e:
return f"Execution Error: {str(e)}"
# Native Tool Calling 宣告
TOOLS_SCHEMA = [
{
"type": "function",
"function": {
"name": "read_file",
"description": "讀取指定路徑的檔案內容",
"parameters": {
"type": "object",
"properties": {"filepath": {"type": "string"}},
"required": ["filepath"]
}
}
},
{
"type": "function",
"function": {
"name": "search_replace_edit",
"description": "使用 Search & Replace 區塊,精準修改檔案中的某段程式碼",
"parameters": {
"type": "object",
"properties": {
"filepath": {"type": "string"},
"search_block": {"type": "string", "description": "檔案中必須完全匹配的舊程式碼區塊"},
"replace_block": {"type": "string", "description": "準備替換進去的新程式碼區塊"}
},
"required": ["filepath", "search_block", "replace_block"]
}
}
},
{
"type": "function",
"function": {
"name": "run_tests",
"description": "執行測試指令以驗證程式碼正確性",
"parameters": {
"type": "object",
"properties": {"test_command": {"type": "string"}},
"required": ["test_command"]
}
}
}
]
# ==========================================
# 2. Coding Agent 引擎 (ReAct + Reflection)
# ==========================================
class CodingAgent:
def __init__(self, model: str = "gpt-4o-mini"):
self.model = model
self.tools = FileSystemTools()
def _dispatch_tool(self, tool_name: str, args: dict) -> str:
if tool_name == "read_file":
return self.tools.read_file(args["filepath"])
elif tool_name == "search_replace_edit":
return self.tools.search_replace_edit(
args["filepath"], args["search_block"], args["replace_block"]
)
elif tool_name == "run_tests":
return self.tools.run_tests(args.get("test_command", "python3 -m unittest"))
return f"Unknown tool: {tool_name}"
def run(self, task: str, max_turns: int = 10) -> str:
print(f"🚀 [Coding Task]: {task}\n")
system_prompt = (
"你是一個高效率且嚴謹的資深 Coding Agent。\n"
"你的目標是根據使用者的需求修改程式碼,並確保通過單元測試。\n\n"
"標準工作流程:\n"
"1. 讀取與分析相關的檔案與測試檔內容 (read_file)。\n"
"2. 使用 search_replace_edit 進行精準修改(避免複製貼上整份檔案)。\n"
"3. 呼叫 run_tests 執行測試。\n"
"4. 如果測試失敗 (Return Code != 0),分析 Error Log,進行 Reflection 並再次修改程式碼,直到所有測試通過!"
)
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": task}
]
for turn in range(1, max_turns + 1):
print(f"🔄 === Turn {turn} ===")
response = client.chat.completions.create(
model=self.model,
messages=messages,
tools=TOOLS_SCHEMA,
tool_choice="auto"
)
msg = response.choices[0].message
messages.append(msg)
# 情況 A:模型決定呼叫 Tool
if msg.tool_calls:
for tool_call in msg.tool_calls:
fn_name = tool_call.function.name
fn_args = json.loads(tool_call.function.arguments)
print(f"🛠️ [Tool Call]: {fn_name}({fn_args})")
# 執行工具
result = self._dispatch_tool(fn_name, fn_args)
print(f"📋 [Observation]:\n{result[:300]}...\n" if len(result) > 300 else f"📋 [Observation]:\n{result}\n")
# 將 Observation 寫回 Context
messages.append({
"role": "tool",
"tool_call_id": tool_call.id,
"content": result
})
# 情況 B:模型給出最終回答
else:
print(f"🤖 [Agent Final Reply]:\n{msg.content}\n")
return msg.content
return "到達最大輪次上限,任務終止。"
我們先在本地建立一個包含缺陷的 calculator.py 與對應的測試檔 test_calculator.py:
# 建立一個測試環境資料夾
os.makedirs("demo_workspace", exist_ok=True)
# 1. 寫入有 Bug 的邏輯檔案
with open("demo_workspace/calculator.py", "w") as f:
f.write("""
class Calculator:
def add(self, a, b):
return a + b
def divide(self, a, b):
# BUG: 沒有處理零除問題 (ZeroDivisionError)
return a / b
def power(self, base, exp):
# BUG: 符號寫錯,寫成 XOR 位元運算
return base ^ exp
""")
# 2. 寫入單元測試檔
with open("demo_workspace/test_calculator.py", "w") as f:
f.write("""
import unittest
from calculator import Calculator
class TestCalculator(unittest.TestCase):
def setUp(self):
self.calc = Calculator()
def test_divide_zero(self):
with self.assertRaises(ValueError):
self.calc.divide(10, 0)
def test_power(self):
self.assertEqual(self.calc.power(2, 3), 8)
if __name__ == '__main__':
unittest.main()
""")
if __name__ == "__main__":
# 切換到 demo 資料夾下執行
os.chdir("demo_workspace")
agent = CodingAgent()
task = (
"請幫我檢查並修復 calculator.py 中的問題。"
"目標是執行 `python3 -m unittest test_calculator.py` 並且所有測試必須 100% 通過。"
)
agent.run(task)
執行後,Console 會印出如下的執行軌跡:
🚀 [Coding Task]: 請幫我檢查並修復 calculator.py 中的問題...
🔄 === Turn 1 ===
🛠️ [Tool Call]: read_file({'filepath': 'calculator.py'})
📋 [Observation]: 讀取內容...
🛠️ [Tool Call]: read_file({'filepath': 'test_calculator.py'})
📋 [Observation]: 發現測試期待 divide(10, 0) 拋出 ValueError,且 power(2, 3) 應為 8。
🔄 === Turn 2 ===
🛠️ [Tool Call]: run_tests({'test_command': 'python3 -m unittest test_calculator.py'})
📋 [Observation]:
Return Code: 1
--- STDERR ---
FAIL: test_power (__main__.TestCalculator)
AssertionError: 1 != 8
ERROR: test_divide_zero (__main__.TestCalculator)
ZeroDivisionError: division by zero
🔄 === Turn 3 === (Reflection: Agent 分析了 Error Log,發現 divide 需要改拋出 ValueError,power 需改成 **)
🛠️ [Tool Call]: search_replace_edit({
'filepath': 'calculator.py',
'search_block': ' def divide(self, a, b):\n # BUG: 沒有處理零除問題 (ZeroDivisionError)\n return a / b',
'replace_block': ' def divide(self, a, b):\n if b == 0:\n raise ValueError("Cannot divide by zero")\n return a / b'
})
📋 [Observation]: Successfully updated 'calculator.py'
🛠️ [Tool Call]: search_replace_edit({
'filepath': 'calculator.py',
'search_block': ' def power(self, base, exp):\n # BUG: 符號寫錯,寫成 XOR 位元運算\n return base ^ exp',
'replace_block': ' def power(self, base, exp):\n return base ** exp'
})
📋 [Observation]: Successfully updated 'calculator.py'
🔄 === Turn 4 ===
🛠️ [Tool Call]: run_tests({'test_command': 'python3 -m unittest test_calculator.py'})
📋 [Observation]:
Return Code: 0
--- STDOUT ---
Ran 2 tests in 0.001s
OK
🔄 === Turn 5 ===
🤖 [Agent Final Reply]:
我已經成功修正了 `calculator.py` 中的兩個問題:
1. 在 `divide` 函數中增加了零除檢查,當 `b == 0` 時主動拋出 `ValueError`。
2. 將 `power` 函數的 XOR 運算子 `^` 修正為次方運算子 `**`。
經由單元測試驗證,所有測試已 100% 通過!
打造支援大型專案(十萬行程式碼等級)的 Coding Agent 時,需要導入以下關鍵技術:
| 技術維度 | 解決的工程痛點 | 實作方案 |
|---|---|---|
| Tree-sitter AST 解析 | 避免過度依賴大範圍文本搜尋或 Token 爆滿 | 將專案轉為 Code Property Graph (CPG),僅向 LLM 提供「該變更相關的 Class / Function 簽名檔」。 |
| Aider-style Search/Replace | 模型輸出完整的 2,000 行檔案既慢又容易出錯 | 約束模型只輸出包含上下文 3 行的 <<<<<<< SEARCH ... ======= ... >>>>>>> REPLACE 區塊。 |
| Git Integration & Rollback | 避免 Agent 把整個專案改壞無路可退 | Agent 在嘗試修正前自動執行 git commit;若連續 3 次 Reflection 依然測試失敗,自動發起 git reset --hard 回滾嘗試新策略。 |
| LSP (Language Server Protocol) | LLM 經常打錯函數名稱或參數名稱 | 將 Agent 接入 LSP,使 Agent 在發起修復後能直接獲得與 IDE 一致的語法紅線(Diagnostics)與型態檢查。 |