Day 15 我們把 Agent 從 FakeLLMClient 接到 Gemini Flash。
接上真正的 LLM 後,evaluation result 開始變得更接近真實情境。
例如原本 fake client 可能只會回:
Fake response for: 請回答 HTTP 狀態碼 404 通常代表什麼
接上 Gemini 後,模型可能會正確回答:
HTTP 狀態碼 404 通常代表找不到請求的資源。
這讓 success rate 從 fake baseline 明顯提升。
但 Day 15 的結果也留下了一些很有價值的失敗案例。
例如 JSON 題的實際輸出可能是:
```json
{
"answer": 15
}
```
人類看得出來它其實是正確 JSON。
但對 json.loads() 來說,整段文字不是合法 JSON,因為外面多了 markdown code fence。
所以 evaluator 會判定:
Output is not valid JSON: Expecting value
到目前為止,平台只知道:
passed: false
failure_reason: Output is not valid JSON: Expecting value
這裡還可以再補一層。
我們要讓平台不只記錄失敗原因,也要記錄失敗類型。
這一篇要做到的是:
在 evaluation result 中加入
failure_type。
會完成幾件事:
evals/evaluators.py 的 EvaluationResult。failure_type。evals/runner.py,讓輸出的 JSON 包含 failure_type。先不碰這些:
範圍先收斂在一件事:
把失敗從一段文字,整理成可以統計的結構化類別。
failure_reason 適合給人看。
例如:
Expected output to contain '任務', but got 'AI Agent 是一種能夠自主感知環境...'
這句話很清楚,但不適合做統計。
如果我們想回答:
這次 eval run 裡,最多的是格式錯誤,還是答案錯誤?
只靠 failure_reason 會很麻煩。
因為每一筆 reason 都可能長得不一樣。
例如:
Expected exactly 'OK', but got 'OK。'
Output is not valid JSON: Expecting value
Missing required key: answer
Unknown tool: search
Agent execution error
這些字串適合除錯,但不適合分析。
所以我們需要一個更穩定的欄位:
{
"passed": false,
"failure_type": "format_error",
"failure_reason": "Output is not valid JSON: Expecting value"
}
failure_type 的用途是讓平台可以統計。
failure_reason 的用途是讓人可以追原因。
兩者不是互相取代,而是互補。
先定義六種失敗類型。
| failure_type | 說明 |
|---|---|
wrong_answer |
Agent 有回答,但答案不符合預期 |
format_error |
輸出格式錯誤,例如不是合法 JSON |
instruction_error |
沒有遵守明確指令,例如要求只回覆 OK |
tool_error |
工具呼叫失敗、工具不存在或工具輸入錯誤 |
execution_error |
Agent 執行過程發生 exception、timeout 或 API error |
unknown_failure |
暫時無法分類的錯誤 |
另外還有一種情況要特別說明:
evaluator_false_positive
這表示 evaluator 判定通過,但其實不該通過。
例如 Day 14 提過的案例:
問題:請回答 Python 中 list 是可變還是不可變資料型別
expected:可變
actual:Fake response for: 請回答 Python 中 list 是可變還是不可變資料型別
因為 actual 裡剛好包含「可變」,所以 contains 判定通過。
但這不是 Agent 真正答對,而是 evaluator 誤判。
這種情況比較難靠單純 rule-based 自動偵測。
所以今天先把它放在概念分類中,後面做 dashboard 和人工分析時再討論。
這次的實作重點先放在:
failed case -> failure_type
這次會修改兩個檔案。
agent-testing-platform/
evals/
__init__.py
cases.json
runner.py
evaluators.py
會修改:
| 檔案 | 修改內容 |
|---|---|
evals/evaluators.py |
讓 EvaluationResult 包含 failure_type,並加入分類邏輯 |
evals/runner.py |
把 failure_type 寫入每一筆 eval result |
不修改:
evals/cases.json
agents/simple_agent.py
agents/gemini_llm.py
agents/client_factory.py
因為這一篇要分類的是 evaluation result,不是改變 Agent 行為。
Day 12 建立的 EvaluationResult 目前長這樣:
@dataclass
class EvaluationResult:
passed: bool
failure_reason: str | None = None
這次要多加一個欄位:
failure_type
修改 evals/evaluators.py:
@dataclass
class EvaluationResult:
passed: bool
failure_reason: str | None = None
failure_type: str | None = None
failure_type 可以是 None。
通過的案例不需要失敗類型。
例如:
{
"passed": true,
"failure_type": null,
"failure_reason": null
}
而失敗的案例會有明確分類:
{
"passed": false,
"failure_type": "format_error",
"failure_reason": "Output is not valid JSON: Expecting value"
}
接著實作一個簡單的分類 function。
這個 function 會接收:
test_case
failure_reason
然後回傳對應的 failure_type。
修改 evals/evaluators.py,新增 classify_failure():
def classify_failure(test_case: dict, failure_reason: str | None) -> str:
grading_method = test_case["grading_method"]
task_type = test_case["task_type"]
reason = failure_reason or ""
if "not valid JSON" in reason:
return "format_error"
if "Missing required key" in reason:
return "format_error"
if "JSON output must be an object" in reason:
return "format_error"
if grading_method == "json_exact":
return "format_error"
if task_type == "instruction_following":
return "instruction_error"
if "Unknown tool" in reason:
return "tool_error"
if "Unsupported grading method" in reason:
return "unknown_failure"
if task_type in {"keyword_qa", "general_qa", "calculation"}:
return "wrong_answer"
return "unknown_failure"
這個 classifier 先採 rule-based。
它不是完美分類器,但很適合 MVP 階段。
例如只要 failure_reason 包含:
not valid JSON
就分類成:
format_error
如果任務類型是:
instruction_following
而且沒有通過,就分類成:
instruction_error
如果是一般問答、關鍵字問答或計算題沒有通過,先分類成:
wrong_answer
要注意的是:wrong_answer 不一定代表模型真的完全錯。
例如 Day 15 的 case_008:
請用一句話說明什麼是 AI Agent
Gemini 的回答語意上可能合理,但因為沒有包含 expected 裡指定的「任務」兩個字,所以被判定失敗。
這時候 wrong_answer 比較精確地說,是:
在目前 evaluator 規則下,答案不符合預期。
後面我們可以再細分成:
但這一篇先不要把分類做太細。
現在要讓每個 evaluator 在失敗時都帶上 failure_type。
一種做法是在每個 evaluator 裡自己判斷。
例如:
return EvaluationResult(
passed=False,
failure_reason="...",
failure_type="format_error",
)
但這樣會讓 evaluate_exact_match()、evaluate_contains()、evaluate_json_exact() 裡到處重複分類邏輯。
所以這裡採用另一種做法:
passed 和 failure_reason。evaluate() 最後補上 failure_type。修改 evals/evaluators.py 的 evaluate():
def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
expected = test_case["expected"]
grading_method = test_case["grading_method"]
if grading_method == "exact_match":
result = evaluate_exact_match(expected, actual)
elif grading_method == "contains":
result = evaluate_contains(expected, actual)
elif grading_method == "json_exact":
result = evaluate_json_exact(expected, actual)
else:
result = EvaluationResult(
passed=False,
failure_reason=f"Unsupported grading method: {grading_method}",
)
if not result.passed:
result.failure_type = classify_failure(
test_case=test_case,
failure_reason=result.failure_reason,
)
return result
這樣設計可以把分類邏輯集中在一個地方。
未來要調整 failure type 規則時,不需要去每個 evaluator 裡面找。
整理後,evals/evaluators.py 會像這樣。
修改 evals/evaluators.py:
import json
from dataclasses import dataclass
from typing import Any
@dataclass
class EvaluationResult:
passed: bool
failure_reason: str | None = None
failure_type: str | None = None
def evaluate_exact_match(expected: Any, actual: str | None) -> EvaluationResult:
if actual is None:
return EvaluationResult(
passed=False,
failure_reason="Actual output is None",
)
expected_text = str(expected).strip()
actual_text = actual.strip()
if actual_text == expected_text:
return EvaluationResult(passed=True)
return EvaluationResult(
passed=False,
failure_reason=f"Expected exactly '{expected_text}', but got '{actual_text}'",
)
def evaluate_contains(expected: Any, actual: str | None) -> EvaluationResult:
if actual is None:
return EvaluationResult(
passed=False,
failure_reason="Actual output is None",
)
expected_text = str(expected).strip()
if expected_text in actual:
return EvaluationResult(passed=True)
return EvaluationResult(
passed=False,
failure_reason=f"Expected output to contain '{expected_text}', but got '{actual}'",
)
def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
if actual is None:
return None, "Actual output is None"
try:
parsed = json.loads(actual)
except json.JSONDecodeError as exc:
return None, f"Output is not valid JSON: {exc.msg}"
if not isinstance(parsed, dict):
return None, "JSON output must be an object"
return parsed, None
def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
if not isinstance(expected, dict):
return EvaluationResult(
passed=False,
failure_reason="Expected value for json_exact must be an object",
)
parsed, error = parse_json_output(actual)
if error:
return EvaluationResult(
passed=False,
failure_reason=error,
)
for key, expected_value in expected.items():
if key not in parsed:
return EvaluationResult(
passed=False,
failure_reason=f"Missing required key: {key}",
)
actual_value = parsed[key]
if actual_value != expected_value:
return EvaluationResult(
passed=False,
failure_reason=(
f"Expected key '{key}' to be '{expected_value}', "
f"but got '{actual_value}'"
),
)
return EvaluationResult(passed=True)
def classify_failure(test_case: dict, failure_reason: str | None) -> str:
grading_method = test_case["grading_method"]
task_type = test_case["task_type"]
reason = failure_reason or ""
if "not valid JSON" in reason:
return "format_error"
if "Missing required key" in reason:
return "format_error"
if "JSON output must be an object" in reason:
return "format_error"
if grading_method == "json_exact":
return "format_error"
if task_type == "instruction_following":
return "instruction_error"
if "Unknown tool" in reason:
return "tool_error"
if "Unsupported grading method" in reason:
return "unknown_failure"
if task_type in {"keyword_qa", "general_qa", "calculation"}:
return "wrong_answer"
return "unknown_failure"
def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
expected = test_case["expected"]
grading_method = test_case["grading_method"]
if grading_method == "exact_match":
result = evaluate_exact_match(expected, actual)
elif grading_method == "contains":
result = evaluate_contains(expected, actual)
elif grading_method == "json_exact":
result = evaluate_json_exact(expected, actual)
else:
result = EvaluationResult(
passed=False,
failure_reason=f"Unsupported grading method: {grading_method}",
)
if not result.passed:
result.failure_type = classify_failure(
test_case=test_case,
failure_reason=result.failure_reason,
)
return result
這份完整版本保留了 Day 12 和 Day 13 的能力:
exact_match
contains
json_exact
只是會在既有結果上多加:
failure_type
接著修改 runner,讓輸出的 eval result 包含 failure_type。
修改 evals/runner.py,找到成功執行 Agent 並 append result 的地方。
原本大概長這樣:
results.append(
{
"case_id": test_case["id"],
"input": test_case["input"],
"expected": test_case["expected"],
"grading_method": test_case["grading_method"],
"task_type": test_case["task_type"],
"status": "completed",
"actual": result.answer,
"passed": evaluation.passed,
"failure_reason": evaluation.failure_reason,
"trace_session_id": result.trace.session_id,
"error": None,
}
)
修改 evals/runner.py,加入 failure_type:
results.append(
{
"case_id": test_case["id"],
"input": test_case["input"],
"expected": test_case["expected"],
"grading_method": test_case["grading_method"],
"task_type": test_case["task_type"],
"status": "completed",
"actual": result.answer,
"passed": evaluation.passed,
"failure_type": evaluation.failure_type,
"failure_reason": evaluation.failure_reason,
"trace_session_id": result.trace.session_id,
"error": None,
}
)
這裡多一個欄位:
"failure_type": evaluation.failure_type
如果測試通過,failure_type 會是:
null
如果測試失敗,則會是:
"format_error"
或:
"wrong_answer"
還有一個地方也要修改。
如果 Agent 執行過程發生 exception,程式會進入 except。
例如:
這些不是 evaluator 判斷出來的錯誤,而是 Agent 執行流程本身失敗。
所以會分類成:
execution_error
修改 evals/runner.py 的 except 區塊:
except Exception as exc:
results.append(
{
"case_id": test_case["id"],
"input": test_case["input"],
"expected": test_case["expected"],
"grading_method": test_case["grading_method"],
"task_type": test_case["task_type"],
"status": "error",
"actual": None,
"passed": False,
"failure_type": "execution_error",
"failure_reason": "Agent execution error",
"trace_session_id": None,
"error": str(exc),
}
)
這裡要分清楚 failure_reason 和 error 的差異。
| 欄位 | 用途 |
|---|---|
failure_reason |
給 evaluation 統計與報告使用 |
error |
保存實際 exception 訊息,方便除錯 |
例如:
{
"status": "error",
"passed": false,
"failure_type": "execution_error",
"failure_reason": "Agent execution error",
"error": "GEMINI_API_KEY is not set"
}
dashboard 可以把它統計成 execution_error,工程師也還是看得到真正的錯誤訊息。
Day 12 的 runner 已經會印出每題 PASS / FAIL。
可以順手把 failure type 也印出來。
修改 evals/runner.py 的 print_summary():
def print_summary(eval_run: dict[str, Any]) -> None:
total_cases = eval_run["total_cases"]
passed_count = sum(1 for result in eval_run["results"] if result["passed"])
failed_count = total_cases - passed_count
print(f"Run ID: {eval_run['run_id']}")
print(f"Total cases: {total_cases}")
print(f"Passed: {passed_count}")
print(f"Failed: {failed_count}")
print()
for result in eval_run["results"]:
label = "PASS" if result["passed"] else "FAIL"
print(f"[{label}] {result['case_id']} - {result['task_type']}")
if result["failure_type"]:
print(f" type: {result['failure_type']}")
if result["failure_reason"]:
print(f" reason: {result['failure_reason']}")
執行後,失敗案例會更容易閱讀。
例如:
[FAIL] case_013 - json_output
type: format_error
reason: Output is not valid JSON: Expecting value
確認你已經設定 Gemini API key:
export GEMINI_API_KEY="你的 Gemini API key"
export GEMINI_MODEL="gemini-3.6-flash"
接著執行:
LLM_PROVIDER=gemini python3 -m evals.runner
這次輸出的結果會多出 failure_type。
例如 JSON 題可能變成:
{
"case_id": "case_013",
"input": "請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer",
"expected": {
"answer": 15
},
"grading_method": "json_exact",
"task_type": "json_output",
"status": "completed",
"actual": "```json\n{\n \"answer\": 15\n}\n```",
"passed": false,
"failure_type": "format_error",
"failure_reason": "Output is not valid JSON: Expecting value",
"trace_session_id": "...",
"error": null
}
跟 Day 15 相比,這次多了一個可分析的欄位:
"failure_type": "format_error"
現在平台不只知道這題失敗,也知道它是格式錯誤。
如果用 Day 15 的 Gemini baseline 來看,失敗案例大致可以這樣分類。
| case_id | task_type | 原因 | failure_type |
|---|---|---|---|
case_008 |
general_qa |
回答語意合理,但沒有包含 expected keyword | wrong_answer |
case_009 |
general_qa |
使用「記錄」而不是 expected 的「紀錄」 | wrong_answer |
case_013 |
json_output |
外層包了 markdown code fence | format_error |
case_014 |
json_output |
外層包了 markdown code fence | format_error |
case_015 |
json_output |
外層包了 markdown code fence | format_error |
統計後可能會得到:
| failure_type | 數量 |
|---|---|
wrong_answer |
2 |
format_error |
3 |
這個結果比單純看:
Failed: 5
更有用。
因為它告訴我們後續改善方向:
format_error 可以用 schema instruction、JSON parsing cleanup、retry 或 guardrail 改善。wrong_answer 可以用 prompt 改善,也可以檢查 evaluator 是否太嚴格。Day 15 的 case_008 和 case_009 很值得討論。
例如 case_009:
input: 請用一句話說明什麼是 Trace
expected: 紀錄
actual: Trace 是指記錄並追蹤程式或請求在系統中執行的完整歷程...
這個答案其實有「記錄」,只是 expected 是「紀錄」。
在人類眼中,這很可能是可接受的答案。
但目前 evaluator 是:
contains
它只會做字串包含檢查。
於是「記錄」不等於「紀錄」,這題就會失敗。
這次先把它歸類成 wrong_answer,因為在目前規則下,actual 沒有符合 expected。
但這也提醒我們一件事:
Failure Analysis 不只是在分析 Agent,也是在檢查 evaluator 是否合理。
後面如果要改善,可以有幾種方向:
但這些都不是這一篇的範圍。
先把可統計的 failure_type 欄位建立起來。
執行完成後,可以到 data/eval_runs/ 找最新的結果檔。
例如:
data/eval_runs/eval_run_20260906_052327.json
使用以下指令格式化查看:
python3 -m json.tool data/eval_runs/eval_run_20260906_052327.json
你應該會看到每筆 result 多出:
"failure_type": null
或:
"failure_type": "format_error"
通過的案例會是:
{
"case_id": "case_011",
"passed": true,
"failure_type": null,
"failure_reason": null
}
失敗的案例會是:
{
"case_id": "case_013",
"passed": false,
"failure_type": "format_error",
"failure_reason": "Output is not valid JSON: Expecting value"
}
這次把 evaluation result 從:
{
"passed": false,
"failure_reason": "Output is not valid JSON: Expecting value"
}
擴充成:
{
"passed": false,
"failure_type": "format_error",
"failure_reason": "Output is not valid JSON: Expecting value"
}
完成的內容包含:
EvaluationResult 加入 failure_type。classify_failure()。evaluate() 統一補上 failure type。evals/runner.py 把 failure type 寫入 JSON。execution_error。做完後,平台可以回答的不只是:
哪幾題失敗?
而是:
這些失敗分別是哪一類?
這是後續 dashboard、prompt A/B testing、retry 與 guardrails 的基礎。
Day 17 會把這些 failure type 視覺化。
下一篇會建立錯誤分析表格與第一版 Failure Dashboard,讓我們可以直接看到:
到那時候,failure_type 就不只是 JSON 裡的一個欄位,而會變成分析 Agent 可靠性的第一個 dashboard 指標。