iT邦幫忙

2026 iThome 鐵人賽

DAY 17
0

前言

Day 16 我們把 evaluation result 從單純的:

{
  "passed": false,
  "failure_reason": "Output is not valid JSON: Expecting value"
}

擴充成:

{
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value"
}

這個改動看起來只多了一個欄位,但它讓後續分析變得更容易。

因為現在我們不只能問:

哪些 cases 失敗?

也能開始問:

這些失敗主要是哪一類?
哪一種 task type 最容易失敗?
失敗案例能不能快速點回 trace?

這一篇要把 Day 16 的 failure_type 用起來。

我們會建立第一版 Failure Dashboard,讓 eval run 的結果不只存在 JSON 檔案裡,而是可以用表格和圖表觀察。


今天要完成什麼?

這一篇要做到的是:

建立一個 Streamlit Failure Dashboard,用來閱讀 eval run 中的失敗案例與錯誤分布。

會完成幾件事:

  1. 新增 ui/failure_dashboard.py。
  2. 新增 failure_dashboard_app.py 作為 Streamlit 入口。
  3. 讀取 data/eval_runs/ 中的 eval run JSON。
  4. 顯示整體成功率。
  5. 顯示 failed cases 表格。
  6. 顯示各 failure type 的數量。
  7. 顯示各 task type 的失敗率。
  8. 保留 trace_session_id,方便回到 Trace Viewer。
  9. 挑出代表性失敗案例,做後續改善方向判斷。

先不做:

  • 即時重新執行 evaluation。
  • 在 dashboard 裡直接打開 trace viewer。
  • 多個 eval runs 的趨勢比較。
  • prompt A/B testing。
  • retry 或自動修正。
  • LLM-as-a-Judge。

範圍先收斂成一件事:

把單次 eval run 的失敗結果整理成可以閱讀、可以比較、可以討論的畫面。


為什麼需要 Failure Dashboard?

目前 eval run 的結果會輸出成 JSON。

例如:

data/eval_runs/eval_run_20260906_052327.json

用 python3 -m json.tool 可以格式化查看:

python3 -m json.tool data/eval_runs/eval_run_20260906_052327.json

這種方式對工程師來說可以接受,但不適合長期分析。

因為當測試案例變多時,JSON 會很快變得難讀。

例如我們想回答:

這次 run 總共有幾題 format_error?
json_output 類型的失敗率是多少?
哪幾題雖然 completed,但 evaluator 判定 failed?

如果每次都靠肉眼看 JSON,很容易漏掉重點。

Failure Dashboard 的價值,是把 eval result 轉成幾個固定視角:

視角 回答的問題
Summary metrics 這次 run 整體表現如何?
Failed cases table 哪些案例失敗?實際輸出是什麼?
Failure type distribution 失敗主要集中在哪些類型?
Task type failure rate 哪些任務類型最容易失敗?
Trace session id 要深入除錯時,能不能回到 trace?

這些資訊會成為後面 prompt A/B testing 和 retry 實驗的基礎。


今天的專案結構

這次會新增兩個檔案。

agent-testing-platform/
  failure_dashboard_app.py
  app.py
  trace_viewer_app.py
  agents/
    __init__.py
    simple_agent.py
    fake_llm.py
    gemini_llm.py
    client_factory.py
  evals/
    __init__.py
    cases.json
    runner.py
    evaluators.py
  ui/
    __init__.py
    trace_viewer.py
    failure_dashboard.py
  data/
    eval_runs/
      eval_run_*.json

會新增:

檔案 用途
ui/failure_dashboard.py Failure Dashboard 的主要畫面與資料整理邏輯
failure_dashboard_app.py Streamlit app 入口

不修改:

  • evals/runner.py
  • evals/evaluators.py
  • evals/cases.json
  • agents/simple_agent.py

因為 Day 16 已經把 failure_type 寫進 eval result。

這次要做的是讀取和呈現,不是改變評測邏輯。


Dashboard 要讀取什麼資料?

Day 16 之後,一份 eval run JSON 大致會長這樣:

{
  "run_id": "eval_run_20260906_052327",
  "created_at": "2026-09-06T05:23:27",
  "llm_provider": "gemini",
  "total_cases": 15,
  "results": [
    {
      "case_id": "case_013",
      "input": "請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer",
      "expected": {
        "answer": 15
      },
      "grading_method": "json_exact",
      "task_type": "json_output",
      "status": "completed",
      "actual": "```json\n{\n  \"answer\": 15\n}\n```",
      "passed": false,
      "failure_type": "format_error",
      "failure_reason": "Output is not valid JSON: Expecting value",
      "trace_session_id": "...",
      "error": null
    }
  ]
}

Dashboard 會使用三個層次的資料。

第一層是 eval run metadata:

欄位 用途
run_id 顯示目前選到哪一次 eval run
created_at 顯示執行時間
llm_provider 顯示這次使用 fake 還是 Gemini
total_cases 計算成功率與失敗率

第二層是每一筆 result:

欄位 用途
case_id 辨識測試案例
task_type 統計不同任務類型的失敗率
input 顯示原始任務
expected 顯示預期結果
actual 顯示 Agent 實際輸出
passed 判斷通過或失敗
failure_type 統計失敗類型
failure_reason 顯示具體失敗原因
trace_session_id 回到 Trace Viewer 除錯

第三層是衍生統計:

指標 計算方式
success rate passed_count / total_cases
failed count total_cases - passed_count
failure type count 依 failure_type group by
task type failure rate 每個 task_type 的 failed count / total count

建立 Failure Dashboard 入口

先新增 Streamlit 入口檔。

新增 failure_dashboard_app.py:

from ui.failure_dashboard import render_failure_dashboard


if __name__ == "__main__":
    render_failure_dashboard()

這個檔案很薄,只負責呼叫畫面 function。

跟 Day 6 的 trace_viewer_app.py 一樣,真正的畫面邏輯放在 ui/ 裡。

這樣拆有幾個實際好處:

  • Streamlit 入口很乾淨。
  • UI 邏輯集中在 ui/failure_dashboard.py。
  • 未來如果要把多個 dashboard 合併,也比較容易重用。

讀取 eval run JSON

接著建立 dashboard 主檔。

新增 ui/failure_dashboard.py,先放入必要 import 與讀檔 function:

import json
from pathlib import Path
from typing import Any

import pandas as pd
import streamlit as st


EVAL_RUNS_DIR = Path("data/eval_runs")


def list_eval_run_files() -> list[Path]:
    if not EVAL_RUNS_DIR.exists():
        return []

    return sorted(
        EVAL_RUNS_DIR.glob("eval_run_*.json"),
        reverse=True,
    )


def load_eval_run(path: Path) -> dict[str, Any]:
    with path.open("r", encoding="utf-8") as file:
        return json.load(file)

這裡先處理兩件事。

list_eval_run_files() 會列出 data/eval_runs/ 底下所有 eval run JSON,並用檔名倒序排列。

因為我們的檔名包含時間:

eval_run_20260906_052327.json

倒序排列後,最新的 run 會出現在最上面。

load_eval_run() 則負責把 JSON 讀成 Python dict。


把 results 轉成 DataFrame

Streamlit 可以直接顯示 list of dict,但如果要做統計,使用 Pandas 會比較方便。

繼續修改 ui/failure_dashboard.py,新增 results_to_dataframe():

def results_to_dataframe(eval_run: dict[str, Any]) -> pd.DataFrame:
    results = eval_run.get("results", [])

    rows = []
    for result in results:
        rows.append(
            {
                "case_id": result.get("case_id"),
                "task_type": result.get("task_type"),
                "grading_method": result.get("grading_method"),
                "status": result.get("status"),
                "passed": result.get("passed", False),
                "failure_type": result.get("failure_type"),
                "failure_reason": result.get("failure_reason"),
                "input": result.get("input"),
                "expected": json.dumps(
                    result.get("expected"),
                    ensure_ascii=False,
                ),
                "actual": result.get("actual"),
                "trace_session_id": result.get("trace_session_id"),
                "error": result.get("error"),
            }
        )

    return pd.DataFrame(rows)

這段程式做了一個小整理。

expected 有時候是字串:

"找不到"

有時候是 dict:

{
  "answer": 15
}

如果直接丟給表格顯示,格式會比較不穩定。

這裡改用:

json.dumps(result.get("expected"), ensure_ascii=False)

把它統一轉成可閱讀的 JSON 字串。


顯示 Summary Metrics

Dashboard 第一區先顯示整體數字。

繼續修改 ui/failure_dashboard.py,新增 render_summary():

def render_summary(eval_run: dict[str, Any], results_df: pd.DataFrame) -> None:
    total_cases = len(results_df)
    passed_count = int(results_df["passed"].sum()) if total_cases else 0
    failed_count = total_cases - passed_count
    success_rate = passed_count / total_cases if total_cases else 0

    st.subheader("Summary")

    col1, col2, col3, col4 = st.columns(4)
    col1.metric("Total cases", total_cases)
    col2.metric("Passed", passed_count)
    col3.metric("Failed", failed_count)
    col4.metric("Success rate", f"{success_rate:.1%}")

    st.caption(
        f"Run ID: {eval_run.get('run_id', 'unknown')} | "
        f"Provider: {eval_run.get('llm_provider', 'unknown')} | "
        f"Created at: {eval_run.get('created_at', 'unknown')}"
    )

這裡用 st.metric() 顯示四個數字:

  • total cases
  • passed
  • failed
  • success rate

這是進 dashboard 後最先需要看到的資訊。

如果 success rate 變高或變低,我們再往下看失敗分布,判斷原因。


顯示 Failed Cases 表格

接著顯示失敗案例。

繼續修改 ui/failure_dashboard.py,新增 render_failed_cases():

def render_failed_cases(results_df: pd.DataFrame) -> None:
    st.subheader("Failed Cases")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.success("這次 eval run 沒有失敗案例。")
        return

    columns = [
        "case_id",
        "task_type",
        "failure_type",
        "failure_reason",
        "input",
        "expected",
        "actual",
        "trace_session_id",
    ]

    st.dataframe(
        failed_df[columns],
        use_container_width=True,
        hide_index=True,
    )

這張表格會是這版 dashboard 最常看的畫面。

它讓我們可以一次看到:

  • 哪個 case 失敗。
  • 屬於哪個 task type。
  • 被歸類成哪種 failure type。
  • evaluator 給出的 failure reason。
  • 原始 input。
  • expected 和 actual 的差異。
  • 對應的 trace session id。

這版先只把 trace_session_id 顯示成文字。

如果要深入看 trace,可以複製這個 id,回到 Day 6 做的 Trace Viewer 查詢。

後面如果要改善,可以把 Failure Dashboard 和 Trace Viewer 整合在同一個 Streamlit app。


注意 Pandas 的布林篩選寫法

上面這行:

failed_df = results_df[results_df["passed"] == False].copy()

在 Python 風格上,有些人會想寫成:

results_df[not results_df["passed"]]

但 Pandas 不能這樣用。

results_df["passed"] 是一整欄 Series,不是單一 bool。

要改用:

results_df["passed"] == False

或:

~results_df["passed"]

這裡用第一種寫法,對初學者比較直覺。


顯示 Failure Type 分布

接著來看失敗類型統計。

繼續修改 ui/failure_dashboard.py,新增 render_failure_type_chart():

def render_failure_type_chart(results_df: pd.DataFrame) -> None:
    st.subheader("Failure Type Distribution")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.info("沒有 failure type 可以統計。")
        return

    counts_df = (
        failed_df["failure_type"]
        .fillna("unknown_failure")
        .value_counts()
        .rename_axis("failure_type")
        .reset_index(name="count")
    )

    st.bar_chart(
        counts_df,
        x="failure_type",
        y="count",
    )

    st.dataframe(
        counts_df,
        use_container_width=True,
        hide_index=True,
    )

這段會把失敗案例依照 failure_type 統計。

例如 Day 16 的 Gemini baseline 可能得到:

failure_type count
format_error 3
wrong_answer 2

這個圖表可以幫助我們快速判斷:

下一步最值得改善的是什麼?

如果 format_error 最多,通常代表可以先改善輸出格式控制。

如果 wrong_answer 最多,可能要檢查 prompt、test case expected、或 evaluator 是否太嚴格。

如果 execution_error 最多,代表系統穩定性或 API 設定可能有問題,這時不應該先改 prompt。


顯示 Task Type 失敗率

只看 failure type 還不夠。

我們也需要知道哪一種任務類型最容易失敗。

例如:

json_output 是否特別容易失敗?
instruction_following 是否已經穩定?
calculation 是否仍然維持高成功率?

繼續修改 ui/failure_dashboard.py,新增 build_task_type_summary():

def build_task_type_summary(results_df: pd.DataFrame) -> pd.DataFrame:
    summary_df = (
        results_df.groupby("task_type")
        .agg(
            total=("case_id", "count"),
            passed=("passed", "sum"),
        )
        .reset_index()
    )

    summary_df["failed"] = summary_df["total"] - summary_df["passed"]
    summary_df["failure_rate"] = summary_df["failed"] / summary_df["total"]

    return summary_df.sort_values(
        by="failure_rate",
        ascending=False,
    )

這裡用 groupby("task_type") 將 results 依任務類型分組。

每一組統計:

  • total:該 task type 有幾題。
  • passed:通過幾題。
  • failed:失敗幾題。
  • failure_rate:失敗率。

接著新增顯示 function:

def render_task_type_summary(results_df: pd.DataFrame) -> None:
    st.subheader("Task Type Failure Rate")

    if results_df.empty:
        st.info("沒有結果可以統計。")
        return

    summary_df = build_task_type_summary(results_df)

    st.bar_chart(
        summary_df,
        x="task_type",
        y="failure_rate",
    )

    display_df = summary_df.copy()
    display_df["failure_rate"] = display_df["failure_rate"].map(
        lambda value: f"{value:.1%}"
    )

    st.dataframe(
        display_df,
        use_container_width=True,
        hide_index=True,
    )

這個區塊會讓我們看到不同任務類型的失敗率。

例如:

task_type total passed failed failure_rate
json_output 3 0 3 100.0%
general_qa 3 1 2 66.7%
calculation 4 4 0 0.0%

這比單純知道:

Failed: 5

更有行動價值。

因為它告訴我們問題集中在哪一種任務。


顯示代表性失敗案例

有了圖表後,還需要回到具體案例。

數字可以告訴我們「哪裡最多」,但改善 Agent 時,還是要看 actual output。

我們可以先做一個簡單版:每種 failure_type 顯示一筆代表案例。

繼續修改 ui/failure_dashboard.py,新增 render_representative_failures():

def render_representative_failures(results_df: pd.DataFrame) -> None:
    st.subheader("Representative Failures")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.info("沒有代表性失敗案例。")
        return

    for failure_type, group_df in failed_df.groupby("failure_type", dropna=False):
        example = group_df.iloc[0]
        label = failure_type or "unknown_failure"

        with st.expander(f"{label} - {example['case_id']}"):
            st.write("Input")
            st.code(example["input"] or "")

            st.write("Expected")
            st.code(example["expected"] or "")

            st.write("Actual")
            st.code(example["actual"] or "")

            st.write("Failure reason")
            st.code(example["failure_reason"] or "")

            st.write("Trace session id")
            st.code(example["trace_session_id"] or "")

這個區塊的用途是讓我們不用一開始就展開所有失敗案例。

例如 dashboard 可能會顯示:

format_error - case_013
wrong_answer - case_008

點開後再看 input、expected、actual 和 failure reason。

這種設計適合 MVP 階段。

等測試案例變多後,我們可以再加上:

  • 依 failure type 篩選。
  • 依 task type 篩選。
  • 搜尋 case id。
  • 顯示完整 trace steps。

組合完整 Dashboard

最後把前面的 function 串起來。

修改 ui/failure_dashboard.py,新增 render_failure_dashboard():

def render_failure_dashboard() -> None:
    st.set_page_config(
        page_title="Agent Failure Dashboard",
        layout="wide",
    )

    st.title("Agent Failure Dashboard")
    st.caption("分析 eval run 的失敗案例、錯誤類型與任務類型失敗率")

    eval_run_files = list_eval_run_files()

    if not eval_run_files:
        st.info("目前沒有 eval run。請先執行 python3 -m evals.runner。")
        return

    selected_file = st.selectbox(
        "選擇 eval run",
        options=eval_run_files,
        format_func=lambda path: path.name,
    )

    eval_run = load_eval_run(selected_file)
    results_df = results_to_dataframe(eval_run)

    render_summary(eval_run, results_df)
    st.divider()

    render_failed_cases(results_df)
    st.divider()

    col1, col2 = st.columns(2)

    with col1:
        render_failure_type_chart(results_df)

    with col2:
        render_task_type_summary(results_df)

    st.divider()
    render_representative_failures(results_df)

這裡的畫面順序是:

  1. 選擇 eval run。
  2. 顯示 summary。
  3. 顯示 failed cases 表格。
  4. 左右兩欄顯示 failure type 和 task type 統計。
  5. 顯示代表性失敗案例。

這樣安排是因為分析時通常會照這個順序閱讀:

整體表現
  -> 失敗清單
  -> 分布統計
  -> 具體案例

ui/failure_dashboard.py 完整版本

整理後,ui/failure_dashboard.py 會長這樣。

新增 ui/failure_dashboard.py:

import json
from pathlib import Path
from typing import Any

import pandas as pd
import streamlit as st


EVAL_RUNS_DIR = Path("data/eval_runs")


def list_eval_run_files() -> list[Path]:
    if not EVAL_RUNS_DIR.exists():
        return []

    return sorted(
        EVAL_RUNS_DIR.glob("eval_run_*.json"),
        reverse=True,
    )


def load_eval_run(path: Path) -> dict[str, Any]:
    with path.open("r", encoding="utf-8") as file:
        return json.load(file)


def results_to_dataframe(eval_run: dict[str, Any]) -> pd.DataFrame:
    results = eval_run.get("results", [])

    rows = []
    for result in results:
        rows.append(
            {
                "case_id": result.get("case_id"),
                "task_type": result.get("task_type"),
                "grading_method": result.get("grading_method"),
                "status": result.get("status"),
                "passed": result.get("passed", False),
                "failure_type": result.get("failure_type"),
                "failure_reason": result.get("failure_reason"),
                "input": result.get("input"),
                "expected": json.dumps(
                    result.get("expected"),
                    ensure_ascii=False,
                ),
                "actual": result.get("actual"),
                "trace_session_id": result.get("trace_session_id"),
                "error": result.get("error"),
            }
        )

    return pd.DataFrame(rows)


def render_summary(eval_run: dict[str, Any], results_df: pd.DataFrame) -> None:
    total_cases = len(results_df)
    passed_count = int(results_df["passed"].sum()) if total_cases else 0
    failed_count = total_cases - passed_count
    success_rate = passed_count / total_cases if total_cases else 0

    st.subheader("Summary")

    col1, col2, col3, col4 = st.columns(4)
    col1.metric("Total cases", total_cases)
    col2.metric("Passed", passed_count)
    col3.metric("Failed", failed_count)
    col4.metric("Success rate", f"{success_rate:.1%}")

    st.caption(
        f"Run ID: {eval_run.get('run_id', 'unknown')} | "
        f"Provider: {eval_run.get('llm_provider', 'unknown')} | "
        f"Created at: {eval_run.get('created_at', 'unknown')}"
    )


def render_failed_cases(results_df: pd.DataFrame) -> None:
    st.subheader("Failed Cases")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.success("這次 eval run 沒有失敗案例。")
        return

    columns = [
        "case_id",
        "task_type",
        "failure_type",
        "failure_reason",
        "input",
        "expected",
        "actual",
        "trace_session_id",
    ]

    st.dataframe(
        failed_df[columns],
        use_container_width=True,
        hide_index=True,
    )


def render_failure_type_chart(results_df: pd.DataFrame) -> None:
    st.subheader("Failure Type Distribution")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.info("沒有 failure type 可以統計。")
        return

    counts_df = (
        failed_df["failure_type"]
        .fillna("unknown_failure")
        .value_counts()
        .rename_axis("failure_type")
        .reset_index(name="count")
    )

    st.bar_chart(
        counts_df,
        x="failure_type",
        y="count",
    )

    st.dataframe(
        counts_df,
        use_container_width=True,
        hide_index=True,
    )


def build_task_type_summary(results_df: pd.DataFrame) -> pd.DataFrame:
    summary_df = (
        results_df.groupby("task_type")
        .agg(
            total=("case_id", "count"),
            passed=("passed", "sum"),
        )
        .reset_index()
    )

    summary_df["failed"] = summary_df["total"] - summary_df["passed"]
    summary_df["failure_rate"] = summary_df["failed"] / summary_df["total"]

    return summary_df.sort_values(
        by="failure_rate",
        ascending=False,
    )


def render_task_type_summary(results_df: pd.DataFrame) -> None:
    st.subheader("Task Type Failure Rate")

    if results_df.empty:
        st.info("沒有結果可以統計。")
        return

    summary_df = build_task_type_summary(results_df)

    st.bar_chart(
        summary_df,
        x="task_type",
        y="failure_rate",
    )

    display_df = summary_df.copy()
    display_df["failure_rate"] = display_df["failure_rate"].map(
        lambda value: f"{value:.1%}"
    )

    st.dataframe(
        display_df,
        use_container_width=True,
        hide_index=True,
    )


def render_representative_failures(results_df: pd.DataFrame) -> None:
    st.subheader("Representative Failures")

    failed_df = results_df[results_df["passed"] == False].copy()

    if failed_df.empty:
        st.info("沒有代表性失敗案例。")
        return

    for failure_type, group_df in failed_df.groupby("failure_type", dropna=False):
        example = group_df.iloc[0]
        label = failure_type or "unknown_failure"

        with st.expander(f"{label} - {example['case_id']}"):
            st.write("Input")
            st.code(example["input"] or "")

            st.write("Expected")
            st.code(example["expected"] or "")

            st.write("Actual")
            st.code(example["actual"] or "")

            st.write("Failure reason")
            st.code(example["failure_reason"] or "")

            st.write("Trace session id")
            st.code(example["trace_session_id"] or "")


def render_failure_dashboard() -> None:
    st.set_page_config(
        page_title="Agent Failure Dashboard",
        layout="wide",
    )

    st.title("Agent Failure Dashboard")
    st.caption("分析 eval run 的失敗案例、錯誤類型與任務類型失敗率")

    eval_run_files = list_eval_run_files()

    if not eval_run_files:
        st.info("目前沒有 eval run。請先執行 python3 -m evals.runner。")
        return

    selected_file = st.selectbox(
        "選擇 eval run",
        options=eval_run_files,
        format_func=lambda path: path.name,
    )

    eval_run = load_eval_run(selected_file)
    results_df = results_to_dataframe(eval_run)

    render_summary(eval_run, results_df)
    st.divider()

    render_failed_cases(results_df)
    st.divider()

    col1, col2 = st.columns(2)

    with col1:
        render_failure_type_chart(results_df)

    with col2:
        render_task_type_summary(results_df)

    st.divider()
    render_representative_failures(results_df)

這份程式碼沒有碰到 Agent 本身。

它只做四件事:

  1. 找到 eval run JSON。
  2. 載入 results。
  3. 用 Pandas 做統計。
  4. 用 Streamlit 顯示。

執行 Failure Dashboard

先確認已經有 eval run JSON。

如果還沒有,先執行:

LLM_PROVIDER=gemini python3 -m evals.runner

如果還沒有設定 Gemini API key,也可以先用 fake client:

python3 -m evals.runner

接著啟動 Failure Dashboard:

streamlit run failure_dashboard_app.py

打開瀏覽器後,應該會看到:

  • Summary metrics。
  • Failed cases table。
  • Failure type distribution。
  • Task type failure rate。
  • Representative failures。

如果畫面顯示:

目前沒有 eval run。請先執行 python3 -m evals.runner。

代表 data/eval_runs/ 裡還沒有任何 eval_run_*.json。

先跑一次 evaluation 即可。


讀 dashboard 時要看什麼?

Failure Dashboard 不是只為了畫圖。

它的目的,是讓我們做下一步工程決策。

建議閱讀順序如下。

第一,看 Success rate。

如果成功率明顯變動,要先確認這次 run 使用的是同一組 dataset、同一個 provider、同一個 prompt。

第二,看 Failure Type Distribution。

如果最多的是:

format_error

代表 Agent 不是不會回答,而是輸出沒有符合機器可解析格式。

這通常適合用:

  • 更明確的 JSON prompt。
  • output cleanup。
  • retry。
  • schema validation guardrail。

第三,看 Task Type Failure Rate。

如果 json_output 是 100% failure,但其他 task type 都正常,那問題很集中。

如果每個 task type 都失敗,可能是 provider、prompt 或 agent runner 本身有問題。

第四,看 Representative Failures。

圖表只能告訴我們類型,不能告訴我們具體原因。

例如同樣是 format_error,可能有幾種不同情況:

外面多了 markdown code fence
回傳 list 而不是 object
漏掉 required key
欄位值型別錯誤

所以最後一定要回到具體 actual output。


用 Day 16 的 Gemini baseline 解讀

假設 Day 16 的 Gemini baseline 有 15 題,其中 10 題通過、5 題失敗。

失敗分布可能是:

failure_type count
format_error 3
wrong_answer 2

task type 失敗率可能是:

task_type total failed failure_rate
json_output 3 3 100.0%
general_qa 3 2 66.7%
calculation 4 0 0.0%
instruction_following 2 0 0.0%
keyword_qa 3 0 0.0%

這份 dashboard 告訴我們幾件事。

第一,Gemini 已經能處理多數一般任務。

這和 fake client 不同。

fake client 大多只會通過 calculation,Gemini 則能通過更多 QA 和指令遵循任務。

第二,json_output 是目前最明顯的問題。

如果三題 JSON 題都失敗,而且 failure type 都是 format_error,代表後續可以優先處理格式控制。

第三,wrong_answer 不一定是真正答錯。

像 Day 16 提到的 case_009:

expected: 紀錄
actual: Trace 是指記錄並追蹤程式或請求在系統中執行的完整歷程...

這題在人類眼中可能可以接受。

但目前 evaluator 是字串包含檢查,所以「記錄」不等於「紀錄」。

Dashboard 會把它放進 wrong_answer,但真正改善時要回頭檢查 evaluator。


Failure Dashboard 和 Trace Viewer 的關係

Day 6 做的是 Trace Viewer。

它回答:

某一次 Agent 執行,每一步發生了什麼?

Day 17 做的是 Failure Dashboard。

它回答:

某一次 eval run,哪些 cases 失敗?失敗分布如何?

兩者關係可以這樣看:

Failure Dashboard
  -> 找到失敗 case
  -> 查看 trace_session_id
  -> 回到 Trace Viewer
  -> 檢查 Agent 執行步驟

例如 dashboard 顯示:

case_013
failure_type: format_error
trace_session_id: 4a2d...

這時可以打開 Trace Viewer:

streamlit run trace_viewer_app.py

再選擇對應 session,檢查:

  • user input 是否正確。
  • LLM 是否有輸出 JSON。
  • final answer 是否被 markdown code fence 包住。
  • 是否有 tool call。
  • 是否有 error step。

這就是 Trace 和 Eval 串起來後能做到的事。

Eval 幫我們縮小問題範圍。

Trace 幫我們看懂單一案例的執行細節。


目前 Dashboard 的限制

今天的 Failure Dashboard 是第一版,所以刻意保持簡單。

目前還有幾個限制。

1. 只能看單次 eval run

今天每次只選一份 JSON。

這適合分析單次 baseline,但還不能比較:

baseline prompt vs improved prompt
fake vs Gemini
retry off vs retry on

Day 18 做 prompt A/B testing 時,就會需要比較兩次或多次 eval runs。

2. trace session 只是文字

今天 dashboard 只顯示 trace_session_id。

它還不能直接點進 trace details。

後面可以把 Trace Viewer 的查詢邏輯整合進來,讓使用者點一個 case 就能展開 steps。

3. failure type 還是 rule-based

Day 16 的 classify_failure() 是規則式分類。

它可以處理明確錯誤,例如:

  • JSON parse error。
  • missing key。
  • instruction following fail。
  • execution error。

但它還不能精準區分:

  • Agent 真正答錯。
  • evaluator 太嚴格。
  • 語意正確但沒有關鍵字。
  • expected 本身寫得不夠好。

所以 dashboard 的分類結果要搭配 actual output 一起看。

4. 圖表還沒有做互動篩選

今天只是顯示 bar chart 和 dataframe。

未來可以加入:

  • 依 failure type 篩選 failed cases。
  • 依 task type 篩選。
  • 只看某個 grading method。
  • 搜尋 case id。
  • 下載 CSV。

但現在先不要加太多功能。

因為 Day 17 的重點是建立分析視角,而不是做完整 BI 工具。


重點整理

這次把 eval run JSON 變成第一版 Failure Dashboard。

新增的檔案有:

failure_dashboard_app.py
ui/failure_dashboard.py

Dashboard 可以顯示:

  • 整體 total / passed / failed / success rate。
  • failed cases table。
  • failure type distribution。
  • task type failure rate。
  • 每種 failure type 的代表性案例。
  • trace session id。

做到這一步後,我們已經不只是把 evaluation 跑完,而是開始建立一個分析迴路:

跑 eval
  -> 產生 results
  -> 分類 failure type
  -> dashboard 觀察分布
  -> 回 trace 看細節
  -> 決定下一步改善

這個迴路是測試 AI Agent 時很實用的基礎。

因為可靠性改善不應該靠感覺,而應該從固定 dataset、可重複的 eval run 和清楚的 failure analysis 開始。


下一步

Day 18 會開始處理 Prompt Versioning 與 Prompt A/B Testing。

現在我們已經能看單次 eval run 的 failure distribution。

下一篇會加入 prompt version 的概念,讓同一組 eval dataset 可以比較:

baseline prompt
improved prompt

到時候 Failure Dashboard 的指標就會派上用場。

我們不只會問:

新 prompt 成功率有沒有變高?

也會問:

format_error 有沒有下降?
wrong_answer 有沒有增加?
某些 task type 是否被新 prompt 影響?

這樣調 prompt 就不再只是憑感覺,而是可以拿 eval run 的結果來比較。


上一篇
Day 16|定義 Agent 失敗類型,並加入評測結果
系列文
從黑盒到可驗證:30 天打造 AI Agent 的 Trace、Eval 與 Guardrails 系統 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言