iT邦幫忙

2026 iThome 鐵人賽

DAY 22
0
AI 自動化

AaaS from Scratch: 從一次性定義,到規模化分析系列 第 22 篇

[Day 22] Agent Behavior Evaluation

  • 分享至 

  • xImage
  •  

前面做到現在,Agent 已經可以:

  • query data
  • read skills
  • write and execute code
  • manage conversation context
  • recover from execution errors
  • rewrite code after OOM
  • move 到 bigger sandbox

但在把一次成功的 conversation 保存成 reusable Analysis 之前,需要先確認一件事:

Agent 實際上的 behavior 到底好不好?

只看 final answer 不太夠。

答案可能是對的,但過程可能:

  • 用錯 tool
  • 做了很多不必要的 query
  • 沒有照 user instruction
  • 遇到 error 只是不斷 retry
  • final answer 裡的 claim 沒有 execution evidence

所以這次先做一次 system-level evaluation。

Evaluate the Whole Agent System

這次不是單獨測 model。

每一個 eval case 都直接跑真正的 application path:

Question
↓
Prompt + Context
↓
Agent
↓
Tools / Skills
↓
Sandbox
↓
Recovery
↓
Final Answer

也就是前面加進去的 prompt、tools、skills、context management、sandbox 和 recovery,都會一起影響結果。

目前先看四個主要 judgment:

  • Correct
  • Grounded
  • Instruction Following
  • Execution Strategy

另外再記錄實際 behavior:

  • Steps
  • Queries
  • Skill Reads
  • Code Executions
  • Tokens
  • Runtime

這兩類東西刻意分開。

Code Check vs LLM-as-Judge

有些東西其實不需要 LLM 判斷。

例如:

用了幾次 query?
有沒有 read SKILL.md?
執行了幾次 code?
用了多少 tokens?
跑了多久?

這些直接從 execution trace 計算就好,屬於 deterministic evaluation。

但有些問題需要 semantic judgment:

Final answer 真的回答 user 的問題嗎?
這些 claim 有沒有 evidence 支持?
Agent 有沒有照 instruction 做?
Execution strategy 合不合理?

這些才交給 LLM-as-Judge。

簡單來說:

Code
→ What happened?

LLM-as-Judge
→ Was that behavior good?

能用 code 判斷的,就不需要再多一個 model call。

每個 Turn 都看 Behavior

目前 Eval page 可以直接拿保存下來的 conversation 做 evaluation。

每一個 turn 都可以看到:

Correct
Grounded
Instruction
Strategy
Steps
Queries
Skill Reads
Code Execs
Tokens
Runtime

這比單純看 final answer 有用很多。

例如一個簡單的 database question 可能只需要:

1 Query
1 Skill Read
0 Code Execution

但如果 Agent 為了一個很簡單的問題跑了 5 次 query、2 次 Python,雖然最後答對,也值得回頭看。

這裡的 Steps、Queries、Tokens、Runtime 本身不是 pass / fail。

它們只是 measurement。

比較多 steps 不代表一定比較差,但至少要知道 Agent 為了得到答案做了什麼。

Correct 不等於 Grounded

這次 evaluation 也真的抓到一個例子。

User 先要求:

Make a bar chart of the top 10 categories by number of distinct videos.

接著 follow-up:

Now make the same chart, but by total views.

這個 turn 的結果是:

Correct      ✓
Grounded     ✗
Instruction  ✓
Strategy     ✓

Final result 本身是對的,Agent 也有照 user instruction 做 chart,execution strategy 也沒有問題。

但 final answer 裡有些 claim,沒有被 Agent 實際觀察到的 execution result 支持。

所以:

Correct
≠
Grounded

如果 code 本身算錯,但 Agent 忠實地回報 execution result:

Wrong computation
→ result = 476
→ Agent says 476

那可能是:

Grounded ✓
Correct  ✗

反過來,如果 Agent 根本沒有 observe 到某個數字,但 final answer 剛好講對:

Correct  ✓
Grounded ✗

所以只做 answer accuracy 不夠。

Unknown 也是結果

目前 verdict 不只有:

pass
fail

還有:

unknown
judge error

例如 follow-up:

What is the median views per video in that category?

如果這個 eval case 沒有 independent reference answer,而且現有 context 也不足以驗證正確性,就不應該硬判 pass 或 fail。

直接回:

Correct: unknown

反而比較合理。

Evaluation 本身也不應該 hallucinate。

Overview

除了逐 turn 看之外,也做了一個 overview。

這次 20 turns 的結果:

Correct       85%
Grounded      95%
Instructions 100%
Strategy     100%

其中 Correct 是:

17 pass
3 unknown

Grounded 則是:

19 pass
1 fail

這個總攬可以快速看到:

哪一類 behavior 開始出問題?

如果之後改了 prompt、tool description、context management 或 execution policy,就可以重新跑同一批 eval,看有沒有 regression。

Eval 也可以驗證之前的設計

有了 system-level eval 之後,也可以拿來重新驗證前面做過的 architecture decision。

例如 Day 16、17 的 Jev Context Management。

這次直接拿 同一組 conversations 跑兩次:

Run A
→ 每個 turn 都送完整 conversation

Run B
→ 每個 turn 只送 Jev 選出的 earlier turns

兩邊回答的是同樣的 questions,差別只有 context strategy。

這樣就可以逐個 conversation 比較:

  • Correct
  • Grounded
  • Tokens
  • Steps
  • Runtime
  • Jev 到底送了哪些 earlier turns

例如其中一個 Python correlation case:

Tokens
50.5K → 34.1K

Steps
6 → 4

Runtime
40.6s → 30.5s

但同時 Grounded 從:

1/1 → 0/1

所以 context filtering 不能只看 token 有沒有下降。

要一起看:

Quality
+
Tokens
+
Steps
+
Runtime

這裡不是要再重新討論一次 Jev,而是現在有了一套 evaluation,可以用同樣的方法檢查前面做過的設計。


上一篇
[Day 21] Sandbox Scaling (2)
系列文
AaaS from Scratch: 從一次性定義,到規模化分析 共 22 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言