前面做到現在,Agent 已經可以:
但在把一次成功的 conversation 保存成 reusable Analysis 之前,需要先確認一件事:
Agent 實際上的 behavior 到底好不好?
只看 final answer 不太夠。
答案可能是對的,但過程可能:
所以這次先做一次 system-level evaluation。
這次不是單獨測 model。
每一個 eval case 都直接跑真正的 application path:
Question
↓
Prompt + Context
↓
Agent
↓
Tools / Skills
↓
Sandbox
↓
Recovery
↓
Final Answer
也就是前面加進去的 prompt、tools、skills、context management、sandbox 和 recovery,都會一起影響結果。
目前先看四個主要 judgment:
另外再記錄實際 behavior:
這兩類東西刻意分開。
有些東西其實不需要 LLM 判斷。
例如:
用了幾次 query?
有沒有 read SKILL.md?
執行了幾次 code?
用了多少 tokens?
跑了多久?
這些直接從 execution trace 計算就好,屬於 deterministic evaluation。
但有些問題需要 semantic judgment:
Final answer 真的回答 user 的問題嗎?
這些 claim 有沒有 evidence 支持?
Agent 有沒有照 instruction 做?
Execution strategy 合不合理?
這些才交給 LLM-as-Judge。
簡單來說:
Code
→ What happened?
LLM-as-Judge
→ Was that behavior good?
能用 code 判斷的,就不需要再多一個 model call。
目前 Eval page 可以直接拿保存下來的 conversation 做 evaluation。
每一個 turn 都可以看到:
Correct
Grounded
Instruction
Strategy
Steps
Queries
Skill Reads
Code Execs
Tokens
Runtime

這比單純看 final answer 有用很多。
例如一個簡單的 database question 可能只需要:
1 Query
1 Skill Read
0 Code Execution
但如果 Agent 為了一個很簡單的問題跑了 5 次 query、2 次 Python,雖然最後答對,也值得回頭看。
這裡的 Steps、Queries、Tokens、Runtime 本身不是 pass / fail。
它們只是 measurement。
比較多 steps 不代表一定比較差,但至少要知道 Agent 為了得到答案做了什麼。
這次 evaluation 也真的抓到一個例子。
User 先要求:
Make a bar chart of the top 10 categories by number of distinct videos.
接著 follow-up:
Now make the same chart, but by total views.
這個 turn 的結果是:
Correct ✓
Grounded ✗
Instruction ✓
Strategy ✓


Final result 本身是對的,Agent 也有照 user instruction 做 chart,execution strategy 也沒有問題。
但 final answer 裡有些 claim,沒有被 Agent 實際觀察到的 execution result 支持。
所以:
Correct
≠
Grounded
如果 code 本身算錯,但 Agent 忠實地回報 execution result:
Wrong computation
→ result = 476
→ Agent says 476
那可能是:
Grounded ✓
Correct ✗
反過來,如果 Agent 根本沒有 observe 到某個數字,但 final answer 剛好講對:
Correct ✓
Grounded ✗
所以只做 answer accuracy 不夠。
目前 verdict 不只有:
pass
fail
還有:
unknown
judge error
例如 follow-up:
What is the median views per video in that category?
如果這個 eval case 沒有 independent reference answer,而且現有 context 也不足以驗證正確性,就不應該硬判 pass 或 fail。
直接回:
Correct: unknown
反而比較合理。
Evaluation 本身也不應該 hallucinate。
除了逐 turn 看之外,也做了一個 overview。
這次 20 turns 的結果:
Correct 85%
Grounded 95%
Instructions 100%
Strategy 100%
其中 Correct 是:
17 pass
3 unknown
Grounded 則是:
19 pass
1 fail

這個總攬可以快速看到:
哪一類 behavior 開始出問題?
如果之後改了 prompt、tool description、context management 或 execution policy,就可以重新跑同一批 eval,看有沒有 regression。
有了 system-level eval 之後,也可以拿來重新驗證前面做過的 architecture decision。
例如 Day 16、17 的 Jev Context Management。
這次直接拿 同一組 conversations 跑兩次:
Run A
→ 每個 turn 都送完整 conversation
Run B
→ 每個 turn 只送 Jev 選出的 earlier turns
兩邊回答的是同樣的 questions,差別只有 context strategy。

這樣就可以逐個 conversation 比較:
例如其中一個 Python correlation case:
Tokens
50.5K → 34.1K
Steps
6 → 4
Runtime
40.6s → 30.5s
但同時 Grounded 從:
1/1 → 0/1
所以 context filtering 不能只看 token 有沒有下降。
要一起看:
Quality
+
Tokens
+
Steps
+
Runtime
這裡不是要再重新討論一次 Jev,而是現在有了一套 evaluation,可以用同樣的方法檢查前面做過的設計。