API Test 只能證明 search_products() 沒壞,不能證明 Agent 遇到「我想買一個適合通勤的耳機」會選它。
Chrome 官方 WebMCP Evals 文件把測試拆得很清楚:Tool 本身仍要寫 deterministic tests;只要碰到模型判斷,就需要 evals,包含 Tool Selection、Arguments、Tool Chain 與完整 User Journey。
今天建立一份 30 組 Prompt Dataset,之後每改 Description/Schema 都能重跑。
不需要模型:
const result = await searchProducts({
keyword: 'keyboard',
maxPrice: 1500
});
console.assert(result.every(item => item.price <= 1500));
測 Business Logic。
Prompt:找 1500 元以下的鍵盤
Expected Tool:search_products
{
"keyword": "鍵盤",
"maxPrice": 1500
}
search_products
→ get_product
→ add_to_cart
→ get_cart
測整條 journey。
01 幫我找 WebMCP 的教學文章。
02 有沒有 WordPress REST API 的介紹?
03 找最近寫的 AI Agent 內容。
04 我想看 WooCommerce 技術文章。
05 找一篇不存在主題 xyz-123 的文章。
Expected:search_posts。
06 找 2000 元以下的耳機。
07 有沒有機械鍵盤?
08 我想買一個適合旅行的滑鼠。
09 找黑色 M 號外套。
10 幫我找一個不存在的商品 xyz-123。
Expected:search_products。
11 我上次買的鍵盤出貨了嗎?
12 查訂單 1234。
13 上個月我買了哪些東西?
14 我之前買過這款耳機嗎?
15 幫我看最近一筆訂單狀態。
Expected:search_orders / get_orders 類 Tool。
16 幫我解釋 JavaScript closure。
17 今天天氣如何?
18 幫我寫一封請假信。
19 1+1 等於多少?
20 你覺得這個網站好看嗎?
Expected:不要硬呼叫無關 WebMCP Tool。
21 找 1500 元以下鍵盤,第二個加入購物車。
22 找 WebMCP 文章,打開最相關的一篇。
23 找耳機,告訴我第一個的規格。
24 看購物車,移除最貴的商品。
25 找商品 A 並加入兩件,再告訴我購物車總額。
Expected:多 Tool chain。
26 找蘋果。 # 可能是商品?文章?
27 我之前那個呢? # context dependent
28 幫我處理訂單。 # 意圖不明
29 買第二個。 # 需要前文
30 幫我全部弄好。 # 不應直接猜高風險行為
這五組最有價值,因為 Production 使用者不會永遠像測試工程師一樣講完整句子。
可以用 JSON:
{
"id": "product-001",
"messages": [
{
"role": "user",
"content": "找 1500 元以下的鍵盤"
}
],
"expectedCall": [
{
"functionName": "search_products",
"arguments": {
"keyword": "鍵盤",
"maxPrice": 1500
}
}
]
}
📸 圖片 1|30 組 WebMCP Regression Dataset
這和 Chrome 官方 Evals 文件展示的 expectedCall 思路一致。
我會至少拆:
Tool selection accuracy
Argument accuracy
Negative-call accuracy
Sequence accuracy
Task completion rate
如果總分 90%,但 confirm_checkout 常常被提前呼叫,那個 90% 沒什麼安慰作用。
高風險 Tool 應該單獨看 False Positive。
例如:
search_products ✅
add_to_cart ✅
apply_coupon ❌
confirm_checkout ???
如果折價券失敗,Agent 是:
A. 停下來告知使用者
B. 默默用原價結帳
兩者差非常多。
官方 Evals 文件也特別提醒要測 Tool Chain 中途失敗時的行為。
這篇的 30 組是測試集,不是測試成績。
真正成績需要記錄:
Agent / Model
Chrome version
Tool definitions version
Run count
Date
📸 圖片 2|Evals 的實際分類結果
因為模型和 WebMCP 都會更新。