iT邦幫忙

2026 iThome 鐵人賽

DAY 19
0
AI Engineering

AI Reliability Lab:30 天用 Python 把「AI 好像很準」變成可以量的工程指標系列 第 19 篇

Day 19|384 次 REAL Requests 跑完了:48 題裡錯的 6 題,全都在 Multi-step Reasoning

  • 分享至 

  • xImage
  •  

Day 18 我把 benchmark 從 24 題擴到 48 題,四個 strata 各 12 題,也把下一次 REAL experiment 的 sampling budget 從每題 3 次提高到 7 次。當時沒有急著直接跑模型,而是先把 benchmark、Gold Answer、scoring、analysis plan 和 agreement < 0.8 這條 descriptive flag 都 freeze。

今天才真的開始跑。

這次的規模和 Day 15 差滿多。48 題,每題 1 個 primary response,再加 7 個 independent samples,所以總共是:

48 × (1 + 7) = 384

也就是 384 個 REAL logical requests。

最後 384 個全部成功,沒有 retry、沒有 duplicated slot,也沒有中途 resume。整個 REAL run ID 是:

20261002T134556Z-6819e3e5

模型維持 gpt-5.6-luna,Benchmark 也是昨天 freeze 的:

day18-stratified-reliability-v2

SHA-256 沒有改:

3d852dd70924a47b7b2c67a2b28f377f1c94959fdd66b8f3c0db998e11102847

這次我特別在意這件事,因為 Day 19 的結果比之前難看不少。如果是看到結果以後再回頭換題目、改 threshold 或調 normalization,那整個 experiment 就沒什麼意義了。所以這次 benchmark、metric、threshold、Gold、policy 全部照昨天的版本直接接受結果。

先看整體:42 / 48

這次 Accuracy 是:

42 / 48 = 87.50%

比 Day 15 的 23/24 低很多,但這兩個數字不能直接拿來說「模型退步了」,因為 Day 19 多了 24 個新 cases,而且 sampling design 也從 3 samples 改成 7 samples。比較合理的說法只是:在這份新的 48-case frozen benchmark 裡,這一次 REAL run 得到 87.5%。

Confidence 的結果則延續 Day 15 那個很明顯的問題。

Signal Mean Confidence Brier 5-bin ECE
Self-reported 1.000000 0.125000 0.125000
Agreement 0.910714 0.040816 0.041667

Self-reported confidence 又是:

1.0

也就是 48 題全部都說自己 100% 有把握,包含 6 題錯誤。

所以 Day 16 當時看到的 saturation 並沒有因為 benchmark 從 24 題擴到 48 題就消失。這次反而更明顯,因為 error 從 1 題增加到 6 題,self-confidence 還是完全沒有變。

Agreement 就有比較多變化,平均是 0.9107,Brier 和 ECE 也都低於 self-report。不過我還是不會直接寫成 Agreement「證明比較好」,目前比較能說的是:在這一次 48-case run 裡,Agreement 至少有比全部卡在 1.0 的 self-confidence 更多 case-level variation。

四個 strata 一拆開,差距很明顯

48 題平均分成四組,每組 12 題。

結果是:
https://ithelp.ithome.com.tw/upload/images/20261003/20184158mQl1wlAybe.png

Stratum n Correct Accuracy Mean Self Mean Agreement Self Brier Agreement Brier
deterministic_short_answer 12 12 100% 1.000000 1.000000 0.000000 0.000000
multi_step_reasoning 12 6 50% 1.000000 0.654762 0.500000 0.161565
insufficient_information 12 12 100% 1.000000 1.000000 0.000000 0.000000
distractor_resistance 12 12 100% 1.000000 0.988095 0.000000 0.001701

這張表很難不注意到 multi_step_reasoning。

其他三組全部都是 12/12,只有 multi-step 是:

6/12 = 50%

而且這 6 個錯誤全部都出現在這一組。

這個結果確實比 Day 15 更值得追,因為 Day 15 當時唯一錯誤也是 multi-step,但那時每組只有 6 題、總共也只有 1 個 error。現在擴到 12 題後,multi-step 又出現 6 個錯誤,至少代表這條線值得繼續研究。

但我還是不會寫成「GPT-5.6 Luna 不擅長 multi-step reasoning」。

這份 benchmark 仍然只有 48 題,而且是人工 curated,不是從完整 task population 隨機抽樣。現在比較安全的說法是:

在這 12 個 frozen multi-step cases 中,本次 REAL run 答對 6 題;其他三個 strata 則全部答對。

6 個錯誤全部還是 Overconfident

Day 13 那套 failure policy 這次完全沒有改,直接套到 48 筆 REAL records。

最後結果:
https://ithelp.ithome.com.tw/upload/images/20261003/20184158X87zEEw0M6.png

Signal Count
incorrect_answer 6
overconfident_error 6
high_agreement_error 1
low_agreement_error 4
insufficient_information_failure 0

Severity:

  • NONE:42
  • HIGH:6
  • MEDIUM:0
  • CRITICAL:0

這裡我覺得最有意思的是:六個錯誤全部都是 overconfident error。

原因也很單純,因為 self-confidence 每一題都是 1.0。

也就是這次模型沒有出現:

「我不確定,而且答錯。」

它出現的是:

「我完全確定,但答錯。」

不過 Agreement 把這六題又拆出了不同樣子。

六個錯誤其實不是同一種 disagreement

先看完整結果:

Case Gold Primary Agreement Additional Signal
D14-009 7 3 0.2857 low_agreement_error
D14-011 2 1 0.4286 low_agreement_error
D18-008 169 64 0.1429 low_agreement_error
D18-009 7 6 0.5714 —
D18-011 3,1 2,1 0.8571 high_agreement_error
D18-012 14 -3 0.2857 low_agreement_error

其中四題 Agreement 很低,所以被標成 low_agreement_error。

例如 D18-008 的七個 samples 是:

64, 144, 256, 729, 361, 81, 324

七個答案全部不同。

Agreement 只有:

1/7 = 0.142857

Normalized entropy 直接是:

1.0

這種 case 很明顯就是生成結果非常散。

D14-009 和 D18-012 也差不多,sample answers 分散得很明顯。

但 D18-011 完全是另一種情況。

它的 Gold 是:

3,1

Primary 回:

2,1

七個 samples 裡:

六次回 3,1

一次回 2,1

所以 Agreement 是:

6/7 = 0.857143

這題最後被標成:

high_agreement_error

它很有意思,因為 primary answer 雖然錯,但 repeated sampling 其實大多數都答對。

也就是說,Agreement 高不一定代表 primary 就可靠;它只是說 samples 本身高度集中。

這也再次提醒我,Agreement 不能被直接當成「primary correctness probability」。

7 Samples 讓 disagreement 看得比以前細很多

Day 15 每題只有 3 samples,所以 Agreement 幾乎就是 1.0 或 0.6667。

這次改成 7 samples 後,能看到的 level 多很多:

1/7, 2/7, 3/7 ... 7/7

實際上這六個 error 就已經出現:

0.1429、0.2857、0.4286、0.5714、0.8571

幾種不同程度。

整體 48 題中:

Agreement < 1.0 的 case 有:

8 題

Agreement < 0.8 的有:

6 題

Mean normalized entropy:

0.095336

Mean vote margin:

0.866071

這些數字本身還不能直接拿來判斷 correctness,但至少 sampling-based signal 不再像三次 sampling 時那麼粗。

昨天先定好的 Agreement < 0.8,這次真的抓到多少錯?

https://ithelp.ithome.com.tw/upload/images/20261003/201841582iVCZ5gHwC.pngDay 18 已經先 freeze 一條 descriptive rule:

agreement < 0.8 → flag

不是看到 Day 19 結果後才調 threshold。

最後 48 題裡:

6 題被 flag

其中:

  • 5 題是真的錯
  • 1 題是正確答案

總共 6 個 errors 裡:

  • 5 個被 flag
  • 1 個沒被 flag

所以 error capture 是:

5 / 6 = 83.33%

乍看不錯,而且只 flag 12.5% 的 case。

但它確實漏掉一題:

D18-011

原因就是剛剛提到的,它的 Agreement 是 0.8571,高於 0.8,所以不會被 flag。

這反而比「100% 全抓到」更有用,因為它很清楚地告訴我這條 rule 的限制。

sampling disagreement 可以抓到很多不穩定 case,但如果 samples 高度一致,而 primary 剛好落在少數答案那邊,Agreement threshold 還是可能漏掉錯誤。

所以目前 agreement < 0.8 仍然只能算 descriptive flag,不是 production safety rule。

Insufficient Information 這次還是 12 / 12

這次 insufficient_information 從 Day 15 的 6 題擴到 12 題。

結果:

12/12 全部正確 abstain

Failed abstention:

0

所以 insufficient_information_failure 仍然是 0。

這是目前兩次 REAL experiment 都很穩定的一個結果,但我還是只會限定在這些 frozen cases 上,不會直接說模型一般情況下都不會亂猜。

另外 distractor_resistance 也是 12/12,只有一題 samples 沒完全一致,所以 Mean Agreement 是 0.9881,但 primary 全部答對。

Legacy 和 New Cases 有差,但我先不硬解讀

Day 18 preregistration 裡有保留 legacy / new provenance,所以這次也能看到:

Legacy:

22/24

New:

20/24

這個差距看起來不大,但我沒有打算拿它做主要結論。

原因是 legacy 24 題在 Day 15 已經跑過,這些 outcome 已知,本來就不是完全 blind 的資料。而且我們也沒有為 legacy vs new 設計正式 statistical comparison。

所以這個數字保留做 descriptive cohort information 就好,不拿來講「新題比較難」之類的故事。

384 個 Requests 跑完後,先驗證再看結果

這次 request accounting 很乾淨:

Item Count
Planned logical requests 384
Completed logical requests 384
API attempts 384
Successful requests 384
Failed attempts 0
Retries 0
Duplicated slots 0

之前還有一次 SDK construction failure,但那次發生在 request 真的送出去以前,所以 API requests 是 0,不會混進這次 canonical REAL run。

完成後,新的 verifier 最後是:

22 / 22 PASS

而且 benchmark hash 仍然和 Day 18 完全一樣。

53 個 protected files 全部 unchanged。

Hidden metadata、Gold Answer、distractor annotation 和 insufficiency proof 也沒有送進模型。

這些東西現在已經有點像每次 REAL Experiment 的固定流程了:先 freeze、再跑、最後 verify。雖然看起來很囉嗦,但做到 Day 19 後,我覺得這反而是整個專案比較有價值的部分。

Day 19 做完後,新的問題反而更清楚了

這次最大的 observation 不只是 Accuracy 從 95.83% 變成 87.5%。

比較值得追的是三件事。

第一,self-reported confidence 到 48 題後還是完全卡在 1.0。現在已經不只是「只有一個錯誤時剛好沒變化」,而是六個錯誤全部也照樣報 100%。

第二,multi-step reasoning 在這 12 個 cases 中只答對 6 題,而且六個錯誤全部集中在這裡。這還不能變成 general conclusion,但至少已經值得專門做 error analysis。

第三,7-sample Agreement 確實比以前提供更細的 case-level variation,但 agreement < 0.8 還是漏掉 D18-011,表示 Agreement 本身不能直接當 correctness detector。

Day 19 完整測試在 REAL execution 前是:

401 passed in 56.17s

最後 canonical verifier:

22 / 22 PASS

這次總共送出:

384 個 REAL API requests

而且從 Day 18 freeze 到 Day 19 完成,沒有做任何 post-hoc benchmark、metric、threshold、normalization 或 policy 修改。

現在手上終於有 48 個 REAL cases、6 個 actual errors,而且每題有 7 次 independent samples。

接下來就不太適合再急著擴 benchmark。

我反而想停下來,好好看這六個 multi-step errors 到底有什麼共同結構。


上一篇
Day 18|24 題真的太少了,所以我把 Reliability Benchmark 擴成 48 題
下一篇
Day 20|模型到底錯在哪一步?我用 Deterministic Trace 拆 12 題 Multi-step Reasoning
系列文
AI Reliability Lab:30 天用 Python 把「AI 好像很準」變成可以量的工程指標 共 24 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言