Day 18 我把 benchmark 從 24 題擴到 48 題,四個 strata 各 12 題,也把下一次 REAL experiment 的 sampling budget 從每題 3 次提高到 7 次。當時沒有急著直接跑模型,而是先把 benchmark、Gold Answer、scoring、analysis plan 和 agreement < 0.8 這條 descriptive flag 都 freeze。
今天才真的開始跑。
這次的規模和 Day 15 差滿多。48 題,每題 1 個 primary response,再加 7 個 independent samples,所以總共是:
48 × (1 + 7) = 384
也就是 384 個 REAL logical requests。
最後 384 個全部成功,沒有 retry、沒有 duplicated slot,也沒有中途 resume。整個 REAL run ID 是:
20261002T134556Z-6819e3e5
模型維持 gpt-5.6-luna,Benchmark 也是昨天 freeze 的:
day18-stratified-reliability-v2
SHA-256 沒有改:
3d852dd70924a47b7b2c67a2b28f377f1c94959fdd66b8f3c0db998e11102847
這次我特別在意這件事,因為 Day 19 的結果比之前難看不少。如果是看到結果以後再回頭換題目、改 threshold 或調 normalization,那整個 experiment 就沒什麼意義了。所以這次 benchmark、metric、threshold、Gold、policy 全部照昨天的版本直接接受結果。
這次 Accuracy 是:
42 / 48 = 87.50%
比 Day 15 的 23/24 低很多,但這兩個數字不能直接拿來說「模型退步了」,因為 Day 19 多了 24 個新 cases,而且 sampling design 也從 3 samples 改成 7 samples。比較合理的說法只是:在這份新的 48-case frozen benchmark 裡,這一次 REAL run 得到 87.5%。
Confidence 的結果則延續 Day 15 那個很明顯的問題。
| Signal | Mean Confidence | Brier | 5-bin ECE |
|---|---|---|---|
| Self-reported | 1.000000 | 0.125000 | 0.125000 |
| Agreement | 0.910714 | 0.040816 | 0.041667 |
Self-reported confidence 又是:
1.0
也就是 48 題全部都說自己 100% 有把握,包含 6 題錯誤。
所以 Day 16 當時看到的 saturation 並沒有因為 benchmark 從 24 題擴到 48 題就消失。這次反而更明顯,因為 error 從 1 題增加到 6 題,self-confidence 還是完全沒有變。
Agreement 就有比較多變化,平均是 0.9107,Brier 和 ECE 也都低於 self-report。不過我還是不會直接寫成 Agreement「證明比較好」,目前比較能說的是:在這一次 48-case run 裡,Agreement 至少有比全部卡在 1.0 的 self-confidence 更多 case-level variation。
48 題平均分成四組,每組 12 題。
結果是:
| Stratum | n | Correct | Accuracy | Mean Self | Mean Agreement | Self Brier | Agreement Brier |
|---|---|---|---|---|---|---|---|
| deterministic_short_answer | 12 | 12 | 100% | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| multi_step_reasoning | 12 | 6 | 50% | 1.000000 | 0.654762 | 0.500000 | 0.161565 |
| insufficient_information | 12 | 12 | 100% | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| distractor_resistance | 12 | 12 | 100% | 1.000000 | 0.988095 | 0.000000 | 0.001701 |
這張表很難不注意到 multi_step_reasoning。
其他三組全部都是 12/12,只有 multi-step 是:
6/12 = 50%
而且這 6 個錯誤全部都出現在這一組。
這個結果確實比 Day 15 更值得追,因為 Day 15 當時唯一錯誤也是 multi-step,但那時每組只有 6 題、總共也只有 1 個 error。現在擴到 12 題後,multi-step 又出現 6 個錯誤,至少代表這條線值得繼續研究。
但我還是不會寫成「GPT-5.6 Luna 不擅長 multi-step reasoning」。
這份 benchmark 仍然只有 48 題,而且是人工 curated,不是從完整 task population 隨機抽樣。現在比較安全的說法是:
在這 12 個 frozen multi-step cases 中,本次 REAL run 答對 6 題;其他三個 strata 則全部答對。
Day 13 那套 failure policy 這次完全沒有改,直接套到 48 筆 REAL records。
最後結果:
| Signal | Count |
|---|---|
incorrect_answer |
6 |
overconfident_error |
6 |
high_agreement_error |
1 |
low_agreement_error |
4 |
insufficient_information_failure |
0 |
Severity:
這裡我覺得最有意思的是:六個錯誤全部都是 overconfident error。
原因也很單純,因為 self-confidence 每一題都是 1.0。
也就是這次模型沒有出現:
「我不確定,而且答錯。」
它出現的是:
「我完全確定,但答錯。」
不過 Agreement 把這六題又拆出了不同樣子。
先看完整結果:
| Case | Gold | Primary | Agreement | Additional Signal |
|---|---|---|---|---|
| D14-009 | 7 | 3 | 0.2857 | low_agreement_error |
| D14-011 | 2 | 1 | 0.4286 | low_agreement_error |
| D18-008 | 169 | 64 | 0.1429 | low_agreement_error |
| D18-009 | 7 | 6 | 0.5714 | — |
| D18-011 | 3,1 |
2,1 |
0.8571 | high_agreement_error |
| D18-012 | 14 | -3 | 0.2857 | low_agreement_error |
其中四題 Agreement 很低,所以被標成 low_agreement_error。
例如 D18-008 的七個 samples 是:
64, 144, 256, 729, 361, 81, 324
七個答案全部不同。
Agreement 只有:
1/7 = 0.142857
Normalized entropy 直接是:
1.0
這種 case 很明顯就是生成結果非常散。
D14-009 和 D18-012 也差不多,sample answers 分散得很明顯。
但 D18-011 完全是另一種情況。
它的 Gold 是:
3,1
Primary 回:
2,1
七個 samples 裡:
六次回 3,1
一次回 2,1
所以 Agreement 是:
6/7 = 0.857143
這題最後被標成:
high_agreement_error
它很有意思,因為 primary answer 雖然錯,但 repeated sampling 其實大多數都答對。
也就是說,Agreement 高不一定代表 primary 就可靠;它只是說 samples 本身高度集中。
這也再次提醒我,Agreement 不能被直接當成「primary correctness probability」。
Day 15 每題只有 3 samples,所以 Agreement 幾乎就是 1.0 或 0.6667。
這次改成 7 samples 後,能看到的 level 多很多:
1/7, 2/7, 3/7 ... 7/7
實際上這六個 error 就已經出現:
0.1429、0.2857、0.4286、0.5714、0.8571
幾種不同程度。
整體 48 題中:
Agreement < 1.0 的 case 有:
8 題
Agreement < 0.8 的有:
6 題
Mean normalized entropy:
0.095336
Mean vote margin:
0.866071
這些數字本身還不能直接拿來判斷 correctness,但至少 sampling-based signal 不再像三次 sampling 時那麼粗。
Day 18 已經先 freeze 一條 descriptive rule:
agreement < 0.8 → flag
不是看到 Day 19 結果後才調 threshold。
最後 48 題裡:
6 題被 flag
其中:
總共 6 個 errors 裡:
所以 error capture 是:
5 / 6 = 83.33%
乍看不錯,而且只 flag 12.5% 的 case。
但它確實漏掉一題:
D18-011
原因就是剛剛提到的,它的 Agreement 是 0.8571,高於 0.8,所以不會被 flag。
這反而比「100% 全抓到」更有用,因為它很清楚地告訴我這條 rule 的限制。
sampling disagreement 可以抓到很多不穩定 case,但如果 samples 高度一致,而 primary 剛好落在少數答案那邊,Agreement threshold 還是可能漏掉錯誤。
所以目前 agreement < 0.8 仍然只能算 descriptive flag,不是 production safety rule。
這次 insufficient_information 從 Day 15 的 6 題擴到 12 題。
結果:
12/12 全部正確 abstain
Failed abstention:
0
所以 insufficient_information_failure 仍然是 0。
這是目前兩次 REAL experiment 都很穩定的一個結果,但我還是只會限定在這些 frozen cases 上,不會直接說模型一般情況下都不會亂猜。
另外 distractor_resistance 也是 12/12,只有一題 samples 沒完全一致,所以 Mean Agreement 是 0.9881,但 primary 全部答對。
Day 18 preregistration 裡有保留 legacy / new provenance,所以這次也能看到:
Legacy:
22/24
New:
20/24
這個差距看起來不大,但我沒有打算拿它做主要結論。
原因是 legacy 24 題在 Day 15 已經跑過,這些 outcome 已知,本來就不是完全 blind 的資料。而且我們也沒有為 legacy vs new 設計正式 statistical comparison。
所以這個數字保留做 descriptive cohort information 就好,不拿來講「新題比較難」之類的故事。
這次 request accounting 很乾淨:
| Item | Count |
|---|---|
| Planned logical requests | 384 |
| Completed logical requests | 384 |
| API attempts | 384 |
| Successful requests | 384 |
| Failed attempts | 0 |
| Retries | 0 |
| Duplicated slots | 0 |
之前還有一次 SDK construction failure,但那次發生在 request 真的送出去以前,所以 API requests 是 0,不會混進這次 canonical REAL run。
完成後,新的 verifier 最後是:
22 / 22 PASS
而且 benchmark hash 仍然和 Day 18 完全一樣。
53 個 protected files 全部 unchanged。
Hidden metadata、Gold Answer、distractor annotation 和 insufficiency proof 也沒有送進模型。
這些東西現在已經有點像每次 REAL Experiment 的固定流程了:先 freeze、再跑、最後 verify。雖然看起來很囉嗦,但做到 Day 19 後,我覺得這反而是整個專案比較有價值的部分。
這次最大的 observation 不只是 Accuracy 從 95.83% 變成 87.5%。
比較值得追的是三件事。
第一,self-reported confidence 到 48 題後還是完全卡在 1.0。現在已經不只是「只有一個錯誤時剛好沒變化」,而是六個錯誤全部也照樣報 100%。
第二,multi-step reasoning 在這 12 個 cases 中只答對 6 題,而且六個錯誤全部集中在這裡。這還不能變成 general conclusion,但至少已經值得專門做 error analysis。
第三,7-sample Agreement 確實比以前提供更細的 case-level variation,但 agreement < 0.8 還是漏掉 D18-011,表示 Agreement 本身不能直接當 correctness detector。
Day 19 完整測試在 REAL execution 前是:
401 passed in 56.17s
最後 canonical verifier:
22 / 22 PASS
這次總共送出:
384 個 REAL API requests
而且從 Day 18 freeze 到 Day 19 完成,沒有做任何 post-hoc benchmark、metric、threshold、normalization 或 policy 修改。
現在手上終於有 48 個 REAL cases、6 個 actual errors,而且每題有 7 次 independent samples。
接下來就不太適合再急著擴 benchmark。
我反而想停下來,好好看這六個 multi-step errors 到底有什麼共同結構。