到底是時勢造英雄
還是英雄造時勢呢
投入了這麼多設計成本
既然 ooo_retire 看起來沒有優勢
我們就得讓他有優勢!!
(題外話: 昨天差兩個 cycle 數不出來,結果是眼睛瞎掉了看錯指令。果然還是不能太晚睡 Orz...
回過頭來看昨天提到的 test case
第一個狀況就是前面有幾道執行時間很長,又互相有 dependency 的指令
很明顯的觀察到 out of order core 在這邊比 in order 的優勢大很多
第二個狀況就是我們對 branch 的限制
在 ooo 的設計中我們為了簡化複雜度,選擇 blocking 直到 branch 在 execution 階段被 resolve
在 ooo_retire 的設計中我們增加了 predict miss panelty,讓後面的指令可以先執行
這些都是 pipeline 在設計上做出的 trade off
也就是說,這邊最後採用的是更複雜的 pipeline control 設計以及更高的 miss panelty
去提升 pipeline 的使用率
從這幾個 case 我們也可以觀察到,實際的 branch prediction miss panelty 影響多大
跟在白算盤上面學到只有一種 panelty 的狀況複雜很多
這時候我們會需要 emulator,讓我們在 RTL 設計出來之前能夠觀察到對架構的影響
像是這次開發的 systemC emulator 是其中一種類型
還有像是 tt-rpm 這種前端有 behavior model,後面用 performance model 來接的設計
addi x10, x0, 120
addi x11, x0, 3
div x1, x10, x11 # x1 = 40
div x2, x1, x11 # Waits for first DIV; x2 = 13
add x3, x2, x2 # Waits for second DIV; x3 = 26
beq x1, x0, done # Needs only first DIV; NOT taken
# Independent of both divides: a chain of useful younger work
addi x4, x0, 7
addi x5, x4, 1
addi x6, x5, 1
addi x7, x6, 1
addi x8, x7, 1
addi x9, x8, 1
addi x10, x9, 1
addi x11, x10, 1
addi x12, x11, 1
addi x13, x12, 1
addi x14, x13, 1
done:
add x20, x14, x3 # Joins both dependency chains; x20 = 43
run log:
/workspace/rtl/build/rtl/rtl_single_inorder_tests
6: [PASS] branch not taken divide and ALU chains [normal] rtl_cycles=148 systemc_cycles=148 retired=18
...
/workspace/rtl/build/rtl/rtl_dual_inorder_tests
7: [PASS] branch not taken divide and ALU chains [normal] rtl_cycles=147 systemc_cycles=147 retired=18
...
/workspace/rtl/build/rtl/rtl_dual_ooo_tests
8: [PASS] branch not taken divide and ALU chains [normal] rtl_cycles=110 systemc_cycles=110 retired=18
...
/workspace/rtl/build/rtl/rtl_dual_ooo_retire_tests
9: [PASS] branch not taken divide and ALU chains [normal] rtl_cycles=104 systemc_cycles=104 retired=18
git link with tag: https://github.com/hsufit/TINY5_OOO/tree/ithome2026_D26