iT邦幫忙

2026 iThome 鐵人賽

DAY 24
0
Software Development

猴子都寫得出來的 RISC-V OoO CPU系列 第 24 篇

Combination of Branch Sequence

  • 分享至 

  • xImage
  •  

既然做了這麼多種 HW
那就要來探討在什麼時候 branch 會變快,什麼時候會變慢了

   Case                    Single    Dual    OOO    OOO4
  ━━━━━━━━━━━━━━━━━━━━━━  ━━━━━━━━  ━━━━━━  ━━━━━  ━━━━━━
  1. Slow not taken           112     111    111     100
  ──────────────────────  ────────  ──────  ─────  ──────
  2. Independent taken         92      92     92     105
  ──────────────────────  ────────  ──────  ─────  ──────
  3. Bypasses stalled add     108     107     92     109
  ──────────────────────  ────────  ──────  ─────  ──────
  4. Discards wrong path       74      74     74      75
  ──────────────────────  ────────  ──────  ─────  ──────
  5. Backward loop             69      69     69      69
  ──────────────────────  ────────  ──────  ─────  ──────
  6. Fast not taken            29      29     29      29

恩.....
可以從整理過的數據上看到
out of order retire 有包含 rename recovery 的實作在很多 case 實際上是沒有優勢的
case 1: 因為可以處理先執行的指令,在 branch 被 stall 的時候執行時間上面有優勢
case 2: 因為 dual_ooo_retire 偷偷執行了後面的指令,猜錯會有更高的 panelty
case 3: 沒有做 rename recovery 的 dual_ooo 領先其他實作
case 4: 其他的都可以及早 redirect,只有 dual_ooo_retire 需要付出 rename recovery 的代價
case 5, case 6: 沒有速度上的差異

case 1: stalled branch, not-taken

  addi x10, x0, 80
  addi x11, x0, 2
  div  x1, x10, x11       # x1 = 40
  beq  x1, x0, done      # waits for DIV; not taken
  addi x2, x0, 7
  addi x3, x0, 9
  mul  x4, x2, x3        # x4 = 63
  add  x5, x4, x2        # x5 = 70
  done: add x6, x5, x0   # x6 = 70

case 2: independent branch, taken

  addi x10, x0, 80
  addi x11, x0, 2
  div  x1, x10, x11       # x1 = 40
  beq  x0, x0, target    # always taken; independent of DIV
  addi x1, x0, 99        # skipped architecturally
  div  x2, x10, x11      # skipped architecturally
  target: mul x3, x10, x11  # x3 = 160

case 3: independent branch, taken

  addi x10, x0, 80
  addi x11, x0, 2
  div  x1, x10, x11       # x1 = 40
  add  x2, x1, x1        # waits for DIV; x2 = 80
  beq  x0, x0, target    # ready without DIV
  addi x2, x0, 99        # skipped architecturally
  target: mul x3, x10, x11  # x3 = 160

case 4: dependent branch, taken

  addi x10, x0, 80
  addi x11, x0, 2
  div  x1, x10, x11       # x1 = 40
  bne  x1, x0, target    # waits for DIV; taken
  div  x2, x10, x11      # wrong path
  addi x1, x0, 99        # wrong-path renamed writer
  target: add x3, x1, x0 # x3 = 40

case 5: loop branch, 2 taken 1 not taken

  addi x1, x0, 3
  loop: addi x1, x1, -1
  bne  x1, x0, loop
  addi x2, x0, 7

case 6: independent, taken

  bne  x0, x0, done      # immediately known not taken
  addi x2, x0, 7
  addi x3, x0, 9
  done: add x4, x2, x3  # x4 = 16

上一篇
CPU Branch Stalled Behavior
系列文
猴子都寫得出來的 RISC-V OoO CPU 共 24 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言