iT邦幫忙

2026 iThome 鐵人賽

DAY 25
0
Software Development

在 AI Compiler 工程師的路上系列 第 25

Day24:Triton NPU Flow:Triton 如何接到不同的 NPU backend

  • 分享至 

  • xImage
  •  

前三天看了 Tenstorrent Wormhole 的 tile / NoC、TT-Metal 的 reader /
compute / writer,以及 TT-MLIR 的 dialect 與 compiler pipeline。

我們在 Day14~Day19 已經提到過 Triton 在 NVIDIA GPU 上的路徑,今天好奇 Triton 如何走向多種不同的加速器,它們會從 Triton IR 的不同位置接出來,轉進各自的 MLIR dialect、runtime 或 device binary。

本篇大綱

  • 先回顧之前 Triton NVIDIA flow
  • 拆解 triton-shared 如何把 Triton 的 block 與 pointer 語意轉成通用 MLIR
  • 再逐條看六個 flow 的 IR
  • 最後說明為什麼接下來會把重點放在 TileLoom
Flow / 元件 GitHub / paper 簡單快速介紹
triton-shared facebookincubatorMicrosoft archive 共用的 Triton-to-MLIR 轉換層,把 block computation 和 pointer access 轉成 Linalg、Tensor 與 MemRef,讓後端可以繼續處理。
Triton-Tenstorrent kernelize-ai/triton-tenstorrent 保留 TTIR / TTGIR 前段,再轉入 Tenstorrent D2M 與 TTNN dialect,最後產生 TTNN FlatBuffer 給 runtime。
TileLoom Loom monorepoloom-dataflowpaper 論文以 Triton / triton-shared 接入通用 MLIR;新版 Loom monorepo 另加入 Helion frontend,再進行 spatial mapping、data reuse 與 NoC 搬移規劃。
Triton-Ascend triton-lang/triton-ascend 把 Triton 轉成 Linalg 與 AscendNPU IR,交給 BiSheng compiler 產生 Ascend device object,並由 CANN 載入執行。
Triton-XDNA amd/Triton-XDNA 把 triton code 轉成 Linalg / Transform IR,再接 MLIR-AIR / MLIR-AIE,將運算與資料搬移配置到 AMD AIE tile array。
Hexagon-MLIR qualcomm/hexagon-mlir 從 Linalg 為主的 MLIR 開始,針對 Qualcomm Hexagon 進行 tiling、vectorization 與 code generation。
Triton-MTIA paper HTMLpaper PDF Meta 在 MTIA 上的 production Triton backend,將 Triton 映射到 PE、fixed-function unit、circular buffer 與 RISC-V core;backend 未完整開源。

回顧之前 Triton NVIDIA flow

下圖把 Triton compiler 分成 frontend、middle-end 與 backend。對 NVIDIA GPU
來說,triton code 先轉成 Triton IR,再降層到帶有 GPU layout
與硬體 mapping 資訊的 TritonGPU IR。NVIDIA backend 接手後繼續
降層到 LLVM IR、PTX 與 cubin。

https://ithelp.ithome.com.tw/upload/images/20260825/20183319P8Er1B5lHT.png

圖片出處:Intel〈Triton architecture:Triton MLIR compiler details:Code generation overview〉)。這是 Intel XPU backend 專案用來說明 Triton 整體 compiler 架構的圖,圖中的 NVIDIA 分支對應本節的 GPU flow。

Python @triton.jit
  → TTIR
  → TTGIR
  → LLVM IR / NVVM
  → PTX
  → ptxas
  → cubin / SASS
  → CUDA Driver API
  → NVIDIA GPU

NVVM、PTX、cubin 與 CUDA Driver API 都是 NVIDIA 路徑的關鍵介面。換到
非 NVIDIA 加速器後,Triton Python 與 TTIR 仍可作為前端入口,後端則需要
改用能描述目標硬體的 IR 和 runtime

Triton → backend bridge → 非 NVIDIA IR / runtime → target device

先認識 triton-shared:Triton 與各家後端之間的轉換層

專案連結先放在這裡

截至 2026 年 8 月,Microsoft repository 已標示為不再維護。facebookincubator 目前另有一份
repository,比較新一點的版本,但是 README 內的 clone 指令仍指向 Microsoft 版本。
閱讀程式碼時要先確認自己看的 owner、commit 與 Triton submodule hash,避免把不同版本的
pass 名稱與支援範圍混在一起。

triton-shared 想提供的介面

Triton IR
  → shared middle layer
     ├─ block value:Linalg / Tensor
     └─ pointer / memory access:MemRef
  → hardware-specific IR

它的價值在於把 Triton block program 還原成 MLIR 生態系較容易分析的
structured computation 與 memory access。README 列出的轉換包括

Triton op / 語意 triton-shared 轉換後
tt.dot linalg.matmul
tt.reduce linalg.reduce
elementwise arith / math linalg.generic
tt.load 與相關 pointer arithmetic memref.reinterpret_castmemref.copy,再以 bufferization.to_tensor 接回 tensor computation
tt.store affine.storememref.tensor_store 相關操作

其中最麻煩的是 pointer,Triton 常用 tt.splattt.make_rangett.addptr
組出 tensor of pointers,硬體 DMA 通常只接受能寫成 base、shape、stride 的
structured access。triton-shared 會做 pointer、use 與 mask analysis,把能
辨識的位址計算轉成 strided MemRef。這也說明了它的限制:任意
scatter / gather 合法於 Triton,卻不一定能被這個 prototype 轉換。

若只想檢查 Triton dialect 轉成什麼,可以從 stand-alone 工具開始

triton-shared-opt --triton-to-linalg input.ttir

它也能藉由 TRITON_SHARED_DUMP_PATH 留下 tt.mlirttshared.mlir
ll.mlirll.ir。這些 artifact 可以讓我們能直接確認某個 backend 重用的是完整 Linalg lowering,還是其中一項 analysis。

Flow 1:Triton-Tenstorrent experimental TTMLIR runtime flow

第一條先看 Triton-Tenstorrent

Triton
=> TTIR
=> TTGIR
=> Tenstorrent D2M Dialect
=> Tenstorrent ttnn Dialect
=> Emit TTNN Flatbuffer

目標硬體:Tenstorrent Tensix core grid

這條 flow 面對的是 Day21 看過的 Tenstorrent 硬體,多個 Tensix core 排成 grid,每個 core 有自己的 L1,資料經由 NoC 在 DRAM tile、worker tile 和 其他介面之間移動。Kernel 能否執行,除了 tile 內的矩陣或向量運算,還要知道可用 core grid、memory 與 system topology。

因此這條 experimental flow 會先使用 ttrt query 產生 system descriptor。
System descriptor 和 kernel IR 分開,負責告訴 compiler / runtime 實際裝置提供哪些硬體資源。

為什麼走 D2M 與 TTNN dialect

這條線保留 Triton lowering 前段,後段進 Tenstorrent 自己的 MLIR dialect。
D2M 的全名是 Direct to Metal,用來把 Triton GPU IR 接到較貼近
Tenstorrent Metal programming model 的表示,後續再 lower 到 TTNN dialect,
交給 TTMLIR runtime 解讀或序列化。

這條線會讓 Triton 進入 Tenstorrent 的 TTNN / TTMLIR runtime 表示。
如果要直接產生 TT-Metalium C API,則是下一條 TileLoom 採用的 backend 形式。

後面如果研究這條線,要看的 artifact 會是

  • TTIR / TTGIR 裡 triton code 被表示成什麼
  • D2M dialect 如何描述 data movement
  • ttnn dialect 如何接到 runtime op
  • Flatbuffer 裡保存哪些執行資訊

這個專案比較偏向實驗性,目前明確把這條路徑標為 experimental。另一條 Metal Runtime
路徑也仍缺少可直接讓 Triton launch generated kernel 的完整 driver。

Flow 2:TileLoom

TileLoom 論文發表於 2026 年的
20th USENIX Symposium on Operating Systems Design and Implementation
(OSDI '26)

作者主要來自新加坡國立大學,合作作者則來自 Arizona State University、Google 與 Lumai Ltd.。這篇工作把 Triton 這類 tile-based language 接到 spatial dataflow accelerator,並在兩代 Tenstorrent 系統上驗證 compiler 規劃的效果。

前幾天提到 Tenstorrent 硬體有很多 compute tile、local memory 和 NoC,TT-Metalium 可以手動描述資料搬移和 kernel,TT-MLIR 則展示 Tenstorrent 如何用多個 dialect 串接 tensor graph、runtime op 與低階 kernel。TileLoom 想做的是把 spatial mapping 和 dataflow planning 納入 compiler pipeline。

OSDI '26 論文與 loom-dataflow 專案展示的 Triton flow 是

Triton
=> TTIR
=> triton-shared
=> affine / linalg / scf / memref
=> TileLoom dataflow-aware MLIR
=> TT-Metalium C API
=> Tenstorrent executable

Helion 有沒有 flow 到 Triton?

有,Helion 是一種嵌入 Python 的 kernel DSL,程式寫法接近 PyTorch with tiles。程式開發者在 hl.tile 迴圈裡使用 PyTorch operator,Helion 會建立 Python AST、補上 type 與 metadata,再降層成
由 FX Graph 和 TorchInductor IR 組成的 Device IR。最後的 codegen 結合 autotuner 選出的 config,產生 Triton code。

Helion kernel
  => Python AST / Extended AST
  => Device IR (FX Graphs + TorchInductor IR)
  => compiler passes + autotuned config
  => Triton code
  => Triton compiler / GPU backend

https://ithelp.ithome.com.tw/upload/images/20260825/20183319fKTE321cqW.png

圖片出處:PyTorch 官方文章〈Helion: A High-Level DSL for Performant and Portable ML Kernels〉的「High-Level Compiler Architecture」

這張官方圖可以確認 Helion 的標準 codegen 路徑會輸出 Triton
code。Loom 的 Helion frontend 選擇在更早的 Device IR 分流,直接轉入
high-level MLIR,這條路徑不需要先產生 Triton code。

Loom 的 Helion compilation pipeline

目前 ecolab-nus/loom 的官方
Architecture
文件以 Helion kernel 為入口。helion-mlir 直接讀取 Helion Device IR
裡的 FX Graph,把 control flow 轉成 affine.for / affine.parallel
並經由 torch-mlir 把 ATen operation 降層成 linalg-on-tensors。後面才進入
Loom 的 dataflow exploration、ETG 解析、block-size solver 與
materialization。

Helion kernel
  => Helion Device IR (FX Graphs)
  => helion-mlir
  => affine control flow + linalg-on-tensors
  => loom-dataflow exploration + ETG JSON
  => loom-mlar resolution
  => loom.solver block-size assignment
  => loom-dataflow materialization
  => bufferized Loom MLIR
  => optional loom2ttkernel / TTKernel lowering
Stage 官方名稱 這一階段做什麼
0 Helion Frontend helion-mlir 把綁定好參數的 Helion kernel 轉成 high-level MLIR。
1 Dataflow Exploration loom-dataflow 列舉硬體 mapping、標記 reuse / copy 選擇,並產生 explored MLIR 與 ETG JSON。
2 ETG Resolution loom-mlar 依據硬體的 performance model 解析各個 ETG variant。
3 Block-Size Solve loom.solver 使用 CPMpy / CP-SAT 搜尋可行且成本較低的 block size。
4 Materialization loom-dataflow 套用選定的 block size,把 tensor IR 降層成 bufferized Loom MLIR。
5 Optional TT Lowering loom2ttkernel 可選擇繼續降層到 TTKernel / tt-mlir / tt-metal code generation。

這裡要保留兩條路徑,Triton flow 說明論文與 loom-dataflow 如何經由 triton-shared 取得通用 MLIR,Helion flow 則是目前 Loom 的端到端入口。兩條路徑會在通用 MLIR 與 Loom dataflow passes 附近會合,它們的 frontend 與前段 IR 不同。

目標硬體:spatial dataflow accelerator

TileLoom 面對的是 spatial dataflow accelerator。它在實驗中使用兩代 Tenstorrent 系統,但設計目標不綁死單一晶片:硬體被拆成 compute core、local memory、global memory 與 interconnect topology。以 Wormhole 為例,這些欄位就會對應到 Tensix grid、每個 core 的 L1、DRAM tile 和 2D torus NoC。

這類硬體的效能取決於 logical tile 放在哪個 physical core、資料是否能在 L1 留住、能否沿 NoC broadcast,以及何時需要回 DRAM。TileLoom 的 hardware representation 因此會記錄 interconnect topology、memory hierarchy 和 compute capability,讓 compiler 能比較不同 mapping。

Triton flow 為什麼先轉通用 MLIR dialect

這條線的重點在 triton-shared。它把 Triton code 轉成比較一般的 MLIR 組合,例如 affine、linalg、scf、memref。這些 dialect 比 TTGIR 更適合做傳統 compiler analysis,也比較容易接 TileLoom 自己的 dataflow planning。

TileLoom 後面的工作包括

  • hardware representation:描述 core array、memory、NoC、compute resource。
  • spatiotemporal mapping:把 logical tile grid 放到 physical core grid 和時間迴圈。
  • reuse analysis:判斷哪些 A/B/C tile 可以跨 core 或跨時間重用。
  • broadcast / copy planning:把資料搬移具體化。
  • backend:轉成 TT-Metalium C API。

TileLoom 和 Triton-Tenstorrent 都能接到 Tenstorrent,但關注範圍不同。 Triton-Tenstorrent 展示 Triton plugin 如何進入 TTMLIR / TTNN runtime,TileLoom 把 spatiotemporal mapping、reuse、broadcast 和 memory allocation 當成主要 研究問題,最後生成 TT-Metalium C API,這條線最適合本系列後面的 paper
與 IR 實戰,因為每個 dataflow planning 階段都有可觀察的表示。

Flow 3:Triton-Ascend

Triton language / compiler / runtime
=> Triton-Ascend
   - Ascend language extension
   - compiler / driver / libdevice
=> TTIR (Triton IR)
=> Linalg IR
=> BiSheng Compiler
=> AscendNPU IR (HFusion Dialect / HIVM Dialect)
=> LLVM IR
=> Back-end code generation
=> triton_xxx_kernel.o
=> Triton-Ascend driver / CANN
=> Ascend NPU

官方圖可分成三塊來讀,最上層保留 Triton 共用的 compiler、language 與 runtime,中間層是 Triton-Ascend 加入的 language extension、compiler、driver 與 libdevice,下層再交給 BiSheng Compiler,從 AscendNPU IR 的 HFusion / HIVM dialect 繼續降層到 LLVM IR 與 backend code generation。

https://ithelp.ithome.com.tw/upload/images/20260825/20183319TPJHrbn16f.png

圖片出處:Triton-Ascend 官方文件〈[架構設計與核心特性](https://triton-ascend.readthedocs.io/zh-cn/latest/architecture_design_and_core_features.html#id2

目標硬體:Ascend NPU 的 AI Core

Ascend NPU 由多個 AI Core 提供平行計算。從 operator 開發角度,
一個 AI Core 內可以看到三類主要 compute resource

  • Cube Unit:處理矩陣類運算
  • Vector Unit:處理向量與 elementwise 運算
  • Scalar Unit:負責 scalar 計算與控制工作

記憶體部分會碰到 global memory、L1,以及 core 內的 Unified Buffer
(UB)。以 A2 系列文件為例,UB 是 192 KB,所以 Triton block 搬進單一
core 的 working set 太大時,compiler 會直接報 ub overflow。這使得
block size、multi-buffer 和資料搬移方式都必須配合硬體容量。

目前 Triton-Ascend 文件列出的硬體涵蓋 Atlas A2、A3 系列產品與 Ascend 950 系列。實際安裝仍要配合裝置型號、CANN 與 Triton-Ascend 版本。

為什麼走 Linalg 與 AscendNPU IR

這裡的路線是先把 Triton IR 轉成 Linalg IR,再由 BiSheng Compiler 接手 AscendNPU IR、LLVM IR 與 backend code generation,產生 device object triton_xxx_kernel.o。Triton-Ascend driver 接著經由 CANN 載入與啟動這個 device kernel。

可以先抓一個觀察,這條線也利用 Linalg 作為中繼層。Linalg 的好處是它把 tensor computation 表示成比較標準的 structured op,後端可以再針對自己的硬體 lowering。

AscendNPU IR 要進一步承接 AI Core 的 multi-core mapping、UB / L1 allocation、資料搬移,以及 Cube / Vector computation。和 TileLoom 相比,Ascend flow 的後半段銜接 BiSheng 與 CANN;TileLoom 則接 dataflow-aware MLIR 和 TT-Metalium。

Flow 4:Triton-XDNA

Triton
=> TTIR
=> triton-shared
=> MLIR Transform dialect
   (affine / linalg / scf / memref)
=> MLIR-AIR / MLIR-AIE
=> XRT binary (aie.xclbin)

XDNA / AIE 這條線也會經過 triton-shared,再進 affine / linalg / scf / memref。後面接 MLIR-AIR / MLIR-AIE,最後產出 XRT binary。

下圖把完整 MLIR-AIR stack 分成六段。Triton 在 frontend 經由
Triton-Shared 接入 MLIR community dialect,與 PyTorch / Torch-MLIR、
TensorFlow / TOSA 等入口匯流。SCF、MemRef 與 tiled Linalg 再降層到
MLIR-AIR,用 launch、segment / herd、channel 和 function call 表示
spatial execution 與 data movement。後端接到 MLIR-AIE,產生 NPU
configuration、DMA 與 runtime sequence,最後由 XRT / ROCr 載入 binary。

https://ithelp.ithome.com.tw/upload/images/20260825/20183319a5q4hgG6ry.png

圖片出處:Erwei Wang et al.,〈From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR〉,Figure 2「MLIR-AIR stack overview」,原始 PDF 第 8 頁。論文採 CC BY 4.0授權。

目標硬體:AMD XDNA 的 AI Engine tile array

Triton-XDNA 目前瞄準 AMD AIE2 與 AIE2P 架構,也就是 Ryzen AI NPU 使用的 AI Engine 路線。AI Engine 由空間排列的 tile array 組成,tile 包含 compute core 或 local memory,tile 之間經由 stream switch 傳資料,並由可程式化 DMA 管理資料搬移。

這個硬體形狀讓 compiler 必須決定三件事,運算切成多大的 tile、每份運算 放到哪個 AIE core、資料如何經由 stream 與 DMA 在 memory tile、compute tile 和 host 之間移動。

為什麼走 Transform、AIR 與 AIE dialect

triton-shared 先把 Triton code 轉成較緊湊的 Linalg compute graph。
MLIR Transform dialect 再描述 tiling、bufferization 和 vectorization;
MLIR-AIR / MLIR-AIE 負責 spatial mapping、array connectivity 與 data
movement,最後產生 aie.xclbin,交由 XRT runtime 載入。

這條線和 TileLoom 都要處理 spatial / array 類硬體。TileLoom 的 target
是 Tenstorrent 的 Tensix grid;Triton-XDNA 接到 AMD AIE tile array 與
XRT。目前 Triton-XDNA README 將專案標為 experimental,已列出的 kernel
涵蓋 matmul、elementwise、softmax 和 layer normalization,硬體與 dtype
支援仍應以專案的 examples dashboard 為準。

Flow 5:Hexagon-MLIR

Hexagon-MLIR flow 是

PyTorch => Torch-MLIR ------\
                              => MLIR IR
Triton  => Triton-to-Linalg -/
=> Hexagon-MLIR
   - quantization / fusion / tiling / buffering
   - multithreading / vectorization / math libraries
   - lowering / optimization
=> LLVM IR
=> Hexagon-LLVM
=> Runtime
=> Qualcomm Hexagon NPU

PyTorch model 經由 Torch-MLIR 轉成 MLIR IR,Triton kernel 則經由
Triton-to-Linalg 轉入同一個 structured MLIR 區域。Hexagon-MLIR 在這個
共用中繼表示上做 quantization、fusion、tiling、buffering、multithreading 與 vectorization,後面再降層到 LLVM IR 與 Hexagon-LLVM。

目標硬體:Qualcomm Hexagon NPU

Hexagon-MLIR 以 Qualcomm Hexagon NPU 為目標,這裡最重要的硬體資源包括

  • HVX(Hexagon Vector eXtensions),負責寬向量運算
  • TCM(Tightly Coupled Memory),提供低延遲的本地資料存放空間
  • DMA 在外部 DDR 與 TCM 之間搬移資料
  • matrix processing 目前公開專案藉由 Hexagon Kernel Library 提供 experimental 支援

因此這條 flow 除了把 scalar operation 轉成 LLVM IR,還要處理資料能否留在 TCM、迴圈如何 tile 成 HVX 適合的寬度,以及 DDR / TCM transfer 能否和 compute 重疊。

為什麼做 fusion、tiling 與 HVX vectorization

Hexagon-MLIR 會先把 Triton code 轉到 Linalg 等 structured IR,再做 fusion、tiling、memory promotion 與 HVX vectorization。Fusion 可以減少 中間 tensor 寫回外部記憶體的次數,並把多個 operation 組成 locality 較好的 mega-kernel,memory pass 則安排 TCM 與 DMA。

這條線會藉由 linalg-hexagon-opt 等工具觀察與執行 Hexagon-MLIR lowering,最後走到 LLVM IR,再產生 Hexagon assembly 與 object code,交給相應 runtime 執行。它的形狀較接近傳統 compiler backend,但最佳化仍由 NPU 的 HVX、TCM 和 DMA 決定。
Hexagon-MLIR 同時接受 Triton code 與 PyTorch model。Figure 1 保留兩個
frontend,可以看到它們如何匯流到 MLIR IR;和本文其他 flow 比較時,
則先聚焦 Triton-to-Linalg 這條分支。

Flow 6:Triton-MTIA

分享這個月剛出爐的論文
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
提供了一個更接近 production 的案例,Meta 把手寫 triton code 與 TorchInductor 產生的 triton code 部署到 MTIA-2i。

先看硬體。MTIA-2i 有 8 × 8 個 processing element(PE),PE 之間以客製化
NoC 相連,周圍再接 on-chip memory 與 LPDDR5 controller。

https://ithelp.ithome.com.tw/upload/images/20260825/20183319ZoTVk5uO2o.png

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
Figure 1(High-level architecture of MTIA-2i),arXiv:2608.00325v1(2026),
CC BY 4.0

https://ithelp.ithome.com.tw/upload/images/20260825/20183319DpZ1sd5vAz.png

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
Figure 2: PE’s internal organization,arXiv:2608.00325v1(2026),
CC BY 4.0

每個 PE 裡有兩顆 RISC-V processor core,以及多種 fixed-function unit
(FFU):DPE 做 GEMM、SE 做 vector / reduction、RE 累加 DPE 結果、MLU 做
layout transformation,FI 則用 DMA 搬移資料。長尾運算若沒有合適的 FFU,
可交給支援 RVV 的 RISC-V vector core。

MTIA-2i 的 local storage(LS)以 circular buffer(CB)管理。RISC-V core
非同步發出 FFU command,CB 的 read / write pointer 負責相依性。Compiler
因此要同時處理「一個 Triton op 應交給哪個 engine」、「tensor 放到哪個
CB」,以及「DMA 與 compute 怎麼 software pipeline」。

Triton-MTIA 的四段 compiler flow

論文 Figure 4 把 pipeline 分成四段

https://ithelp.ithome.com.tw/upload/images/20260825/20183319esaL8AUCSA.png

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
Figure 4(High-level flow of Triton-MTIA Compiler),arXiv:2608.00325v1(2026),
CC BY 4.0

把圖中的資訊展開後,可以寫成

手寫 triton kernel
  → upstream Triton parser(加上 MTIA language extension)
  → TTIR
  → Middle-end IR
     ├─ structured memory access → FI DMA
     ├─ block-wise compute → DPE / SE / MLU / RE
     └─ 其他運算 → RISC-V core / RVV
  → Backend IR
     ├─ tensor bufferization
     ├─ LS / circular buffer allocation and reuse
     └─ architecture-specific lowering
  → LLVM IR 或可讀的 C++
  → Clang compile and link
  → MTIA device binary

這條 flow 沒有沿用 GPU 的 TTGIR → NVVM / PTX 路線。團隊保留 upstream Triton parser 產生 TTIR,middle-end 再把 op 配到 MTIA 的硬體 engine,backend 才加入 LS、CB 與更低階的架構資訊。最後可產生 LLVM IR,也能產生容易對照 Triton source 的 C++,讓 kernel developer 修改與除錯後再交給 Clang。

triton-shared 在 MTIA flow 中負責什麼

MTIA 是理解 triton-shared 邊界的好例子。論文只明確指出它重用
triton-shared 的 Triton IR pointer analysis

Triton tensor-of-pointers
  → 分析 pointer arithmetic
  → 可表示為 base + shape + strides?
     ├─ 可以:改寫成 structured access,交給 DMA
     └─ 不行:由 RISC-V core 執行 scalar load / store

DMA 的頻寬利用率較高,但只接受 structured memory access。這個分析結果會直接影響一段原本為 GPU 寫的 Triton pointer code,到了 MTIA 後能不能走快的DMA path。它也呼應 triton-shared README 中的限制:pointer analysis 能處理strided pattern,尚未涵蓋任意 memory access。

這裡也要避免過度解讀,論文沒有說 MTIA backend 直接採用 triton-shared 的完整 Triton-to-Linalg pipeline,可以確認的重用範圍是 pointer analysis。MTIA 的 middle-end、backend dialect 與 production runtime 細節也沒有在開源 repository 中完整提供。

這篇 paper 的證據與限制

這篇 systems paper 的研究問題是,Triton 這種原本以 GPU 為中心的 kernel DSL,能否支援程式模型差異很大的 custom ML accelerator,並達到 production 所需的 coverage 與 performance?它提出 MTIA backend、TorchInductor codegen 調整和少量 MTIA language extension,再用 MTIA-2i silicon 與 production deployment 回答。

論文摘要報告的最終 production footprint 是約 60 種 model type、50% 的 layer,以及 47% 的 non-GEMM execution time。Figure 12 顯示從 2025 Q1 到 Q2,以 layer 與 runtime 計算的 coverage 都大約增加一倍。

https://ithelp.ithome.com.tw/upload/images/20260825/20183319PvKwQdzFTK.png

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
Figure 12(Triton’s production footprint in models on MTIA),
arXiv:2608.00325v1(2026),
CC BY 4.0

閱讀 Figure 12 時要留意,圖中的分母排除 GEMM,不能解讀成 Triton 已涵蓋整個 model execution time。

移植經驗也很具體,16 個原本跑在 GPU 的 triton code 中,11 個可直接移植
並提供足夠效能,另外 5 個需要修改 pointer arithmetic、numeric handling 或
MTIA-specific optimization。這組資料支持 Triton source portability,也顯示
高效能 kernel 不會自動跨硬體保持同一份寫法。

這份結果來自 Meta 內部 MTIA software stack、模型與 production workload,
compiler backend 也沒有完整開源。它能證明 Triton-MTIA 已大規模部署,無法讓
外部讀者獨立重現全部數字。文章中提到的 GEMM 超過 80% architecture roofline、
coverage 與 speedup,都應保留論文限定的 shape、model 與 MTIA-2i 條件。

六條 flow 的共同問題

可以把這六條 flow 放成一張表

Flow 目標硬體 硬體映射重點 最後接到哪裡
Triton-Tenstorrent TTMLIR Tensix core grid L1、NoC、system topology TTMLIR / TTNN runtime、Flatbuffer
TileLoom Spatial dataflow accelerator;實驗使用 Tenstorrent core mapping、reuse、broadcast、memory allocation TT-Metalium C API
Triton-Ascend Ascend AI Core Cube / Vector、UB / L1、multi-core mapping BiSheng object、CANN
Triton-XDNA AMD AIE2 / AIE2P tile array tiling、core placement、stream / DMA aie.xclbin、XRT
Hexagon-MLIR Qualcomm Hexagon NPU fusion、HVX、TCM、DMA LLVM object、Hexagon runtime
Triton-MTIA MTIA-2i 8 × 8 PE array op-to-engine mapping、LS / CB、software pipelining、PID-to-PE mapping LLVM IR 或 C++、Clang、MTIA device binary

共同問題像是

  • 什麼資訊要保留到後端?
  • layout、tiling、memory movement 要在哪一層決定?
  • 哪些最佳化可以用通用 MLIR dialect 做?
  • 哪些最佳化必須進硬體 dialect 才能表達?

為什麼接下來看 TileLoom

第一,它和 Tenstorrent 硬體形狀很貼近。Wormhole 是 core grid + local memory + NoC,TileLoom 的 paper 也正是在處理 spatiotemporal mapping 和 data movement planning。

第二,它保留了 Triton 作為前端。這讓我們可以從比較熟悉的 triton code 出發,看 compiler 怎麼把 tile-level program 變成 dataflow-aware program。

第三,它有公開的 mm IR 階段,可以拿來逐步對照。後面的實戰篇會用 test/Passes/mm/IR 裡的九個 .mlir 檔案,看一個 matmul 如何從 frontend IR 逐步變成接近 Tenstorrent backend 的表示。

今天先走到這裡

今天先把幾條 Triton flow 排在一起

  • affine / linalg / scf / memref 常被拿來當可分析、可轉換的 MLIR 中繼層。
  • Tenstorrent 有兩條值得區分的方向:TTMLIR / TTNN runtime flow,以及 TileLoom 這種 dataflow planning flow。
  • Triton-Ascend 要把 tile 對應到 AI Core、UB / L1 和 Cube / Vector Unit,最後交給 BiSheng 與 CANN。
  • Triton-XDNA 要把運算與資料搬移排到 AIE tile array,最後由 XRT 載入 aie.xclbin
  • Hexagon-MLIR 的重點是 fusion、HVX vectorization、TCM locality 與 DMA。
  • Triton-MTIA 保留 TTIR 前端,middle-end 決定 op 要交給 FFU、DMA 或 RISC-V core,backend 再管理 LS 與 circular buffer。
  • TileLoom 的特色是把 tile program 的 logical grid 放到 spatial dataflow accelerator 的 physical core grid。

明天開始讀 TileLoom paper,好奇 spatial dataflow accelerator 會把哪些 mapping 與 data movement 決策交給 compiler?

參考資料


上一篇
Day23:TT-MLIR:Tenstorrent 的 MLIR dialect 與 compiler flow
下一篇
Day25:TileLoom :spatial dataflow accelerator 的編譯問題
系列文
在 AI Compiler 工程師的路上32
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言