前三天看了 Tenstorrent Wormhole 的 tile / NoC、TT-Metal 的 reader /
compute / writer,以及 TT-MLIR 的 dialect 與 compiler pipeline。
我們在 Day14~Day19 已經提到過 Triton 在 NVIDIA GPU 上的路徑,今天好奇 Triton 如何走向多種不同的加速器,它們會從 Triton IR 的不同位置接出來,轉進各自的 MLIR dialect、runtime 或 device binary。
| Flow / 元件 | GitHub / paper | 簡單快速介紹 |
|---|---|---|
| triton-shared | facebookincubator、Microsoft archive | 共用的 Triton-to-MLIR 轉換層,把 block computation 和 pointer access 轉成 Linalg、Tensor 與 MemRef,讓後端可以繼續處理。 |
| Triton-Tenstorrent | kernelize-ai/triton-tenstorrent | 保留 TTIR / TTGIR 前段,再轉入 Tenstorrent D2M 與 TTNN dialect,最後產生 TTNN FlatBuffer 給 runtime。 |
| TileLoom | Loom monorepo、loom-dataflow、paper | 論文以 Triton / triton-shared 接入通用 MLIR;新版 Loom monorepo 另加入 Helion frontend,再進行 spatial mapping、data reuse 與 NoC 搬移規劃。 |
| Triton-Ascend | triton-lang/triton-ascend | 把 Triton 轉成 Linalg 與 AscendNPU IR,交給 BiSheng compiler 產生 Ascend device object,並由 CANN 載入執行。 |
| Triton-XDNA | amd/Triton-XDNA | 把 triton code 轉成 Linalg / Transform IR,再接 MLIR-AIR / MLIR-AIE,將運算與資料搬移配置到 AMD AIE tile array。 |
| Hexagon-MLIR | qualcomm/hexagon-mlir | 從 Linalg 為主的 MLIR 開始,針對 Qualcomm Hexagon 進行 tiling、vectorization 與 code generation。 |
| Triton-MTIA | paper HTML、paper PDF | Meta 在 MTIA 上的 production Triton backend,將 Triton 映射到 PE、fixed-function unit、circular buffer 與 RISC-V core;backend 未完整開源。 |
下圖把 Triton compiler 分成 frontend、middle-end 與 backend。對 NVIDIA GPU
來說,triton code 先轉成 Triton IR,再降層到帶有 GPU layout
與硬體 mapping 資訊的 TritonGPU IR。NVIDIA backend 接手後繼續
降層到 LLVM IR、PTX 與 cubin。

圖片出處:Intel〈Triton architecture:Triton MLIR compiler details:Code generation overview〉)。這是 Intel XPU backend 專案用來說明 Triton 整體 compiler 架構的圖,圖中的 NVIDIA 分支對應本節的 GPU flow。
Python @triton.jit
→ TTIR
→ TTGIR
→ LLVM IR / NVVM
→ PTX
→ ptxas
→ cubin / SASS
→ CUDA Driver API
→ NVIDIA GPU
NVVM、PTX、cubin 與 CUDA Driver API 都是 NVIDIA 路徑的關鍵介面。換到
非 NVIDIA 加速器後,Triton Python 與 TTIR 仍可作為前端入口,後端則需要
改用能描述目標硬體的 IR 和 runtime
Triton → backend bridge → 非 NVIDIA IR / runtime → target device
專案連結先放在這裡
截至 2026 年 8 月,Microsoft repository 已標示為不再維護。facebookincubator 目前另有一份
repository,比較新一點的版本,但是 README 內的 clone 指令仍指向 Microsoft 版本。
閱讀程式碼時要先確認自己看的 owner、commit 與 Triton submodule hash,避免把不同版本的
pass 名稱與支援範圍混在一起。
triton-shared 想提供的介面
Triton IR
→ shared middle layer
├─ block value:Linalg / Tensor
└─ pointer / memory access:MemRef
→ hardware-specific IR
它的價值在於把 Triton block program 還原成 MLIR 生態系較容易分析的
structured computation 與 memory access。README 列出的轉換包括
| Triton op / 語意 | triton-shared 轉換後 |
|---|---|
tt.dot |
linalg.matmul |
tt.reduce |
linalg.reduce |
elementwise arith / math |
linalg.generic |
tt.load 與相關 pointer arithmetic |
memref.reinterpret_cast、memref.copy,再以 bufferization.to_tensor 接回 tensor computation |
tt.store |
affine.store 或 memref.tensor_store 相關操作 |
其中最麻煩的是 pointer,Triton 常用 tt.splat、tt.make_range、tt.addptr
組出 tensor of pointers,硬體 DMA 通常只接受能寫成 base、shape、stride 的
structured access。triton-shared 會做 pointer、use 與 mask analysis,把能
辨識的位址計算轉成 strided MemRef。這也說明了它的限制:任意
scatter / gather 合法於 Triton,卻不一定能被這個 prototype 轉換。
若只想檢查 Triton dialect 轉成什麼,可以從 stand-alone 工具開始
triton-shared-opt --triton-to-linalg input.ttir
它也能藉由 TRITON_SHARED_DUMP_PATH 留下 tt.mlir、ttshared.mlir、ll.mlir 與 ll.ir。這些 artifact 可以讓我們能直接確認某個 backend 重用的是完整 Linalg lowering,還是其中一項 analysis。
第一條先看 Triton-Tenstorrent
Triton
=> TTIR
=> TTGIR
=> Tenstorrent D2M Dialect
=> Tenstorrent ttnn Dialect
=> Emit TTNN Flatbuffer
這條 flow 面對的是 Day21 看過的 Tenstorrent 硬體,多個 Tensix core 排成 grid,每個 core 有自己的 L1,資料經由 NoC 在 DRAM tile、worker tile 和 其他介面之間移動。Kernel 能否執行,除了 tile 內的矩陣或向量運算,還要知道可用 core grid、memory 與 system topology。
因此這條 experimental flow 會先使用 ttrt query 產生 system descriptor。
System descriptor 和 kernel IR 分開,負責告訴 compiler / runtime 實際裝置提供哪些硬體資源。
這條線保留 Triton lowering 前段,後段進 Tenstorrent 自己的 MLIR dialect。
D2M 的全名是 Direct to Metal,用來把 Triton GPU IR 接到較貼近
Tenstorrent Metal programming model 的表示,後續再 lower 到 TTNN dialect,
交給 TTMLIR runtime 解讀或序列化。
這條線會讓 Triton 進入 Tenstorrent 的 TTNN / TTMLIR runtime 表示。
如果要直接產生 TT-Metalium C API,則是下一條 TileLoom 採用的 backend 形式。
後面如果研究這條線,要看的 artifact 會是
這個專案比較偏向實驗性,目前明確把這條路徑標為 experimental。另一條 Metal Runtime
路徑也仍缺少可直接讓 Triton launch generated kernel 的完整 driver。
TileLoom 論文發表於 2026 年的
20th USENIX Symposium on Operating Systems Design and Implementation
(OSDI '26)。
作者主要來自新加坡國立大學,合作作者則來自 Arizona State University、Google 與 Lumai Ltd.。這篇工作把 Triton 這類 tile-based language 接到 spatial dataflow accelerator,並在兩代 Tenstorrent 系統上驗證 compiler 規劃的效果。
前幾天提到 Tenstorrent 硬體有很多 compute tile、local memory 和 NoC,TT-Metalium 可以手動描述資料搬移和 kernel,TT-MLIR 則展示 Tenstorrent 如何用多個 dialect 串接 tensor graph、runtime op 與低階 kernel。TileLoom 想做的是把 spatial mapping 和 dataflow planning 納入 compiler pipeline。
OSDI '26 論文與 loom-dataflow 專案展示的 Triton flow 是
Triton
=> TTIR
=> triton-shared
=> affine / linalg / scf / memref
=> TileLoom dataflow-aware MLIR
=> TT-Metalium C API
=> Tenstorrent executable
有,Helion 是一種嵌入 Python 的 kernel DSL,程式寫法接近 PyTorch with tiles。程式開發者在 hl.tile 迴圈裡使用 PyTorch operator,Helion 會建立 Python AST、補上 type 與 metadata,再降層成
由 FX Graph 和 TorchInductor IR 組成的 Device IR。最後的 codegen 結合 autotuner 選出的 config,產生 Triton code。
Helion kernel
=> Python AST / Extended AST
=> Device IR (FX Graphs + TorchInductor IR)
=> compiler passes + autotuned config
=> Triton code
=> Triton compiler / GPU backend

圖片出處:PyTorch 官方文章〈Helion: A High-Level DSL for Performant and Portable ML Kernels〉的「High-Level Compiler Architecture」
這張官方圖可以確認 Helion 的標準 codegen 路徑會輸出 Triton
code。Loom 的 Helion frontend 選擇在更早的 Device IR 分流,直接轉入
high-level MLIR,這條路徑不需要先產生 Triton code。
目前 ecolab-nus/loom 的官方
Architecture
文件以 Helion kernel 為入口。helion-mlir 直接讀取 Helion Device IR
裡的 FX Graph,把 control flow 轉成 affine.for / affine.parallel,
並經由 torch-mlir 把 ATen operation 降層成 linalg-on-tensors。後面才進入
Loom 的 dataflow exploration、ETG 解析、block-size solver 與
materialization。
Helion kernel
=> Helion Device IR (FX Graphs)
=> helion-mlir
=> affine control flow + linalg-on-tensors
=> loom-dataflow exploration + ETG JSON
=> loom-mlar resolution
=> loom.solver block-size assignment
=> loom-dataflow materialization
=> bufferized Loom MLIR
=> optional loom2ttkernel / TTKernel lowering
| Stage | 官方名稱 | 這一階段做什麼 |
|---|---|---|
| 0 | Helion Frontend | helion-mlir 把綁定好參數的 Helion kernel 轉成 high-level MLIR。 |
| 1 | Dataflow Exploration | loom-dataflow 列舉硬體 mapping、標記 reuse / copy 選擇,並產生 explored MLIR 與 ETG JSON。 |
| 2 | ETG Resolution | loom-mlar 依據硬體的 performance model 解析各個 ETG variant。 |
| 3 | Block-Size Solve | loom.solver 使用 CPMpy / CP-SAT 搜尋可行且成本較低的 block size。 |
| 4 | Materialization | loom-dataflow 套用選定的 block size,把 tensor IR 降層成 bufferized Loom MLIR。 |
| 5 | Optional TT Lowering | loom2ttkernel 可選擇繼續降層到 TTKernel / tt-mlir / tt-metal code generation。 |
這裡要保留兩條路徑,Triton flow 說明論文與 loom-dataflow 如何經由 triton-shared 取得通用 MLIR,Helion flow 則是目前 Loom 的端到端入口。兩條路徑會在通用 MLIR 與 Loom dataflow passes 附近會合,它們的 frontend 與前段 IR 不同。
TileLoom 面對的是 spatial dataflow accelerator。它在實驗中使用兩代 Tenstorrent 系統,但設計目標不綁死單一晶片:硬體被拆成 compute core、local memory、global memory 與 interconnect topology。以 Wormhole 為例,這些欄位就會對應到 Tensix grid、每個 core 的 L1、DRAM tile 和 2D torus NoC。
這類硬體的效能取決於 logical tile 放在哪個 physical core、資料是否能在 L1 留住、能否沿 NoC broadcast,以及何時需要回 DRAM。TileLoom 的 hardware representation 因此會記錄 interconnect topology、memory hierarchy 和 compute capability,讓 compiler 能比較不同 mapping。
這條線的重點在 triton-shared。它把 Triton code 轉成比較一般的 MLIR 組合,例如 affine、linalg、scf、memref。這些 dialect 比 TTGIR 更適合做傳統 compiler analysis,也比較容易接 TileLoom 自己的 dataflow planning。
TileLoom 後面的工作包括
TileLoom 和 Triton-Tenstorrent 都能接到 Tenstorrent,但關注範圍不同。 Triton-Tenstorrent 展示 Triton plugin 如何進入 TTMLIR / TTNN runtime,TileLoom 把 spatiotemporal mapping、reuse、broadcast 和 memory allocation 當成主要 研究問題,最後生成 TT-Metalium C API,這條線最適合本系列後面的 paper
與 IR 實戰,因為每個 dataflow planning 階段都有可觀察的表示。
Triton language / compiler / runtime
=> Triton-Ascend
- Ascend language extension
- compiler / driver / libdevice
=> TTIR (Triton IR)
=> Linalg IR
=> BiSheng Compiler
=> AscendNPU IR (HFusion Dialect / HIVM Dialect)
=> LLVM IR
=> Back-end code generation
=> triton_xxx_kernel.o
=> Triton-Ascend driver / CANN
=> Ascend NPU
官方圖可分成三塊來讀,最上層保留 Triton 共用的 compiler、language 與 runtime,中間層是 Triton-Ascend 加入的 language extension、compiler、driver 與 libdevice,下層再交給 BiSheng Compiler,從 AscendNPU IR 的 HFusion / HIVM dialect 繼續降層到 LLVM IR 與 backend code generation。

圖片出處:Triton-Ascend 官方文件〈[架構設計與核心特性](https://triton-ascend.readthedocs.io/zh-cn/latest/architecture_design_and_core_features.html#id2
Ascend NPU 由多個 AI Core 提供平行計算。從 operator 開發角度,
一個 AI Core 內可以看到三類主要 compute resource
記憶體部分會碰到 global memory、L1,以及 core 內的 Unified Buffer
(UB)。以 A2 系列文件為例,UB 是 192 KB,所以 Triton block 搬進單一
core 的 working set 太大時,compiler 會直接報 ub overflow。這使得
block size、multi-buffer 和資料搬移方式都必須配合硬體容量。
目前 Triton-Ascend 文件列出的硬體涵蓋 Atlas A2、A3 系列產品與 Ascend 950 系列。實際安裝仍要配合裝置型號、CANN 與 Triton-Ascend 版本。
這裡的路線是先把 Triton IR 轉成 Linalg IR,再由 BiSheng Compiler 接手 AscendNPU IR、LLVM IR 與 backend code generation,產生 device object triton_xxx_kernel.o。Triton-Ascend driver 接著經由 CANN 載入與啟動這個 device kernel。
可以先抓一個觀察,這條線也利用 Linalg 作為中繼層。Linalg 的好處是它把 tensor computation 表示成比較標準的 structured op,後端可以再針對自己的硬體 lowering。
AscendNPU IR 要進一步承接 AI Core 的 multi-core mapping、UB / L1 allocation、資料搬移,以及 Cube / Vector computation。和 TileLoom 相比,Ascend flow 的後半段銜接 BiSheng 與 CANN;TileLoom 則接 dataflow-aware MLIR 和 TT-Metalium。
Triton
=> TTIR
=> triton-shared
=> MLIR Transform dialect
(affine / linalg / scf / memref)
=> MLIR-AIR / MLIR-AIE
=> XRT binary (aie.xclbin)
XDNA / AIE 這條線也會經過 triton-shared,再進 affine / linalg / scf / memref。後面接 MLIR-AIR / MLIR-AIE,最後產出 XRT binary。
下圖把完整 MLIR-AIR stack 分成六段。Triton 在 frontend 經由
Triton-Shared 接入 MLIR community dialect,與 PyTorch / Torch-MLIR、
TensorFlow / TOSA 等入口匯流。SCF、MemRef 與 tiled Linalg 再降層到
MLIR-AIR,用 launch、segment / herd、channel 和 function call 表示
spatial execution 與 data movement。後端接到 MLIR-AIE,產生 NPU
configuration、DMA 與 runtime sequence,最後由 XRT / ROCr 載入 binary。

圖片出處:Erwei Wang et al.,〈From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR〉,Figure 2「MLIR-AIR stack overview」,原始 PDF 第 8 頁。論文採 CC BY 4.0授權。
Triton-XDNA 目前瞄準 AMD AIE2 與 AIE2P 架構,也就是 Ryzen AI NPU 使用的 AI Engine 路線。AI Engine 由空間排列的 tile array 組成,tile 包含 compute core 或 local memory,tile 之間經由 stream switch 傳資料,並由可程式化 DMA 管理資料搬移。
這個硬體形狀讓 compiler 必須決定三件事,運算切成多大的 tile、每份運算 放到哪個 AIE core、資料如何經由 stream 與 DMA 在 memory tile、compute tile 和 host 之間移動。
triton-shared 先把 Triton code 轉成較緊湊的 Linalg compute graph。
MLIR Transform dialect 再描述 tiling、bufferization 和 vectorization;
MLIR-AIR / MLIR-AIE 負責 spatial mapping、array connectivity 與 data
movement,最後產生 aie.xclbin,交由 XRT runtime 載入。
這條線和 TileLoom 都要處理 spatial / array 類硬體。TileLoom 的 target
是 Tenstorrent 的 Tensix grid;Triton-XDNA 接到 AMD AIE tile array 與
XRT。目前 Triton-XDNA README 將專案標為 experimental,已列出的 kernel
涵蓋 matmul、elementwise、softmax 和 layer normalization,硬體與 dtype
支援仍應以專案的 examples dashboard 為準。
Hexagon-MLIR flow 是
PyTorch => Torch-MLIR ------\
=> MLIR IR
Triton => Triton-to-Linalg -/
=> Hexagon-MLIR
- quantization / fusion / tiling / buffering
- multithreading / vectorization / math libraries
- lowering / optimization
=> LLVM IR
=> Hexagon-LLVM
=> Runtime
=> Qualcomm Hexagon NPU
PyTorch model 經由 Torch-MLIR 轉成 MLIR IR,Triton kernel 則經由
Triton-to-Linalg 轉入同一個 structured MLIR 區域。Hexagon-MLIR 在這個
共用中繼表示上做 quantization、fusion、tiling、buffering、multithreading 與 vectorization,後面再降層到 LLVM IR 與 Hexagon-LLVM。
Hexagon-MLIR 以 Qualcomm Hexagon NPU 為目標,這裡最重要的硬體資源包括
因此這條 flow 除了把 scalar operation 轉成 LLVM IR,還要處理資料能否留在 TCM、迴圈如何 tile 成 HVX 適合的寬度,以及 DDR / TCM transfer 能否和 compute 重疊。
Hexagon-MLIR 會先把 Triton code 轉到 Linalg 等 structured IR,再做 fusion、tiling、memory promotion 與 HVX vectorization。Fusion 可以減少 中間 tensor 寫回外部記憶體的次數,並把多個 operation 組成 locality 較好的 mega-kernel,memory pass 則安排 TCM 與 DMA。
這條線會藉由 linalg-hexagon-opt 等工具觀察與執行 Hexagon-MLIR lowering,最後走到 LLVM IR,再產生 Hexagon assembly 與 object code,交給相應 runtime 執行。它的形狀較接近傳統 compiler backend,但最佳化仍由 NPU 的 HVX、TCM 和 DMA 決定。
Hexagon-MLIR 同時接受 Triton code 與 PyTorch model。Figure 1 保留兩個
frontend,可以看到它們如何匯流到 MLIR IR;和本文其他 flow 比較時,
則先聚焦 Triton-to-Linalg 這條分支。
分享這個月剛出爐的論文
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
提供了一個更接近 production 的案例,Meta 把手寫 triton code 與 TorchInductor 產生的 triton code 部署到 MTIA-2i。
先看硬體。MTIA-2i 有 8 × 8 個 processing element(PE),PE 之間以客製化
NoC 相連,周圍再接 on-chip memory 與 LPDDR5 controller。

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators,
Figure 1(High-level architecture of MTIA-2i),arXiv:2608.00325v1(2026),
CC BY 4.0。

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators,
Figure 2: PE’s internal organization,arXiv:2608.00325v1(2026),
CC BY 4.0。
每個 PE 裡有兩顆 RISC-V processor core,以及多種 fixed-function unit
(FFU):DPE 做 GEMM、SE 做 vector / reduction、RE 累加 DPE 結果、MLU 做
layout transformation,FI 則用 DMA 搬移資料。長尾運算若沒有合適的 FFU,
可交給支援 RVV 的 RISC-V vector core。
MTIA-2i 的 local storage(LS)以 circular buffer(CB)管理。RISC-V core
非同步發出 FFU command,CB 的 read / write pointer 負責相依性。Compiler
因此要同時處理「一個 Triton op 應交給哪個 engine」、「tensor 放到哪個
CB」,以及「DMA 與 compute 怎麼 software pipeline」。
論文 Figure 4 把 pipeline 分成四段

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators,
Figure 4(High-level flow of Triton-MTIA Compiler),arXiv:2608.00325v1(2026),
CC BY 4.0。
把圖中的資訊展開後,可以寫成
手寫 triton kernel
→ upstream Triton parser(加上 MTIA language extension)
→ TTIR
→ Middle-end IR
├─ structured memory access → FI DMA
├─ block-wise compute → DPE / SE / MLU / RE
└─ 其他運算 → RISC-V core / RVV
→ Backend IR
├─ tensor bufferization
├─ LS / circular buffer allocation and reuse
└─ architecture-specific lowering
→ LLVM IR 或可讀的 C++
→ Clang compile and link
→ MTIA device binary
這條 flow 沒有沿用 GPU 的 TTGIR → NVVM / PTX 路線。團隊保留 upstream Triton parser 產生 TTIR,middle-end 再把 op 配到 MTIA 的硬體 engine,backend 才加入 LS、CB 與更低階的架構資訊。最後可產生 LLVM IR,也能產生容易對照 Triton source 的 C++,讓 kernel developer 修改與除錯後再交給 Clang。
MTIA 是理解 triton-shared 邊界的好例子。論文只明確指出它重用triton-shared 的 Triton IR pointer analysis
Triton tensor-of-pointers
→ 分析 pointer arithmetic
→ 可表示為 base + shape + strides?
├─ 可以:改寫成 structured access,交給 DMA
└─ 不行:由 RISC-V core 執行 scalar load / store
DMA 的頻寬利用率較高,但只接受 structured memory access。這個分析結果會直接影響一段原本為 GPU 寫的 Triton pointer code,到了 MTIA 後能不能走快的DMA path。它也呼應 triton-shared README 中的限制:pointer analysis 能處理strided pattern,尚未涵蓋任意 memory access。
這裡也要避免過度解讀,論文沒有說 MTIA backend 直接採用 triton-shared 的完整 Triton-to-Linalg pipeline,可以確認的重用範圍是 pointer analysis。MTIA 的 middle-end、backend dialect 與 production runtime 細節也沒有在開源 repository 中完整提供。
這篇 systems paper 的研究問題是,Triton 這種原本以 GPU 為中心的 kernel DSL,能否支援程式模型差異很大的 custom ML accelerator,並達到 production 所需的 coverage 與 performance?它提出 MTIA backend、TorchInductor codegen 調整和少量 MTIA language extension,再用 MTIA-2i silicon 與 production deployment 回答。
論文摘要報告的最終 production footprint 是約 60 種 model type、50% 的 layer,以及 47% 的 non-GEMM execution time。Figure 12 顯示從 2025 Q1 到 Q2,以 layer 與 runtime 計算的 coverage 都大約增加一倍。

圖片來源:Haishan Zhu et al.,
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators,
Figure 12(Triton’s production footprint in models on MTIA),
arXiv:2608.00325v1(2026),
CC BY 4.0
閱讀 Figure 12 時要留意,圖中的分母排除 GEMM,不能解讀成 Triton 已涵蓋整個 model execution time。
移植經驗也很具體,16 個原本跑在 GPU 的 triton code 中,11 個可直接移植
並提供足夠效能,另外 5 個需要修改 pointer arithmetic、numeric handling 或
MTIA-specific optimization。這組資料支持 Triton source portability,也顯示
高效能 kernel 不會自動跨硬體保持同一份寫法。
這份結果來自 Meta 內部 MTIA software stack、模型與 production workload,
compiler backend 也沒有完整開源。它能證明 Triton-MTIA 已大規模部署,無法讓
外部讀者獨立重現全部數字。文章中提到的 GEMM 超過 80% architecture roofline、
coverage 與 speedup,都應保留論文限定的 shape、model 與 MTIA-2i 條件。
可以把這六條 flow 放成一張表
| Flow | 目標硬體 | 硬體映射重點 | 最後接到哪裡 |
|---|---|---|---|
| Triton-Tenstorrent TTMLIR | Tensix core grid | L1、NoC、system topology | TTMLIR / TTNN runtime、Flatbuffer |
| TileLoom | Spatial dataflow accelerator;實驗使用 Tenstorrent | core mapping、reuse、broadcast、memory allocation | TT-Metalium C API |
| Triton-Ascend | Ascend AI Core | Cube / Vector、UB / L1、multi-core mapping | BiSheng object、CANN |
| Triton-XDNA | AMD AIE2 / AIE2P tile array | tiling、core placement、stream / DMA | aie.xclbin、XRT |
| Hexagon-MLIR | Qualcomm Hexagon NPU | fusion、HVX、TCM、DMA | LLVM object、Hexagon runtime |
| Triton-MTIA | MTIA-2i 8 × 8 PE array | op-to-engine mapping、LS / CB、software pipelining、PID-to-PE mapping | LLVM IR 或 C++、Clang、MTIA device binary |
共同問題像是
第一,它和 Tenstorrent 硬體形狀很貼近。Wormhole 是 core grid + local memory + NoC,TileLoom 的 paper 也正是在處理 spatiotemporal mapping 和 data movement planning。
第二,它保留了 Triton 作為前端。這讓我們可以從比較熟悉的 triton code 出發,看 compiler 怎麼把 tile-level program 變成 dataflow-aware program。
第三,它有公開的 mm IR 階段,可以拿來逐步對照。後面的實戰篇會用 test/Passes/mm/IR 裡的九個 .mlir 檔案,看一個 matmul 如何從 frontend IR 逐步變成接近 Tenstorrent backend 的表示。
今天先把幾條 Triton flow 排在一起
affine / linalg / scf / memref 常被拿來當可分析、可轉換的 MLIR 中繼層。aie.xclbin。明天開始讀 TileLoom paper,好奇 spatial dataflow accelerator 會把哪些 mapping 與 data movement 決策交給 compiler?