Skip to main navigation Skip to search Skip to main content

From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization

  • Shuaijiang Li
  • , Jiacheng Zhao
  • , Ying Liu
  • , Shuoming Zhang
  • , Lei Chen
  • , Yijin Li
  • , Yangyu Zhang
  • , Zhicheng Lin
  • , Runyu Zhou
  • , Xiyu Shi
  • , Chunwei Xia
  • , Yuan Wen
  • , Xiaobing Feng
  • , Huimin Cui
  • University of Chinese Academy of Sciences
  • Chinese Academy of Sciences
  • University of Leeds

Research output: Chapter in Book/Report/Conference proceedingPublished conference contribution

Abstract

CUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators.
Original languageEnglish
Title of host publication2026 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
EditorsStephen M. Blackburn, Albert Cohen, Timothy M. Jones
Place of PublicationPassau, Germany
PublisherIEEE Explore
Pages362-374
Number of pages13
ISBN (Electronic)979-8-3315-9288-2
ISBN (Print)979-8-3315-9289-9
DOIs
Publication statusPublished - 23 Feb 2026
EventIEEE/ACM International Symposium on Code Generation and Optimization - Sydney, Australia
Duration: 30 Jan 20264 Feb 2026
https://2026.cgo.org/

Conference

ConferenceIEEE/ACM International Symposium on Code Generation and Optimization
Abbreviated titleCGO 2026
Country/TerritoryAustralia
CitySydney
Period30/01/264/02/26
Internet address

Funding

We thank all the reviewers for their valuable comments and suggestions. This work was supported in part by the National Key R&D Program of China, Grant No.2023YFB3001502, the Jiangsu Provincial Key R&D Program, Grant No.BG2024028 and the National Natural Science Foundation of China, Grant No.U23B2020, No.62090024, No.62302479, No.62232015.

FundersFunder number
National Key Research and Development Program of China2023YFB3001502
National Natural Science Foundation of ChinaU23B2020, 62090024, 62302479, 62232015

    Keywords

    • MLIR
    • CIDA
    • auto-vectorization
    • heterogeneous computing
    • AI accelerators

    Fingerprint

    Dive into the research topics of 'From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization'. Together they form a unique fingerprint.

    Cite this