Tile Language (tile-lang) is a Pythonic domain‑specific language that compiles high‑performance kernels for GPU, CPU and NPU targets via a TVM‑based compiler infrastructure. Recent work added native support for Huawei Ascend 950 NPUs, enabling automatic scheduling, synchronization and SIMD/SIMT vector code generation. An open‑sourced Language Server Protocol implementation supplies inlay hints for buffer shapes, dtypes and scopes, hover details and precise diagnostics. Release v0.1.13 introduced a multi‑backend dialect, source‑location reporting in compiler errors, new CUDA and Metal hardware paths and a suite of correctness fixes while retiring several legacy APIs. Additional features include an SM120 NVF4 block‑scaled MMA path for Blackwell, Metal 4 cooperative‑tensor GEMM for Apple M5 (with a simdgroup fallback), source‑aware diagnostics carried into TIRX, block‑causal attention examples for diffusion language models, an IR lower‑trace debugging tool, adaptive launch‑width selection for DeepSeek V3.2 sparse MLA backward, and a top‑k optimizer that yields ≈1.9× speedup in reported benchmarks. The project also added compiler‑pass profiling, IKET profiler integration, LLVM CPU lowering, a tile‑scheduler, backend registry, pass visualizer, cross‑host CUDA binary cache and arbitrary‑layout TMA lowering. Benchmarks show strong MLA decoding and FlashAttention performance on H100, competitive GEMM throughput on RTX 4090, A100, H100 and MI300X, and efficient dequantized matmul on A100.

Read original