What problem does it solve? Optimizing CATLASS matmul kernels on Ascend NPUs requires choosing the right kernel family, DispatchPolicy, tile shapes, and swizzle parameters, and wrong choices lead to idle AIC cores, L1 overflow, or no measurable speedup. ## Core Features & Use Cases - Kernel Family Selection: Decision table mapping M/N/K shape, alignment, and epilogue needs to BasicMatmul, MatmulEpilogue, padding, Split-K, or Preload kernel families. - Tile and Load-Balance Tuning: Guidance for adjusting L1TileShape and L0TileShape so block counts match AIC core counts without exceeding L1/L0 capacity. - Swizzle and DispatchPolicy Configuration: Rules for GemmIdentityBlockSwizzle direction and offset, plus when MmadAtlasA2Preload beats Pingpong. - Use Case: When an AR task shows profile time dominated by Transdata instead of MMAD, the guide directs edits to catlass_torch.cpp to remove redundant npu_format_cast calls rather than wasting cycles on tile changes. ## Quick Start Ask the agent to tune the CATLASS matmul kernel in catlass_kernel.asc for your reference shape by adjusting DispatchPolicy, tile shapes, and swizzle, then verify compilation and profile results.