Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Who is this for?
This is an advanced topic for software developers and performance engineers who want to integrate a KleidiAI Scalable Matrix Extension 2 (SME2) microkernel into an existing AI inference framework.
What will you learn?
Upon completion of this Learning Path, you will be able to:
- Identify the XNNPACK qd8_f16_qc4w operand formats and select a KleidiAI microkernel with a matching quantization contract.
- Pack XNNPACK QD8 activations and QC4W weights into the layouts required by a KleidiAI SME2 kernel.
- Add runtime SME2 dispatch with a safe fallback path.
- Build and validate the integration on an Arm-based Android device.
Prerequisites
Before starting, you will need the following:
- Familiarity with C or C++ and basic matrix multiplication
- Android Debug Bridge (
adb) and an Arm-based Android device with SME2 support
Summary
This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
qd8_f16_qc4w fully connected operator. First, you’ll examine operand formats and select a kernel with matching quantization. Next, you’ll examine how static QC4W weights and dynamic QD8 activations are packed while preserving their numerical contracts. You’ll then inspect how SME2 runtime dispatch is configured with the existing XNNPACK fallback. Finally, you’ll build the Android test binary, run correctness tests on an SME2 device, and verify the fallback build.Frequently asked questions
These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
f16 output, a qai8dxp left-hand side with asymmetric int8 quantization per row, a qsi4cxp right-hand side with signed int4 quantization per channel, and sme2_mopa instructions. Reject the kernel if any part of that contract differs.matmul_clamp_f16_qsi8d32p_qai4c32p because it requires symmetric int8 quantization for each block of 32 K values. XNNPACK QD8 uses asymmetric int8 quantization for each row, so you’d have to dequantize and requantize the left-hand side. Instead, select the qai8dxp and qsi4cxp kernel that matches the existing operand contract.xnn_create_fully_connected_nc_qd8_f16_qc4w. The packed int4 weights, weight sums, per-channel scales, and bias are stored in operator memory or the XNNPACK weights cache so that you can reuse them across inference runs.pack_lhs and lhs_packed_size helpers added by patch 0003-pack-qd8-lhs-for-kai-sme2.patch in src/qd8-f16-qc4w-gemm/qd8-f16-qc4w-gemm-minmax-16x64c4-neonsme2.c. The patch preserves the QD8 values, stores each row’s negative zero point and scale, and pads K with that row’s zero point. Patch 0004-dispatch-qd8-f16-qc4w-through-kai-sme2.patch queries the kernel parameters and calls these helpers at runtime.FULLY_CONNECTED_NC_QD8_F16_QC4W correctness suite on an SME2 device and confirm that all 15 tests pass. Separately, build the //:packing target with xnn_enable_kleidiai=false to verify that the compile-time guards preserve the non-KleidiAI configuration.