Who is this for?

This is an advanced topic for software developers and performance engineers who want to integrate a KleidiAI Scalable Matrix Extension 2 (SME2) microkernel into an existing AI inference framework.

What will you learn?

Upon completion of this Learning Path, you will be able to:

  • Identify the XNNPACK qd8_f16_qc4w operand formats and select a KleidiAI microkernel with a matching quantization contract.
  • Pack XNNPACK QD8 activations and QC4W weights into the layouts required by a KleidiAI SME2 kernel.
  • Add runtime SME2 dispatch with a safe fallback path.
  • Build and validate the integration on an Arm-based Android device.

Prerequisites

Before starting, you will need the following:

  • Familiarity with C or C++ and basic matrix multiplication
  • Android Debug Bridge (adb) and an Arm-based Android device with SME2 support

Summary

AI-assisted

This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
You’ll integrate a KleidiAI SME2 microkernel into XNNPACK’s qd8_f16_qc4w fully connected operator. First, you’ll examine operand formats and select a kernel with matching quantization. Next, you’ll examine how static QC4W weights and dynamic QD8 activations are packed while preserving their numerical contracts. You’ll then inspect how SME2 runtime dispatch is configured with the existing XNNPACK fallback. Finally, you’ll build the Android test binary, run correctness tests on an SME2 device, and verify the fallback build.

Frequently asked questions

AI-assisted

These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
How do I read the KleidiAI kernel name to verify that it matches qd8_f16_qc4w?
Read the name as a description of the output type, operand formats, and instruction family. For this operator, confirm f16 output, a qai8dxp left-hand side with asymmetric int8 quantization per row, a qsi4cxp right-hand side with signed int4 quantization per channel, and sme2_mopa instructions. Reject the kernel if any part of that contract differs.
Which kernel should I avoid, and why is it incompatible?
Avoid a kernel such as matmul_clamp_f16_qsi8d32p_qai4c32p because it requires symmetric int8 quantization for each block of 32 K values. XNNPACK QD8 uses asymmetric int8 quantization for each row, so you’d have to dequantize and requantize the left-hand side. Instead, select the qai8dxp and qsi4cxp kernel that matches the existing operand contract.
When should I pack QC4W weights, and what gets stored?
Pack the QC4W right-hand side once during xnn_create_fully_connected_nc_qd8_f16_qc4w. The packed int4 weights, weight sums, per-channel scales, and bias are stored in operator memory or the XNNPACK weights cache so that you can reuse them across inference runs.
How do I pack the QD8 activations without requantizing?
Use the pack_lhs and lhs_packed_size helpers added by patch 0003-pack-qd8-lhs-for-kai-sme2.patch in src/qd8-f16-qc4w-gemm/qd8-f16-qc4w-gemm-minmax-16x64c4-neonsme2.c. The patch preserves the QD8 values, stores each row’s negative zero point and scale, and pads K with that row’s zero point. Patch 0004-dispatch-qd8-f16-qc4w-through-kai-sme2.patch queries the kernel parameters and calls these helpers at runtime.
How do I validate the SME2 integration and the fallback build?
Build the Android test binary with KleidiAI enabled, then run the filtered FULLY_CONNECTED_NC_QD8_F16_QC4W correctness suite on an SME2 device and confirm that all 15 tests pass. Separately, build the //:packing target with xnn_enable_kleidiai=false to verify that the compile-time guards preserve the non-KleidiAI configuration.
Next