Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Check the patch series
With all four patches applied, run the commands from the XNNPACK checkout directory on your development host. You’ll build on the host, then use adb to run the correctness tests on the Android device with Scalable Matrix Extension 2 (SME2) support.
Confirm that the patches were applied to the baseline XNNPACK revision in the expected order:
git log --oneline -4
The output is similar to:
Dispatch QD8 F16 QC4W through KAI SME2
Pack QD8 LHS for KAI SME2
Support transposed KAI QC4W weights
Prepare QD8 F16 QC4W SME2 kernel
The example output lists commit subjects. Your output will also include commit identifiers.
Regenerate the microkernel source lists:
python3 tools/update-microkernels.py
git diff --check
Depending on your Python version, you might see DeprecationWarning: codecs.open() is deprecated. Use open() instead. from the generator. This warning alone doesn’t indicate a failure. Continue if the script completes successfully without duplicate microkernel messages and git diff --check passes.
Build the test binary for Android
To build the test binary for Android, install the Android Native Development Kit (Android NDK) r29.
From the XNNPACK checkout, set WORKSPACE to its parent directory so that the NDK is installed alongside XNNPACK:
export WORKSPACE="$(dirname "$PWD")"
Download and extract the NDK for your development host:
cd "$WORKSPACE"
wget https://dl.google.com/android/repository/android-ndk-r29-linux.zip
unzip android-ndk-r29-linux.zip
cd "$WORKSPACE"
wget https://dl.google.com/android/repository/android-ndk-r29-darwin.zip
unzip android-ndk-r29-darwin.zip
Both archives extract to android-ndk-r29. Set ANDROID_NDK to that directory, then return to the XNNPACK checkout and build with KleidiAI and tests enabled:
export ANDROID_NDK="$WORKSPACE/android-ndk-r29"
cd "$WORKSPACE/XNNPACK"
scripts/build-android-arm64.sh \
-DXNNPACK_ENABLE_KLEIDIAI=ON \
-DXNNPACK_BUILD_BENCHMARKS=OFF \
-DXNNPACK_BUILD_TESTS=ON
Build the fully connected operator test:
cmake --build build/android/arm64-v8a \
--target fully-connected-nc-test -- parallel
Run the correctness tests on an SME2 device
From the development host, copy the test binary to an Android device with SME2 support . Make the binary executable, and run the filtered correctness suite:
adb push build/android/arm64-v8a/test/operators/fully-connected-nc-test \
/data/local/tmp/xnnpack-fc-test
adb shell "chmod 755 /data/local/tmp/xnnpack-fc-test"
adb shell "/data/local/tmp/xnnpack-fc-test \
--gtest_filter='FULLY_CONNECTED_NC_QD8_F16_QC4W.*'"
The expected output is:
[==========] Running 15 tests from 1 test suite.
[ PASSED ] 15 tests.
The suite covers the following:
- Normal and small batches
- Minimum and maximum clamp ranges
- Input and output stride
- Optional bias
- Transposed weights
- Weights-cache reuse
Check the fallback build
Ensure that the KleidiAI path doesn’t break builds where KleidiAI is disabled. On a development host, run:
bazel build //:packing --define=xnn_enable_kleidiai=false
This validates that the XNN_ENABLE_KLEIDIAI guards preserve the non-KleidiAI configuration.
What you’ve accomplished
You’ve built the patched Android test binary, passed the filtered QD8 F16 QC4W correctness suite, and built the packing target with KleidiAI disabled.
You can extend the workflow and patches to integrate a KleidiAI SME2 microkernel for your own use case into an existing AI inference framework. Explore the Understand KleidiAI SME2 matmul microkernels and KleidiAI on Android with MediaPipe and XNNPACK Learning Paths.
For workloads that split a single activation matrix across many N tiles, consider a workspace-based LHS pre-pack stage:
Pack LHS once per run
-> reuse packed LHS across all RHS N tiles
Keep the following rules when optimizing:
- Preserve the QD8 per-row zero point and scale.
- Pad K with the quantized zero point.
- Query KleidiAI tile dimensions instead of hard-coding vector-length-dependent values.
- Keep the non-SME2 XNNPACK fallback.