Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Integrate a KleidiAI SME2 kernel into XNNPACK
Introduction
Prepare the XNNPACK baseline
Understand XNNPACK QD8 and QC4W operand formats
Select a compatible KleidiAI microkernel
Inspect QC4W weight packing for the KleidiAI SME2 kernel
Inspect QD8 activation packing without requantizing for KleidiAI
Inspect SME2 kernel configuration and dispatch
Build and validate the integration
Next Steps
Start with the math
The fully connected operator computes the following:
C[M,N] = A[M,K] x B[N,K]^T + bias[N]
The variables are as follows:
Mis the number of input rows or tokens.Kis the input feature dimension, also called the reduction dimension.Nis the output feature dimension.Ais the activation matrix.Bis the weight matrix.Cis the FP16 output matrix.
Although the logical right-hand side (RHS) matrix is written as B[N,K], the multiplication uses its transpose. Each output has one row of weights with K values.
Left-hand side: QD8 activation data
XNNPACK stores the activation matrix as signed 8-bit values. Each row has separate dynamic quantization parameters:
struct xnn_qd8_quantization_params {
int32_t zero_point;
float inv_scale;
};
The real value represented by one element is:
A_real[m,k] = (A_q[m,k] - zero_point[m]) * inv_scale[m]
This is asymmetric, per-row quantization:
- The int8 values are stored row-major with the input stride.
- Each row can have a different zero point.
- Each row can have a different scale.
The important point is that this left-hand side (LHS) QD8 representation already contains the quantized activation values. A framework integration should preserve these values whenever possible.
Right-hand side: QC4W weights
The XNNPACK QC4W input consists of:
- Packed 4-bit weights
- One FP32 scale for each output channel
- A kernel zero point of
0or8 - Optional FP32 bias values
Two 4-bit weights occupy one byte:
bits [3:0] = K element 0
bits [7:4] = K element 1
For the normal weight layout, the raw tensor is N x K. Each output channel has a contiguous packed row.
XNNPACK also supports XNN_FLAG_TRANSPOSE_WEIGHTS. In that case, the source tensor is K x N. The integration must therefore convert the tensor to the N x K form expected by the selected KleidiAI RHS packer.
Safe K padding
The selected KleidiAI Scalable Matrix Extension 2 (SME2) microkernel rounds its internal K dimension up to a multiple of 32. The original model doesn’t need to have a K dimension that’s divisible by 32.
The integration creates a padded representation:
K_padded = round_up(K, 32)
Padding must represent real value zero:
- For the QD8 LHS, each row is padded with its own
zero_point. - For signed int4 weights, the padding is zero.
- For unsigned int4 weights with zero point 8, the RHS packer converts padding to signed zero.
Padding with any other value changes the dot product and produces incorrect output.
Where the data enters XNNPACK
The public operator entry points are:
xnn_create_fully_connected_nc_qd8_f16_qc4w(...)
xnn_reshape_fully_connected_nc_qd8_f16_qc4w(...)
xnn_setup_fully_connected_nc_qd8_f16_qc4w(...)
The create step receives the static weights, scales, and bias, and packs the persistent RHS representation. The reshape step receives the runtime batch size and plans the general matrix multiplication (GEMM) tiles, parallel work, and any required workspace. The setup step receives the QD8 activation values, the FP16 output buffer, and the per-row xnn_qd8_quantization_params array.
What you’ve learned and what’s next
You’ve identified the per-row QD8 activation parameters, per-channel QC4W weight data, and zero-valued padding requirements. These define the operand contract that the KleidiAI integration must preserve.
Next, you’ll use these facts to examine the choice of KleidiAI microkernel.