Start with the math

The fully connected operator computes the following:

    

        
        
C[M,N] = A[M,K] x B[N,K]^T + bias[N]

    

The variables are as follows:

  • M is the number of input rows or tokens.
  • K is the input feature dimension, also called the reduction dimension.
  • N is the output feature dimension.
  • A is the activation matrix.
  • B is the weight matrix.
  • C is the FP16 output matrix.

Although the logical right-hand side (RHS) matrix is written as B[N,K], the multiplication uses its transpose. Each output has one row of weights with K values.

Left-hand side: QD8 activation data

XNNPACK stores the activation matrix as signed 8-bit values. Each row has separate dynamic quantization parameters:

    

        
        
struct xnn_qd8_quantization_params {
  int32_t zero_point;
  float inv_scale;
};

    

The real value represented by one element is:

    

        
        
A_real[m,k] = (A_q[m,k] - zero_point[m]) * inv_scale[m]

    

This is asymmetric, per-row quantization:

  • The int8 values are stored row-major with the input stride.
  • Each row can have a different zero point.
  • Each row can have a different scale.

The important point is that this left-hand side (LHS) QD8 representation already contains the quantized activation values. A framework integration should preserve these values whenever possible.

Right-hand side: QC4W weights

The XNNPACK QC4W input consists of:

  • Packed 4-bit weights
  • One FP32 scale for each output channel
  • A kernel zero point of 0 or 8
  • Optional FP32 bias values

Two 4-bit weights occupy one byte:

    

        
        
bits [3:0] = K element 0
bits [7:4] = K element 1

    

For the normal weight layout, the raw tensor is N x K. Each output channel has a contiguous packed row.

XNNPACK also supports XNN_FLAG_TRANSPOSE_WEIGHTS. In that case, the source tensor is K x N. The integration must therefore convert the tensor to the N x K form expected by the selected KleidiAI RHS packer.

Safe K padding

The selected KleidiAI Scalable Matrix Extension 2 (SME2) microkernel rounds its internal K dimension up to a multiple of 32. The original model doesn’t need to have a K dimension that’s divisible by 32.

The integration creates a padded representation:

    

        
        
K_padded = round_up(K, 32)

    

Padding must represent real value zero:

  • For the QD8 LHS, each row is padded with its own zero_point.
  • For signed int4 weights, the padding is zero.
  • For unsigned int4 weights with zero point 8, the RHS packer converts padding to signed zero.

Padding with any other value changes the dot product and produces incorrect output.

Where the data enters XNNPACK

The public operator entry points are:

    

        
        
xnn_create_fully_connected_nc_qd8_f16_qc4w(...)
xnn_reshape_fully_connected_nc_qd8_f16_qc4w(...)
xnn_setup_fully_connected_nc_qd8_f16_qc4w(...)

    

The create step receives the static weights, scales, and bias, and packs the persistent RHS representation. The reshape step receives the runtime batch size and plans the general matrix multiplication (GEMM) tiles, parallel work, and any required workspace. The setup step receives the QD8 activation values, the FP16 output buffer, and the per-row xnn_qd8_quantization_params array.

What you’ve learned and what’s next

You’ve identified the per-row QD8 activation parameters, per-channel QC4W weight data, and zero-valued padding requirements. These define the operand contract that the KleidiAI integration must preserve.

Next, you’ll use these facts to examine the choice of KleidiAI microkernel.

Back
Next