# Benchmark the LiteRT model

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/)
- [Explore LiteRT, XNNPACK, KleidiAI, and SME2](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/1-litert-kleidiai-sme2/)
- [Create LiteRT models](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/2-build-model/)
- [Build the LiteRT benchmark tool](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/3-build-tool/)
- [Benchmark the LiteRT model](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/4-benchmark/)
- [Next Steps](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/litert-sme/_next-steps/)

## Prerequisites for Benchmarking on SME2
Before you begin benchmarking LiteRT models on an SME2-capable Android device, make sure you have the following components prepared:

Once you have:
- A LiteRT model (for example, `fc_fp32.tflite`)
- The `benchmark_model` binary built with and without KleidiAI and Scalable Matrix Extension version 2 (SME2)

you can run benchmarks directly on an SME2-capable Android device.

## Verify SME2 support on the device
First, check if your Android device supports SME2.  
On the device (via Android Debug Bridge, ADB shell), run:

```bash
cat /proc/cpuinfo
```

Look for a `Features` line similar to:
```
processor	: 7
BogoMIPS	: 2000.00
Features	: fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm ssbs sb dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh bti mte ecv afp mte3 sme smei8i32 smef16f32 smeb16f32 smef32f32 wfxt rprfm sme2 smei16i32 smebi32i32 hbc lrcpc3
```
If you see `sme2` in the features, your CPU supports SME2.

## Run benchmark_model on an SME2 core
Next, run the benchmark tool and bind execution to a core that supports SME2.  
For example, to pin to CPU 7, use a single thread, and run enough iterations for stable timing:

```bash
taskset 80 ./benchmark_model --graph=./fc_fp32.tflite --num_runs=1000 --num_threads=1 --use_cpu=true --use_profiler=true
```

This command uses `taskset` to run the benchmark on core 7, sets `--num_threads=1`, and runs 1000 inferences. The `--use_profiler=true` flag enables operator-level profiling.

You should see output similar to:

```
INFO: [litert/runtime/accelerators/auto_registration.cc:148] CPU accelerator registered.
INFO: [litert/runtime/compiled_model.cc:415] Flatbuffer model initialized directly from incoming litert model.
INFO: Initialized TensorFlow Lite runtime.
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
VERBOSE: Replacing 1 out of 1 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0.
INFO: The input model file size (MB): 3.27774
INFO: Initialized session in 4.478ms.
INFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.
INFO: count=1055 first=1033 curr=473 min=443 max=1033 avg=465.319 std=18 p5=459 median=463 p95=478
...
```

You will see the time spent on model initialization, warm-up, and inference, as well as memory usage. With the profiler enabled, the output also reports the execution time of each operator.

Because the model contains only a single fully connected layer, the node type `Fully Connected (NC, PF32) GEMM` shows the average execution time and its percentage of total inference time.

**Note**: To verify that KleidiAI SME2 micro-kernels are invoked for the FullyConnected operator during inference, run `simpleperf record -g -- <workload>` to capture the calling graph. If you’re using the `benchmark_model`, build it with the `-c dbg` option.

## Measure the performance impact of KleidiAI SME2 micro-kernels
To compare the performance of the KleidiAI SME2 implementation with the original XNNPACK implementation, run the `benchmark_model` tool without KleidiAI enabled using the same parameters:

```bash
taskset 80 ./benchmark_model --graph=./fc_fp32.tflite --num_runs=1000 --num_threads=1 --use_cpu=true --use_profiler=true
```

The output should look like:

```
INFO: [litert/runtime/accelerators/auto_registration.cc:148] CPU accelerator registered.
INFO: [litert/runtime/compiled_model.cc:415] Flatbuffer model initialized directly from incoming litert model.
INFO: Initialized TensorFlow Lite runtime.
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
VERBOSE: Replacing 1 out of 1 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0.
INFO: The input model file size (MB): 3.27774
INFO: Initialized session in 4.488ms.
INFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.
...
```

You should notice significant throughput uplift and speedup in inference time when KleidiAI SME2 micro-kernels are enabled.

### Interpreting node type names for KleidiAI
For the same model, the XNNPACK node type name is different. For the non-KleidiAI implementation, the node type is `Fully Connected (NC, F32) GEMM`, whereas for the KleidiAI implementation, it is `Fully Connected (NC, PF32) GEMM`.

For other operators supported by KleidiAI, the per-operator profiling node types differ between the implementations with and without KleidiAI enabled in XNNPACK as follows:

| Operator                                   | Node Type (KleidiAI Enabled)    | Node Type (KleidiAI Disabled)  |
|--------------------------------------------|----------------------------------|---------------------------------|
| Fully Connected / Conv2D (Pointwise)      | Fully Connected (NC, PF32)       | Fully Connected (NC, F32)       |
| Fully Connected                             | Dynamic Fully Connected (NC, PF32)| Dynamic Fully Connected (NC, F32)|
| Fully Connected / Conv2D (Pointwise)      | Fully Connected (NC, PF16)       | Fully Connected (NC, F16)       |
| Fully Connected                             | Dynamic Fully Connected (NC, PF16)| Dynamic Fully Connected (NC, F16)|
| Fully Connected                             | Fully Connected (NC, QP8, F32, QC4W)| Fully Connected (NC, QD8, F32, QC4W)|
| Fully Connected / Conv2D (Pointwise)      | Fully Connected (NC, QP8, F32, QC8W)| Fully Connected (NC, QD8, F32, QC8W)|
| Fully Connected / Conv2D (Pointwise)      | Fully Connected (NC, PQS8, QC8W) | Fully Connected (NC, QS8, QC8W) |
| Conv2D                                     | Convolution (NHWC, PF32)        | Convolution (NHWC, F32)         |
| Conv2D                                     | Convolution (NHWC, PF16)        | Convolution (NHWC, F16)         |
| Conv2D                                     | Convolution (NHWC, PQS8, QS8, QC8W)| Convolution (NHWC, QC8)         |
| TransposeConv                              | Deconvolution (NHWC, PQS8, QS8, QC8W)| Deconvolution (NC, QS8, QC8W)   |
| Batch Matrix Multiply                       | Batch Matrix Multiply (NC, PF32) | Batch Matrix Multiply (NC, F32)  |
| Batch Matrix Multiply                       | Batch Matrix Multiply (NC, PF16) | Batch Matrix Multiply (NC, F16)  |
| Batch Matrix Multiply                       | Batch Matrix Multiply (NC, QP8, F32, QC8W)| Batch Matrix Multiply (NC, QD8, F32, QC8W)|

The letter “P” in the node type indicates a KleidiAI implementation.

For example, `Convolution (NHWC, PQS8, QS8, QC8W)` represents a Conv2D operator computed by a KleidiAI micro-kernel, where the tensor is in NHWC layout, the input is packed INT8 quantized, the weights are per-channel INT8 quantized, and the output is INT8 quantized.

By comparing `benchmark_model` runs with and without KleidiAI and SME2, and inspecting the profiled node types (PF32, PF16, QP8, PQS8), you can confirm that LiteRT is dispatching to SME2-optimized KleidiAI micro-kernels and quantify their performance impact on your Android device.

## What you’ve accomplished and what’s next
In this section, you learned how to benchmark LiteRT models on SME2-capable Android devices, verify SME2 support, and interpret performance results with and without KleidiAI SME2 micro-kernels. You also discovered how to identify which micro-kernels are used during inference by examining node type names in the profiler output.

You are now ready to further analyze performance, experiment with different models, or explore advanced profiling and optimization techniques for on-device AI with LiteRT, XNNPACK, and KleidiAI. Continue to the next section to deepen your understanding or apply these skills to your own projects.
