# [Run Vision LLM inference on Android with KleidiAI and MNN](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/)
- [Background](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/background/)
- [Environment setup and prepare model](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/1-devenv-and-model/)
- [Benchmark the Vision Transformer performance with KleidiAI](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/2-generate-apk/)
- [Build the MNN Command-line ViT Demo](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/3-benchmark/)
- [Next Steps](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/vision-llm-inference-on-android-with-kleidiai-and-mnn/_next-steps/)

## About this Learning Path

| Skill level:     | Introductory                 |
|------------------|------------------------------|
| Reading time:    | 30 min                       |
| Last updated:    | 21 Aug 2026                  |

| Authors:                       | Shuheng Deng, Arm<br/>Yiyang Fan, Arm                                    |
|--------------------------------|-------------------------------------------------------------------------|
| Arm IP:                        | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors)|
| Tags:                          | ML, Android, Android Studio, KleidiAI                                   |

### Who is this for?
This Learning Path is for developers who want to run Vision Transformers (ViT) efficiently on Android.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Download a Vision Large Language Model (LLM) from Hugging Face.
- Convert the model to the Mobile Neural Network (MNN) framework.
- Install an Android demo application using the model to run an inference.
- Compare inference performance with and without KleidiAI Arm-optimized micro-kernels.

### Prerequisites
Before starting, you will need the following:
- A development machine with [Android Studio](https://developer.android.com/studio) installed.
- A smartphone running Android with support for `i8mm` and `dotprod` instructions.

### Summary
You’ll run the Qwen2.5-VL-3B-Instruct-MNN vision model on an Android device with MNN and KleidiAI. First, you’ll install the Android tools, download a pre-quantized MNN model, and build the Android Studio demo. Then, you’ll prepare an image, build and run the MNN command-line demo, enable KleidiAI, rebuild the binaries, and compare the reported benchmark timings.

### Frequently asked questions
<details>
<summary>Which NDK and CMake do I need, and how do I install them?</summary>
To match the tested setup, use Android NDK `28.0.12916984` and CMake `4.0.0-rc1`. In Android Studio, select **Tools > SDK Manager**, open **SDK Tools**, then select **NDK (Side by side)** and **CMake**. On Ubuntu or Debian, install `cmake` and `git-lfs` with the provided command.
</details>

<details>
<summary>How do I clone and open the project in Android Studio?</summary>
Run `git clone https://gitlab.arm.com/kleidi/kleidi-examples/vision-language-models`. In Android Studio, select **File > Open**, choose the `vision-language-models` directory, and select **Open**. Android Studio then builds the project.
</details>

<details>
<summary>Where should I put the example image and what name should it have?</summary>
Rename the image to `example.png`, then run `adb push example.png /data/local/tmp/` to copy it to your device.
</details>

<details>
<summary>How do I compare inference with and without KleidiAI micro-kernels?</summary>
Build and run the MNN command-line demo without KleidiAI first. Then add the `CPU_ENABLE_KLEIDIAI` runtime hint, rebuild the binaries, replace them on your device, and run the same inference command again to compare the reported timings.
</details>

<details>
<summary>When is the model converted to MNN, and how do I know it worked?</summary>
Model conversion is optional. By default, you download the pre-quantized `Qwen2.5-VL-3B-Instruct-MNN` model. If you convert the model with `llmexport`, verify that the `Qwen2.5-VL-3B-Instruct-convert-4bit-64qblock` directory is at least 2 GB.
</details>
