# Build an Android chat app with Llama, KleidiAI, ExecuTorch, and XNNPACK

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/)
- [Create a development environment](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/1-dev-env-setup/)
- [ExecuTorch Setup](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/2-executorch-setup/)
- [Understanding Llama models](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/3-understanding-llama-models/)
- [Prepare Llama models for ExecuTorch](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/4-prepare-llama-models/)
- [Run Benchmark on Android phone](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/5-run-benchmark-on-android/)
- [Build and Run Android chat app](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/6-build-android-chat-app/)
- [Next Steps](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-llama3-chat-android-app-using-executorch-and-xnnpack/_next-steps/)

## About this Learning Path

| Skill level:      | Introductory            |
|-------------------|-------------------------|
| Reading time:     | 1 hr                    |
| Last updated:     | 05 Aug 2026             |

| Authors:          | Varun Chari, Arm<br/>Pareena Verma, Arm [GitHub](https://github.com/pareenaverma) [LinkedIn](https://linkedin.com/in/pareena-verma-7853607) |
|-------------------|-------------------------|
| Arm IP:           | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags:             | [ML](https://learn.arm.com/tag/ml) [macOS](https://learn.arm.com/tag/macos) [Android](https://learn.arm.com/tag/android) [Java](https://learn.arm.com/tag/java) [CPP](https://learn.arm.com/tag/cpp) [Python](https://learn.arm.com/tag/python) [Hugging Face](https://learn.arm.com/tag/hugging-face) [ExecuTorch](https://learn.arm.com/tag/executorch) |

### Who is this for?

This is an introductory topic for software developers interested in learning how to build an Android chat app with Llama, KleidiAI, ExecuTorch, and XNNPACK.

### What will you learn?

Upon completion of this Learning Path, you will be able to:

- Set up an ExecuTorch development environment.
- Describe how ExecuTorch uses KleidiAI kernels to accelerate performance on Arm-based platforms.
- Describe how 4-bit groupwise PTQ quantization reduces model size without significantly sacrificing model accuracy.
- Build and run Llama models using ExecuTorch on your development machine.
- Build and run an Android chat app with different Llama models using ExecuTorch on an Arm-based smartphone.

### Prerequisites

Before starting, you will need the following:

- An Apple M1/M2 development machine with Android Studio installed or a Linux machine with at least 16GB of RAM.
- An Arm-powered smartphone with the i8mm feature running Android, with 16GB of RAM.
- A USB cable to connect your smartphone to your development machine.
- Android Debug Bridge (adb) installed on your device. Follow the steps in [adb](https://developer.android.com/tools/adb) to install Android SDK Platform Tools. The adb tool is included in this package.
- Java 17 JDK. Follow the steps in [Java 17 JDK](https://www.oracle.com/java/technologies/javase/jdk17-archive-downloads.html) to download and install JDK for host.
- Python 3.10.

### Summary

You’ll build and deploy an Android LLM chat app with ExecuTorch, XNNPACK, and KleidiAI on an Arm smartphone. First, you’ll set up an isolated Python environment, prepare a Llama 3.2 1B Instruct model, and enable KleidiAI through XNNPACK for supported Arm chips. Then, you’ll cross-compile the runner and JNI libraries with the Android NDK, integrate them, deploy the app, and run benchmarks.

### Frequently asked questions

<details>
<summary>Which Python environment should I use to install ExecuTorch dependencies?</summary>
Use an isolated environment. You can choose either a Python virtual environment or a Conda environment. You need only one environment.
</details>

<details>
<summary>How do I obtain and prepare the Llama 3.2 1B Instruct model for ExecuTorch?</summary>
Request access on Meta’s Llama Downloads page and use the time-limited download link you receive. Install the `llama-stack` package from `pip`, then run the provided command to download the model using your download link.
</details>

<details>
<summary>How do I know my Android NDK is set correctly before cross-compiling?</summary>
Set the `ANDROID_NDK` environment variable to your NDK path. Confirm that `$ANDROID_NDK/build/cmake/android.toolchain.cmake` exists so CMake can locate the Android toolchain.
</details>

<details>
<summary>What gets built when I compile for Android with KleidiAI enabled?</summary>
You’ll build the ExecuTorch runtime and a Llama runner binary for Android, along with JNI libraries for the app. Use these artifacts to run the model and execute benchmarks on the Android device.
</details>

<details>
<summary>Can I use a different Llama model instead of 3.2 1B Instruct?</summary>
Yes. The same instructions apply to other Llama options with minimal modification.
</details>
