# [Run ERNIE-4.5 Mixture of Experts model on Armv9 with llama.cpp](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/)
- [Understand Mixture of Experts architecture for edge deployment](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/1_mixture_of_experts/)
- [Set up llama.cpp on an Armv9 development board](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/2_llamacpp_installation/)
- [Compare ERNIE model behavior and expert routing](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/3_ernie_moe/)
- [Optimize performance with Armv9 hardware features](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/4_v9_optimization/)
- [Next Steps](https://learn.arm.com/learning-paths/cross-platform/ernie_moe_v9/_next-steps/)

## About this Learning Path

| Skill level: | Advanced |
|--------------|----------|
| Reading time: | 1 hr |
| Last updated: | 06 Jul 2026 |

| Author: | Odin Shen, Arm [GitHub](https://github.com/odincodeshen) [LinkedIn](https://linkedin.com/in/odin-shen-lmshen) |
|----------|---------------------------------------------------------------------------------------------------------------------|
| Arm IP: | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags: | [ML](https://learn.arm.com/tag/ml) [Linux](https://learn.arm.com/tag/linux) [Python](https://learn.arm.com/tag/python) [CPP](https://learn.arm.com/tag/cpp) [Bash](https://learn.arm.com/tag/bash) [llama.cpp](https://learn.arm.com/tag/llama.cpp) |

### Who is this for?
This is an advanced topic for developers and engineers who want to deploy Mixture of Experts (MoE) models, such as ERNIE 4.5, on edge devices. MoE architectures allow large LLMs with 21 billion or more parameters to run with only a fraction of their weights active per inference, making them ideal for resource constrained environments.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Deploy MoE models like ERNIE-4.5 on edge devices using llama.cpp
- Compare inference behavior between ERNIE-4.5 PT and Thinking versions
- Measure performance impact of Armv9-specific hardware optimizations

### Prerequisites
Before starting, you will need the following:
- An Armv9 device with at least 32 GB of available disk space, for example, Radxa Orion O6

### Summary
You’ll use `llama.cpp` on an Armv9 Linux development board to deploy ERNIE-4.5 Mixture of Experts models, validate inference, and confirm multilingual output with the Thinking variant. You’ll install the PT and Thinking models, run the same task across both, and examine internal expert routing to see how only a subset of parameters is activated at runtime. Then, you’ll compare a baseline CPU build with an Armv9-optimized build that enables SVE, i8mm, and dotprod, and benchmark them under identical conditions to measure the impact of Armv9-specific optimizations.

### Frequently asked questions
<details>
<summary>What result should I expect when verifying the setup on the Armv9 board?</summary>
You should see successful model inference and multilingual output from the ERNIE-4.5 Thinking variant. This confirms the `llama.cpp` build and runtime environment are working end to end.
</details>

<details>
<summary>Which ERNIE-4.5 variant should I use for the comparison step?</summary>
Use both PT and Thinking on the same task and with the same settings. This makes the differences in response style and reasoning easier to observe.
</details>

<details>
<summary>How do I inspect and interpret MoE expert routing during generation?</summary>
Use the inspection method shown in the steps to view which experts are selected per token. Compare activation patterns between PT and Thinking to understand routing differences.
</details>

<details>
<summary>How do I set up baseline and optimized builds to benchmark Armv9 features?</summary>
Build a regular CPU version and a separate Armv9-specific version with SVE, i8mm, and dotprod enabled. Run the same benchmarks on both builds and compare results under identical conditions.
</details>

<details>
<summary>What should I check if a model download or inference run fails?</summary>
Verify the device uses an Armv9 CPU and that at least 32 GB of disk space is available. Re-run the setup verification on the board before proceeding to comparisons or benchmarking.
</details>
