# Accelerate matrix multiplication performance with SME2

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/)
- [Overview](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/overview/)
- [Set up your SME2 development environment](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/1-get-started/)
- [Test your SME2 development environment](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/2-check-your-environment/)
- [Streaming mode and ZA state in SME](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/3-streaming-mode/)
- [Vanilla matrix multiplication](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/4-vanilla-matmul/)
- [Outer product](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/5-outer-product/)
- [SME2 assembly matrix multiplication](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/6-sme2-matmul-asm/)
- [Matrix multiplication using SME2 intrinsics in C](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/7-sme2-matmul-intr/)
- [Benchmarking](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/8-benchmarking/)
- [Debugging](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/9-debugging/)
- [Going further](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/10-going-further/)
- [Next Steps](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/_next-steps/)

## About this Learning Path

| Skill level:                | Advanced         |
|-----------------------------|-------------------|
| Reading time:               | 1 hr              |
| Last updated:               | 19 Aug 2026       |
| Author:                     | Arnaud de Grandmaison, Arm [GitHub](https://github.com/Arnaud-de-Grandmaison-ARM) [LinkedIn](https://linkedin.com/in/arnauddegrandmaison) |
| Arm IP:                     | [Arm C1](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags:                       | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture), [Linux](https://learn.arm.com/tag/linux), [macOS](https://learn.arm.com/tag/macos), [Windows](https://learn.arm.com/tag/windows), [C](https://learn.arm.com/tag/c), [Clang](https://learn.arm.com/tag/clang), [LLVM](https://learn.arm.com/tag/llvm), [SME2](https://learn.arm.com/tag/sme2) |

### Who is this for?
This Learning Path is an advanced topic for developers who want to accelerate the performance of matrix multiplication using Arm's Scalable Matrix Extension Version 2 (SME2).

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Implement a baseline matrix multiplication kernel in C without SME2
- Use SME2 assembly instructions to accelerate matrix multiplication performance
- Use SME2 intrinsics to vectorize and optimize matrix multiplication
- Compile code with SME2 intrinsics and assembly
- Benchmark and validate SME2-accelerated matrix multiplication on Arm hardware or in a Linux-based emulation environment
- Compare performance metrics between baseline and SME2-optimized implementations

### Prerequisites
Before starting, you will need the following:
- Working knowledge of Arm’s SVE and SME2 instruction sets
- Intermediate proficiency with the C programming language and the Armv9-A assembly language
- A computer running Linux, macOS, or Windows
- Installations of Git, CMake and Ninja for project setup
- A platform that supports SME2 - see the list of [devices with SME2 support](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/1-get-started/#devices) or an emulator to run code with SME2 instructions
- Installation of Docker for SME2 emulation (if you don’t have SME2 available)
- Installation of Android Development Studio and adb (if you’re targeting an Android phone with SME2 support)
- Compiler support for SME2 instructions (for example, LLVM 18 or later with SME2 backend support)

### Summary
You’ll start with a baseline C matrix-multiplication kernel, then build SME2-accelerated versions with intrinsics and assembly either on supported Arm hardware or an emulator. You’ll verify the toolchain with CMake, learn how SME streaming mode and ZA state work through Arm C Language Extensions, and implement a row-major reference kernel. Then, you’ll benchmark SME2 versions and validate them against the baseline.

### Frequently asked questions

<details>
<summary>Should I use native SME2 hardware or emulation?</summary>
Use native SME2 hardware if you have a supported device, such as a Mac with an M4 chip or some Android phones. Otherwise, use the emulation option and check the setup section’s device list for native support.
</details>

<details>
<summary>CMake is finding the wrong Clang. What should I do?</summary>
Point CMake to the intended Clang when configuring the project. The system default might not support SME2, so select the compiler as described in the environment check.
</details>

<details>
<summary>How do I verify the environment is ready before continuing?</summary>
Build the examples in `code-examples/learning-paths/cross-platform/multiplying-matrices-with-sme2` using CMake. A successful build and run confirms that the toolchain, compiler, and hardware or emulator are configured correctly.
</details>

<details>
<summary>When should I enable streaming mode, and do I need to handle ZA state myself?</summary>
Enable streaming mode on functions that use SME features by applying the appropriate Arm C Language Extensions. The compiler manages streaming transitions and ZA save and restore, so you don’t implement those operations manually.
</details>

<details>
<summary>What result should I expect from the baseline C matrix multiplication?</summary>
The baseline row-major implementation computes the reference output matrix. Keep that result and compare it with the intrinsics and assembly versions during benchmarking and validation.
</details>
