# [Write SIMD code on Arm using Rust](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/)
- [Introduction to Rust](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/intro-to-rust/)
- [Arm SIMD on Rust](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/simd-on-rust-part1/)
- [Inlining Intrinsics](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/simd-on-rust-part2/)
- [Matrix transpose](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/simd-on-rust-part3/)
- [A more complicated example](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/simd-on-rust-part4/)
- [Conclusion](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/conclusion/)
- [Next Steps](https://learn.arm.com/learning-paths/cross-platform/simd-on-rust/_next-steps/)

## About this Learning Path

| Skill level:    | Advanced      |
|------------------|---------------|
| Reading time:    | 30 min        |
| Last updated:    | 03 Aug 2026   |

| Author:          | Konstantinos Margaritis, VectorCamp |
|------------------|--------------------------------------|
| Arm IP:          | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:            | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture) [Linux](https://learn.arm.com/tag/linux) [GCC](https://learn.arm.com/tag/gcc) [Clang](https://learn.arm.com/tag/clang) [Rust](https://learn.arm.com/tag/rust) [Runbook](https://learn.arm.com/tag/runbook) |

### Who is this for?
This is an advanced topic for software developers who want to take advantage of SIMD code on Arm systems using Rust.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Write SIMD code with Rust using `std::arch` and Neon intrinsics on Arm
- Use portable SIMD abstractions with `std::simd` for cross-platform code
- Apply feature detection and target attributes for architecture-specific optimizations
- Compare C and Rust SIMD implementations and disassembly output

### Prerequisites
Before starting, you will need the following:
- An Arm-based computer with recent versions of a C compiler (Clang or GCC) and a Rust compiler installed.

### Summary
You’ll write SIMD code for Arm with Rust by translating familiar C examples into Rust. Starting with Arm Advanced SIMD (Neon) intrinsics in C, you’ll mirror them in Rust with `std::arch` and explore portable SIMD with `std::simd`. You’ll implement pairwise averaging, a dot-product sum of absolute differences, a 4x4 matrix transpose, and a DCT butterfly. Then, you’ll compare C and Rust output and disassembly.

### Frequently asked questions
<details>
<summary>Which Rust SIMD API should I use when porting the C Neon intrinsics examples?</summary>
Use `std::arch` for a one-to-one match with the C Neon intrinsics. Use `std::simd` when you prefer a portable abstraction.
</details>

<details>
<summary>How do I know the pairwise average example produced the right result?</summary>
The program averages corresponding elements from two arrays. Compare its output with the scalar calculation `(A[i] + B[i]) / 2` for each index.
</details>

<details>
<summary>What output should I expect from the dot product (`vdotq_u32`) example?</summary>
The example prints the input arrays and then reports one sum of absolute differences (SAD) value. Compare the arrays and SAD total with the expected calculation.
</details>

<details>
<summary>What should I check if the Rust intrinsics code does not compile?</summary>
Confirm that the imports and architecture-specific attributes match the example, and target an Arm platform that supports the intrinsics. Mismatched names or attributes can cause build errors.
</details>

<details>
<summary>How do I compare the C and Rust implementations for the transpose or butterfly steps?</summary>
Build and run both versions, then compare their printed outputs. Review the disassembly to see how each implementation maps to Arm Neon instructions.
</details>
