# Optimize SIMD code with vectorization-friendly data layout

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/)
- [What exactly is data layout?](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/data-layout-basics/)
- [Improve data alignment](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/data-layout-basics-2/)
- [Increase complexity](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/a-more-complex-problem/)
- [Write hand optimized SIMD code](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/a-more-complex-problem-manual-simd/)
- [Structure of arrays](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/a-more-complex-problem-revisited/)
- [Migrate to the Scalable Vector Extension (SVE)](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/a-more-complex-problem-revisited-sve/)
- [Next Steps](https://learn.arm.com/learning-paths/cross-platform/vectorization-friendly-data-layout/_next-steps/)

## About this Learning Path

| Skill level:     | Advanced          |
|-------------------|-------------------|
| Reading time:     | 45 min            |
| Last updated:     | 04 Aug 2026       |

| Author:                  | Konstantinos Margaritis, VectorCamp |
|--------------------------|--------------------------------------|
| Arm IP:                  | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors), [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags:                    | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture), [Linux](https://learn.arm.com/tag/linux), [GCC](https://learn.arm.com/tag/gcc), [Clang](https://learn.arm.com/tag/clang), [Runbook](https://learn.arm.com/tag/runbook) |

### Who is this for?
This is an advanced topic for C/C++ developers who are interested in improving the performance of SIMD code.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Comprehend the importance of data layout when writing SIMD code

### Prerequisites
Before starting, you will need the following:
- An Arm computer running Linux and a recent version of Clang or the GNU compiler (gcc) installed.

### Summary
You’ll learn how data layout affects SIMD vectorization on Arm. Starting with an Array-of-Structures model that stores x, y, and z values in 12-byte triplets, you’ll examine how strided access limits four-wide floating-point operations. You’ll add padding for four-element vectors, handle boundary checks, and write a hand-optimized Arm NEON version when auto-vectorization is insufficient. Then, you’ll compare a Structure-of-Arrays layout with contiguous component arrays and explore an SVE implementation using the simulation examples.

### Frequently asked questions

<details>
<summary>How do I know the current data layout is blocking SIMD vectorization?</summary>
If your structure stores 3-element vectors (x, y, and z), the 12-byte grouping creates strided access that can make 4-wide SIMD vectorization of 32-bit floats difficult. Use compiler output and the example programs to see how this layout affects vectorization.
</details>

<details>
<summary>What should I change in the object struct to improve SIMD access?</summary>
Replace the three-element `vec3` fields with four-element `vec4` fields and initialize the additional element as padding. The 16-byte layout makes four-wide operations easier for the compiler. The later Structure-of-Arrays example stores each component contiguously instead.
</details>

<details>
<summary>Which files do I modify as complexity increases?</summary>
Copy `simulation1.c` to `simulation2.c` to add bounding-box logic, including the updated `simulate_objects()` function, the `ctr4` structure, and the `box` constant. Then copy `simulation2.c` to `simulation3.c` for the hand-written NEON version. Create `simulation4.c` for the Structure-of-Arrays approach, or `simulation4_sve.c` for the SVE example.
</details>

<details>
<summary>What should I expect to see in the hand-optimized SIMD version?</summary>
`simulation3.c` includes `arm_neon.h` and uses types such as `float32x4_t` to process data in 4-wide chunks. The math is expressed with NEON intrinsics, so vector operations are explicit in the source.
</details>

<details>
<summary>How does the Structure-of-Arrays version change the way the code runs?</summary>
The Structure-of-Arrays version stores each component, such as all x values, in its own contiguous array. This improves sequential access and makes 4-wide operations straightforward, reducing the penalties of interleaved fields.
</details>
