# [Optimize the application](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/optimize-1/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/)
- [Introduction to performance analysis](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/intro/)
- [Build an example application](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/application-1/)
- [Gather performance metrics](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/analysis-1/)
- [Optimize the application](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/optimize-1/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/top-down-n1/_next-steps/)

## Performance Optimization
You can use software prefetching to improve performance.

The code to enable prefetching is:
```c
#if defined(ENABLE_PREFETCH) && defined (DIST)
    const int prefetch_distance = DIST * kStride * kLineSize;
    __builtin_prefetch(&buffer[position + prefetch_distance], 0, 0);
#endif
```

To enable data prefetching, recompile the application with 2 defines. You can experiment with values of `DIST` to see how performance is impacted.

The white paper shows a graph of various values of DIST and explains how the performance saturates at a `DIST` value of 40 on the N1SDP hardware. The example below uses 100 for `DIST`.

To compile with prefetching:
```bash
g++ -g -O3 -DENABLE_PREFETCH -DDIST=100 stride.cpp -o stride
```

Run the `perf` command again to count instructions and cycles.
```bash
perf stat -e instructions,cycles ./stride
```

The output is similar to:
```
__output__Performance counter stats for './stride':
__output__
__output__14,002,762,166      instructions:u            #    0.63  insn per cycle
__output__22,106,858,662      cycles:u
__output__
   9.895357134 seconds time elapsed
__output__
   9.874959000 seconds user
   0.020003000 seconds sys
```

The time to run the original application was more than 20 seconds and with prefetching the time drops to under 10 seconds.

The table below shows the improvements in instructions, cycles, and IPC.

| Metrics      | Baseline        | Optimized       |
|--------------|-----------------|-----------------|
| Instructions | 10,002,762,164  | 14,002,762,166  |
| Cycles       | 45,157,063,927  | 22,106,858,662  |
| IPC          | 0.22            | 0.63            |
| Runtime      | 20.1 sec        | 9.9 sec         |

As outlined in the white paper, the IPC nearly triples and the runtime halves.

You can review how the other metrics change with prefetching enabled.

**Note:** Neoverse N1 servers and cloud instances will not demonstrate as much performance improvement because they have higher memory bandwidth, but you will be able to see performance improvement as a result of enabling prefetching.
