# Analyze cache behavior with Perf C2C on Arm

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/)
- [Arm Statistical Profiling Extension and false sharing](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/how-to-1/)
- [Set up your environment for Arm SPE and Perf C2C profiling](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/how-to-2/)
- [False sharing example](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/how-to-3/)
- [Perform root cause analysis with Perf C2C](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/how-to-4/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/false-sharing-arm-spe/_next-steps/)

## About this Learning Path

| Skill level:         | Introductory                         |
|----------------------|-------------------------------------|
| Reading time:        | 15 min                              |
| Last updated:        | 11 Sep 2026                         |

| Author:              | Kieran Hejmadi, Arm [GitHub](https://github.com/kieranhejmadi01) [LinkedIn](https://linkedin.com/in/kieran-hejmadi-88920815b) |
|----------------------|-----------------------------------------------------|
| Arm IP:              | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors)                |
| Tags:                | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture), [AWS Graviton](https://learn.arm.com/tag/aws-graviton), [Microsoft Azure Cobalt](https://learn.arm.com/tag/microsoft-azure-cobalt), [Google Axion](https://learn.arm.com/tag/google-axion), [Oracle Cloud Infrastructure (OCI) Ampere Compute](https://learn.arm.com/tag/oracle-cloud-infrastructure-oci-ampere-compute), [Linux](https://learn.arm.com/tag/linux), [perf](https://learn.arm.com/tag/perf), [Runbook](https://learn.arm.com/tag/runbook) |

### Who is this for?

This topic is for performance-oriented developers working on Arm-based cloud or server systems who want to optimize memory access patterns and investigate cache inefficiencies using Perf C2C and Arm SPE.

### What will you learn?

Upon completion of this Learning Path, you will be able to:

- Identify and fix false sharing issues using Perf C2C, a cache line analysis tool.
- Enable and use the Arm Statistical Profiling Extension (SPE) on Linux systems.
- Investigate cache line performance with Perf C2C.

### Prerequisites

Before starting, you will need the following:

- Access to an Arm-based cloud instance with support for the Arm Statistical Profiling Extension (SPE).
- A basic understanding of cache coherency and its impact on performance.
- Familiarity with Linux Perf tools.

### Summary

You’ll analyze cache behavior on Arm cloud systems with Arm SPE and Perf C2C. First, you’ll verify SPE and Linux `perf` access, then build aligned and unaligned versions of a multithreaded C example. Next, you’ll compare runtimes with `perf stat` and use Perf C2C to identify contended cache lines and map them to source code. You’ll finish by applying alignment changes to reduce false sharing.

### Frequently asked questions

<details><summary>How do I confirm that my Arm-based instance supports Arm SPE before collecting data?</summary>

Check that both the hardware and the kernel support Arm SPE and verify that Linux `perf` can access the relevant events. Run `sudo modprobe arm_spe_pmu` to confirm if the SPE kernel module is loaded. Run `ls /sys/bus/event_source/devices/ | grep arm_spe` to check if SPE is included in the kernel.

</details>

<details><summary>Which build should I profile first with Perf C2C to observe false sharing?</summary>

Start with the unaligned version of the example to expose false sharing. Then, profile the cache-aligned version to compare results and confirm the impact of alignment.

</details>

<details><summary>What result should I expect when comparing the aligned and unaligned binaries?</summary>

Expect a runtime and counter difference between the two builds, with the unaligned version typically showing more cache-related contention. Exact values vary by system and load.

</details>

<details><summary>How do I know Perf C2C captured useful cache line information?</summary>

Look for a report that highlights shared or contended cache lines and maps addresses back to symbols or source locations. You should be able to identify which structures or lines of code are involved in the contention.

</details>

<details><summary>How can I troubleshoot similar performance between the aligned and unaligned builds?</summary>

Verify that the alignment changes are present in the compiled binaries and that multiple threads are actively updating shared data. Repeat measurements with `perf stat -r 3` for each binary and confirm that your SPE and `perf` setup is correct.

</details>
