# Characterize the memory subsystem of an Arm Linux system using ASCT

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/)
- [Identify Arm CPU topology, cache hierarchy, and NUMA configuration](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/system-overview/)
- [Analyze Arm cache hierarchy and performance characteristics](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/cache-hierarchy/)
- [Measure Arm cache and memory latency using ASCT pointer chase](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/pointer-chase-latency/)
- [Measure Arm single-core memory bandwidth with ASCT](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/streaming-bandwidth/)
- [Measure Arm multi-core memory bandwidth and loaded latency with ASCT](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/multicore-bandwidth/)
- [Compare Arm memory subsystem performance across systems](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/comparative-analysis/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/_next-steps/)

## About this Learning Path

| Skill level:    | Advanced    |
|------------------|-------------|
| Reading time:    | 1 hr        |
| Last updated:    | 29 Sep 2026 |

| Author:          | Jason Andrews, Arm [GitHub](https://github.com/jasonrandrews) [LinkedIn](https://linkedin.com/in/jason-andrews-7b05a8) |
|------------------|-------------|
| Arm IP:          | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:            | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture), [Linux](https://learn.arm.com/tag/linux), [ASCT](https://learn.arm.com/tag/asct), [Perf](https://learn.arm.com/tag/perf) |

### Who is this for?
This is an advanced topic for software developers and performance engineers who want to understand and characterize the CPU-side memory subsystem of Arm Linux systems.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Identify the core topology, cluster layout, and cache hierarchy of an Arm Linux system using standard tools.
- Measure cache and memory latency using a pointer-chase benchmark.
- Measure single-core and multi-core streaming bandwidth at each level of the memory hierarchy.
- Evaluate latency behavior under bandwidth pressure.
- Compare results across Arm systems and draw conclusions.

### Prerequisites
Before starting, you will need the following:
- Two or more Arm Linux systems with root or sudo access
- Arm System Characterization Tool (ASCT) installed on each system
- A good understanding of CPU memory subsystems, including cache hierarchies, cache lines, and DRAM in the memory hierarchy.

### Summary
You’ll characterize an Arm Linux server’s memory subsystem with ASCT. First, you’ll identify the CPU, cache, and Non-Uniform Memory Access (NUMA) topology, then measure cache and DRAM latency with pointer chasing. You’ll then profile single-core bandwidth across working sets and test multi-core scaling, saturation, and loaded latency. Finally, you’ll compare systems to recognize latency plateaus, bandwidth limits, and shared-resource effects.

### Frequently asked questions
<details>
<summary>Which ASCT benchmarks should I run to cover latency and bandwidth?</summary>
Use `pointer-chase` for cache and memory latency. Run the single-core bandwidth sweep for per-core throughput, then use `peak-bandwidth` and `loaded-latency` to assess multi-core scaling and latency under bandwidth pressure.
</details>

<details>
<summary>What should I verify about the system topology before running benchmarks?</summary>
Confirm the core count, cluster layout, private and shared cache levels, and NUMA configuration. Use the baseline to interpret performance cliffs and bandwidth scaling.
</details>

<details>
<summary>What result should I expect from the pointer-chase latency test?</summary>
Look for stepwise latency increases as the working set exceeds L1, L2, any shared cache, and finally DRAM. Use the plateaus and jumps to estimate the effective latency of each level.
</details>

<details>
<summary>How should I compare results across Arm systems, such as different Graviton generations?</summary>
Run the same ASCT benchmarks with consistent settings and similar conditions on each system. Compare latency plateaus, bandwidth peaks, and multi-core scaling while accounting for differences in core topology and cache hierarchy.
</details>

<details>
<summary>How should I interpret the NOP count in loaded-latency results?</summary>
Treat higher NOP counts as lower bandwidth pressure and zero NOPs as saturation. Look for the knee where latency starts climbing steeply as bandwidth increases.
</details>
