# Compare Arm memory subsystem performance across systems

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/)
- [Identify Arm CPU topology, cache hierarchy, and NUMA configuration](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/system-overview/)
- [Analyze Arm cache hierarchy and performance characteristics](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/cache-hierarchy/)
- [Measure Arm cache and memory latency using ASCT pointer chase](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/pointer-chase-latency/)
- [Measure Arm single-core memory bandwidth with ASCT](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/streaming-bandwidth/)
- [Measure Arm multi-core memory bandwidth and loaded latency with ASCT](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/multicore-bandwidth/)
- [Compare Arm memory subsystem performance across systems](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/comparative-analysis/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/memory-subsystem/_next-steps/)

## Complete memory subsystem analysis workflow
You now have a complete set of memory subsystem measurements for each of your test systems. This section shows how to collect all results in one pass, compare systems using ASCT’s built-in diff tool, and draw meaningful architectural conclusions.

## Collect all memory benchmarks in one command
Throughout the previous sections, you ran individual ASCT benchmarks. To characterize a new system from scratch, run all memory benchmarks at once.

Run all tests on each system and save the results.
```
sudo asct run memory loaded-latency --output-dir results_$(hostname)
```
This runs `latency-sweep`, `bandwidth-sweep`, `idle-latency`, `peak-bandwidth`, `c2c-latency`, and `loaded-latency` in a single pass, saving all results and plots to the output directory. The directory name includes the hostname so you can easily identify which system produced the data.

## Compare systems with the diff command
ASCT includes a `diff` command that compares output directories from different runs. It loads the results from each directory, groups fields by benchmark, and reports percentage differences for measurements and raw values for system configuration.

After collecting results on both instances, copy the output directories to the same machine and run:
```
asct diff results_graviton2-c6g/ results_graviton4-c8g/
```
ASCT uses the first directory as the baseline.

To specify a baseline explicitly:
```
asct diff results_graviton2-c6g/ --baseline results_graviton4-c8g/
```

### Save diff output
Save the comparison as CSV for further analysis:
```
asct diff results_graviton2-c6g/ results_graviton4-c8g/ --format=csv --output-dir comparison/ --benchmarks ^system-info
```
ASCT writes `diff.csv` to the output directory.

### Filter the comparison
You can also focus on specific benchmarks.

To compare only peak bandwidth and latency sweep:
```
asct diff results_graviton2-c6g/ results_graviton4-c8g/ --benchmarks peak-bandwidth latency-sweep
```
To compare everything except system-info:
```
asct diff results_graviton2-c6g/ results_graviton4-c8g/ --benchmarks ^system-info
```

The output for everything except the system info is:
```
__output__,field,recipe,run,comparator,delta,delta_percent,baseline
__output__0,loaded-latency.data.0.Bandwidth [GB/s],loaded-latency,results_graviton4-c8g,459.65,291.03,172.59%,168.62
__output__1,loaded-latency.data.0.Loaded latency [ns],loaded-latency,results_graviton4-c8g,197.43,-146.73,-42.63%,344.16
__output__2,loaded-latency.data.10.Bandwidth [GB/s],loaded-latency,results_graviton4-c8g,464.27,298.39,179.88%,165.88
...
__output__53,idle-latency.data.Node 0.Node 0,idle-latency,results_graviton4-c8g,115.16,19.52,20.41%,95.64
```

## Organize your results
Collect the key measurements from each system into a comparison table:

| Metric | Graviton2 (Neoverse N1) | Graviton4 (Neoverse V2) | Graviton4 vs Graviton2 |
|--------|--------------------------|--------------------------|------------------------|
| L1D latency | 1.6 ns | 1.4 ns | 11% faster |
| L2 latency | 5.4 ns | 4.0 ns | 27% faster |
| LLC latency | 28.8 ns | 21.8 ns | 24% faster |
| DRAM latency (unloaded) | 95.5 ns | 114.6 ns | 20% slower |
| Core-to-core latency (local) | 41.8 ns | 25.8 ns | 38% faster |
| Loaded latency at idle (~3000 NOPs) | 96.7 ns | 115.1 ns | 19% slower |
| Loaded latency at saturation (0 NOPs) | 344.2 ns | 197.4 ns | 43% faster |
| L1 bandwidth (1 core) | 159.4 GB/s | 321.1 GB/s | 101% higher |
| L2 bandwidth (1 core) | 73.4 GB/s | 95.0 GB/s | 29% higher |
| LLC bandwidth (1 core) | 35.8 GB/s | 79.6 GB/s | 123% higher |
| DRAM bandwidth (1 core) | 20.9 GB/s | 37.0 GB/s | 77% higher |
| Peak bandwidth (all cores, all reads) | 169.7 GB/s | 465.6 GB/s | 174% higher |

## Analyze the differences
The conclusions below are from running ASCT on AWS EC2 instances that are easily available for you to try so you can learn the process and learn how to think about the results. There are many more details that go into CPU microarchitecture and system design that are not covered here, but this is a good way to get an initial understanding of the memory system of an Arm Linux server.

### Latency analysis across generations
The latency measurements from `latency-sweep` reveal how each generation’s cache hierarchy performs. L1 latency is similar across both generations (1.6 ns vs 1.4 ns) because L1 caches are designed for very low latency, so the small difference reflects clock speed rather than cache microarchitecture. L2 latency improves by 27% on Graviton4 (4.0 ns vs 5.4 ns): Neoverse V2 has a 2 MB private L2 compared to Neoverse N1’s 1 MB, and despite the larger size, the improved cache pipeline delivers lower latency. LLC latency improves by 24% on Graviton4 (21.8 ns vs 28.8 ns) even though the L3 is only slightly larger (36 MB vs 32 MB), which reflects microarchitectural improvements in the Neoverse V2 cache interconnect. DRAM latency is higher on Graviton4 (114.6 ns vs 95.5 ns, +20%) because DDR5 trades slightly higher access latency for significantly more bandwidth per channel compared to DDR4.

### Loaded latency comparison
The `loaded-latency` results reveal the latency-bandwidth tradeoff. At low load (~3000 NOPs), latency matches the idle DRAM latency from `latency-sweep` (96.7 ns for Graviton2, 115.1 ns for Graviton4), confirming the two benchmarks are consistent. The knee on Graviton2 occurs between 70 and 50 NOPs, where latency nearly doubles from 134 ns to 266 ns before climbing to 344 ns at saturation. On Graviton4 the knee occurs between 30 and 20 NOPs at a much higher bandwidth (~340 GB/s vs ~126 GB/s on Graviton2), and latency plateaus around 190–198 ns at saturation rather than continuing to climb, suggesting the DDR5 memory controllers handle queue saturation more gracefully. Despite Graviton4’s higher idle DRAM latency, its latency at saturation is 43% lower than Graviton2’s (197 ns vs 344 ns), making it the better choice for workloads that operate near memory bandwidth limits.

### Bandwidth comparison across systems
Single-core bandwidth from `bandwidth-sweep` shows the throughput capacity at each cache level. L1 bandwidth doubles on Graviton4 (321 vs 159 GB/s, +101%), likely reflecting a combination of higher clock speed and microarchitectural throughput improvements in Neoverse V2. L2 bandwidth improves by 29% (95 vs 73 GB/s) due to the wider L2 fill path, and LLC bandwidth more than doubles (80 vs 36 GB/s, +123%), reflecting improvements in the interconnect and shared cache design. Single-core DRAM bandwidth improves by 77% (37 vs 21 GB/s) because Neoverse V2 supports more outstanding memory requests than Neoverse N1 and DDR5 provides more bandwidth per channel than DDR4.

Peak bandwidth from `peak-bandwidth` shows the system-level memory controller capacity. Graviton4 delivers 2.7x the peak all-reads bandwidth of Graviton2 (465.6 vs 169.7 GB/s), the most dramatic difference between the two generations. Compare the “All Reads” figure with the theoretical peak bandwidth from `asct system-info` to see how efficiently each system uses its available bandwidth — well-configured systems typically achieve 85–95% of theoretical peak.

## Cross-validate your findings
The credibility of this type of analysis comes from cross-validation. Check that:
- Latency steps match cache sizes: the `latency-sweep` boundaries should align with the cache sizes reported by `sysfs` and the TRM.
- Bandwidth plateaus match cache sizes: the `bandwidth-sweep` should show transitions at the same data sizes that `latency-sweep` identified.
- Peak bandwidth matches documented limits: compare the “All Reads” figure from `peak-bandwidth` against the theoretical peak from `asct system-info`. A well-configured system typically achieves 85-95% of theoretical peak.
- Idle loaded latency matches unloaded latency: the lowest-load row from `loaded-latency` should be close to the DRAM latency from `latency-sweep`.
- Results are repeatable: run each benchmark at least twice. If results vary by more than 5%, investigate sources of noise (CPU frequency scaling, background processes).

### Tip
Check if your system allows you to control CPU frequency scaling during benchmarks for more consistent results:
```
echo performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
```

## Adapting this analysis to any Arm Linux system
The examples here used AWS Graviton instances, but the methodology works on any Arm Linux machine with a NUMA-enabled kernel. To adapt it:
- Install ASCT on the target system following the [install guide](https://learn.arm.com/install-guides/asct/).
- Run all memory benchmarks: `sudo asct run memory loaded-latency --format=csv --output-dir results_$(hostname)`
- Compare with previous systems: `asct diff results_new/ --baseline results_reference/`
- Cross-validate: check that latency boundaries, bandwidth plateaus, and peak numbers are consistent with each other and with the hardware documentation.

## What you’ve learned
In this section you:
- Identified the topology and cache hierarchy of Arm Linux systems using `sysfs` and `asct system-info`
- Measured cache and memory latency with the ASCT `latency-sweep` benchmark
- Measured single-core streaming bandwidth with `bandwidth-sweep`
- Measured peak system bandwidth with `peak-bandwidth` and the latency-bandwidth tradeoff with `loaded-latency`
- Compared Graviton2 and Graviton4 to understand how generational improvements in core microarchitecture and memory technology affect real-world performance
- Developed a portable methodology that works on any Arm Linux system with a single command
