Optimize AArch64 code with LLVM link-time optimization and profile-guided optimization
Introduction
Understand PGO and LTO for AArch64 code
Prepare your AArch64 environment and verify LLVM tool availability
Build AArch64 code with LTO
Optimize AArch64 code with S-PGO
Optimize AArch64 code with FE-PGO
Optimize AArch64 code with IR-PGO
Optimize AArch64 code with CSIR-PGO
Next Steps
Optimize AArch64 code with LLVM link-time optimization and profile-guided optimization
What CSIR-PGO is
After an initial intermediate representation-level profile-guided optimization (IR-PGO) build, context-sensitive IR-PGO (CSIR-PGO) adds a second, context-sensitive profiling pass.
The first pass is the standard IR-PGO instrumentation pass. The second pass instruments the program after inlining, enabling LLVM to distinguish execution counts from different calling contexts.
The additional context can improve optimization when a function’s behavior depends on its call site. It doesn’t guarantee better performance for every program.
Use CSIR-PGO when you want to provide LLVM with more detailed profile information and can afford an extra build and training run.
Build the context-sensitive instrumented binary
First, generate prof/ir.profdata by completing the
IR-PGO workflow
. Use that profile to guide Clang’s usual PGO-driven optimization decisions while it builds a second instrumented binary:
clang++ -O3 -flto -fuse-ld=lld \
-fprofile-use=prof/ir.profdata \
-fcs-profile-generate=prof/csir \
bsort.cpp -o out/bsort.csirpgo.instr
The -fcs-profile-generate option then adds context-sensitive counters after inlining.
Run the context-sensitive instrumented binary:
./out/bsort.csirpgo.instr
Confirm that the training run created at least one context-sensitive raw profile:
ls prof/csir/*.profraw
Merge and inspect the profiles
Merge the context-sensitive raw profiles with the existing prof/ir.profdata profile:
Don’t merge only the .profraw files: the final CSIR-PGO profile needs counts from both instrumentation passes.
llvm-profdata merge prof/ir.profdata prof/csir -output=prof/csir.profdata
Inspect the context-sensitive execution counts recorded for sort_array. For this example, sort_array should have six context-sensitive counters:
llvm-profdata show --showcs --counts --function=sort_array prof/csir.profdata
The output is similar to:
Counters:
ld-temp.o;_Z10sort_arrayPi:
Hash: 0x18c2aba34f0cfff9
Counters: 6
Block counts: [24763682, 25224415, 9999, 9882, 1, 9881]
Instrumentation level: IR entry_first = 0 instrument_loop_entries = 0
Functions shown: 1
Total functions: 12
Maximum function count: 24763682
Maximum internal block count: 25224415
Total number of blocks: 32
Total count: 75242276
The --showcs option selects context-sensitive records from the merged profile. Exact hashes and individual counter values can vary with the LLVM version and training workload.
Build with CSIR-PGO and LTO
Build the optimized binary using the merged CSIR-PGO profile:
clang++ -O3 -flto -fuse-ld=lld \
-fprofile-use=prof/csir.profdata \
bsort.cpp -o out/bsort.csirpgo.opt
Run the optimized binary:
./out/bsort.csirpgo.opt
What you’ve accomplished
You’ve added a second context-sensitive profiling pass on top of IR-PGO and built a CSIR-PGO optimized binary with LTO.
You can now apply the workflow that best matches your profiling environment to a representative workload from your own application.