# Compare Neon and SVE with the Arm Performix Instruction Mix recipe

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/)
- [Understand profiling with Arm Performix Instruction Mix](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-1/)
- [Set up and run a GPT-2 baseline](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-2/)
- [Find GPT-2 hotspots and profile with the Arm Performix Instruction Mix recipe](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-3/)
- [(Optional) Optimize matmul with vector intrinsics](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-4/)
- [Compare Neon and SVE with the Arm Performix Instruction Mix recipe](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-5/)
- [Accelerate execution with KleidiAI](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/how-to-6/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/_next-steps/)

## Verify vectorization with the Instruction Mix recipe
In the earlier build step, you created `gpt2_neon` and `gpt2_sve`. These binaries use the reference solutions in `matmul_neon.cpp` and `matmul_sve.cpp`, respectively.

Run the `gpt2_neon` binary with the following command to observe the speedup.
```
./build/gpt2_neon --model gpt2-medium "Once upon a time" -n 20
```
![GPT-2 Neon runtime output on Arm Linux](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/gpt_neon.gif)  
*GPT-2 Neon runtime output on Arm Linux*

The Neon implementation delivers a noticeable increase in token generation throughput. To evaluate the SVE implementation, run the same workload using the `gpt2_sve` binary instead of `gpt2_neon`. Performance gains will vary across systems, largely depending on the Scalable Vector Length (SVL) supported by the Arm Linux platform.

Rerun the Instruction Mix recipe with the `gpt2_neon` binary using the same recipe settings and workload arguments as baseline. After the run completes, select both runs and select **Compare** to open a comparison view.

> **Tip**  
> Rename each run with a descriptive name, such as baseline and Neon, so you can identify and compare results quickly.

The baseline profile is mostly scalar instructions. After you add Neon intrinsics, the instruction mix shifts toward Advanced SIMD (Neon) instructions, showing that the code is using Arm Neon hardware more effectively.

![Neon versus scalar instruction mix](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/neon_scalar_instruction_mix.webp)  
*Neon versus scalar instruction mix*

You can also compare SVE variants in the same way. The increase in SVE operations shows that this path is now using SVE hardware.

![SVE versus baseline instruction mix](https://learn.arm.com/learning-paths/servers-and-cloud-computing/performix-instruction-mix/sve_vs_baseline.webp)  
*SVE versus baseline instruction mix*

## Compare throughput across kernels
Compare throughput across Neon and SVE.

### Neon kernel
You can also inspect the Neon intrinsic implementation using Compiler Explorer, where the hot accumulation step (`vacc`) runs in ASIMD (Neon) registers such as `v0`:

### SVE kernel
For variable-length vectorization, compare with an explicit SVE implementation that assumes SVE support, where the hot accumulation step (`vacc`) runs in SVE z registers with predicate-controlled loads and multiply-accumulate:

For a full-page view, open a [Godbolt session with all three `matmul` kernels](https://godbolt.org/z/E4a7Wxh8K).

## Measure speedup
Run the provided comparison script to measure tokens per second across all available binaries:
```
./compare_gpt2_variants.sh
__output__ 
Model: gpt2-medium 
__output__ 
Prompt: Once upon a time 
__output__ 
Tokens: 20 
__output__ 
Runs: 1 
__output__ 
== gpt2 == 
__output__ 
run 1: 3.04976 tok/s 
__output__ 
avg: 3.049760 tok/s 
__output__ 
== gpt2_neon == 
__output__ 
run 1: 11.3649 tok/s 
__output__ 
avg: 11.364900 tok/s 
__output__ 
== gpt2_sve == 
__output__ 
run 1: 13.907 tok/s 
__output__ 
avg: 13.907000 tok/s 
__output__ 
== gpt2_user == 
__output__ 
run 1: 3.04859 tok/s 
__output__ 
avg: 3.048590 tok/s 
```
These results show that intrinsics increase throughput from about 3 tok/s in the scalar baseline to about 13.9 tok/s with SVE.

## What you’ve accomplished and what’s next
You’ve now verified vectorization using the Instruction Mix recipe, compared throughput across Neon and SVE, and measured tokens per second across the available binaries.

Next, you’ll use optimized libraries to push performance further.
