# Unlock quantized LLM performance on Arm-based NVIDIA DGX Spark

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/)
- [Explore Grace Blackwell architecture for efficient quantized LLM inference](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/1_gb10_introduction/)
- [Verify your Grace Blackwell system readiness for AI inference](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/1a_gb10_setup/)
- [Build the GPU version of llama.cpp on GB10](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/2_gb10_llamacpp_gpu/)
- [Build the CPU version of llama.cpp on GB10](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/3_gb10_llamacpp_cpu/)
- [Analyze CPU instruction mix using Process Watch](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/4_gb10_processwatch/)
- [Next Steps](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/_next-steps/)

## About this Learning Path

| Skill level:       | Introductory     |
|--------------------|------------------|
| Reading time:      | 1 hr             |
| Last updated:      | 29 Jul 2026      |

| Author:            | Odin Shen, Arm [GitHub](https://github.com/odincodeshen) [LinkedIn](https://linkedin.com/in/odin-shen-lmshen)  |
|--------------------|---------------------------------------------------|
| Arm IP:            | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) [Cortex-X](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags:              | [ML](https://learn.arm.com/tag/ml) [Linux](https://learn.arm.com/tag/linux) [Python](https://learn.arm.com/tag/python) [C](https://learn.arm.com/tag/c) [Bash](https://learn.arm.com/tag/bash) [llama.cpp](https://learn.arm.com/tag/llama.cpp) |

### Who is this for?

This is an introductory topic for AI practitioners, performance engineers, and system architects who want to learn how to deploy and optimize quantized large language models (LLMs) on NVIDIA DGX Spark systems powered by the Grace-Blackwell (GB10) architecture.

### What will you learn?

Upon completion of this Learning Path, you will be able to:

- Describe the Grace–Blackwell (GB10) architecture and its support for efficient AI inference
- Build CUDA-enabled and CPU-only versions of llama.cpp for flexible deployment
- Validate the functionality of both builds on the DGX Spark platform
- Analyze how Armv9 SIMD instructions accelerate quantized LLM inference on the Grace CPU

### Prerequisites

Before starting, you will need the following:

- Access to an NVIDIA DGX Spark system with at least 15 GB of available disk space
- Familiarity with command-line interfaces and basic Linux operations
- Understanding of CUDA programming basics, as well as GPU and CPU compute concepts
- Basic knowledge of quantized large language models (LLMs) and machine learning inference
- Experience building software from source using CMake and make

### Summary

You’ll prepare NVIDIA DGX Spark with its Grace CPU and Blackwell GPU, build `llama.cpp` for CUDA and CPU execution, and inspect Armv9 vector instructions during quantized LLM inference. You’ll verify CUDA, compile both variants, and use Process Watch to examine Neon activity and the current lack of SVE and SVE2. You’ll finish with GPU and CPU binaries.

### Frequently asked questions

<details>
<summary>How do I know my DGX Spark is ready before building?</summary>
Confirm the Grace CPU configuration, operating system, Blackwell GPU, and CUDA drivers are active, and verify the CUDA 13 toolkit is installed.
</details>

<details>
<summary>Which build should I start with, GPU-enabled or CPU-only?</summary>
If the Blackwell GPU and CUDA are available, start with the GPU-enabled `llama.cpp` build. Then, build the CPU-only version to run on the Grace CPU and to keep a flexible deployment option.
</details>

<details>
<summary>What result should I expect after a successful build?</summary>
The build produces a compiled `llama.cpp` binary targeting either the GPU or the CPU that runs quantized LLM inference. A quick test run should complete without errors on the DGX Spark.
</details>

<details>
<summary>How do I validate that the CPU-only build uses Armv9 vector features?</summary>
Run an inference with the CPU-only binary and analyze it with Process Watch to observe the instruction mix. Expect to see Neon SIMD activity; SVE and SVE2 might remain inactive under the current kernel configuration.
</details>

<details>
<summary>Why don’t I see SVE or SVE2 instructions in my Process Watch results?</summary>
Although the Grace CPU supports SVE and SVE2, the current kernel exposes a fixed 16-byte (128-bit) SVE vector length. The `llama.cpp` workload therefore uses Neon instructions instead. Future kernel updates might enable SVE2 instructions.
</details>
