# Run a DeepSeek-R1 chatbot on Arm servers

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/deepseek-cpu/)
- [Run a DeepSeek-R1 chatbot on Arm servers](https://learn.arm.com/learning-paths/servers-and-cloud-computing/deepseek-cpu/deepseek-chatbot/)
- [Access the chatbot using the OpenAI-compatible API](https://learn.arm.com/learning-paths/servers-and-cloud-computing/deepseek-cpu/deepseek-server/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/deepseek-cpu/_next-steps/)

## Overview and what you’ll build

The instructions in this Learning Path are for any Arm server running Ubuntu 24.04 LTS. You need an Arm server instance with at least 64 cores and 512GB of RAM to run this example. Configure disk storage up to at least 400 GB. The instructions have been tested on an AWS Graviton4 r8g.24xlarge instance.

Arm CPUs are increasingly used for machine learning and AI inference workloads due to their efficiency and scalability. In this Learning Path, you’ll deploy a generative AI chatbot using the [DeepSeek-R1 671B LLM](https://huggingface.co/bartowski/DeepSeek-R1-GGUF) on your Arm-based CPU, leveraging the `llama.cpp` inference engine optimized for Arm architecture.

You’ll learn how to do the following:

- Build and run `llama.cpp` with Arm-specific performance optimizations.
- Download a quantized GGUF model from Hugging Face.
- Run and benchmark inference performance on a large Arm instance, such as AWS Graviton4.

[`llama.cpp`](https://github.com/ggerganov/llama.cpp) is an open source C/C++ project developed by Georgi Gerganov that enables efficient LLM inference on a variety of hardware - both locally, and in the cloud.

## Understanding the DeepSeek-R1 model and GGUF format

The [DeepSeek-R1 model](https://huggingface.co/deepseek-ai/DeepSeek-R1) from DeepSeek-AI available on Hugging Face, is released under the [MIT License](https://github.com/deepseek-ai/DeepSeek-R1/blob/main/LICENSE) and free to use for research and commercial purposes.

The DeepSeek-R1 model has 671 billion parameters, based on Mixture of Experts (MoE) architecture. This improves inference speed and maintains model quality. For this example, the full 671 billion (671B) model is used for retaining quality chatbot capability while also running efficiently on your Arm-based CPU.

Traditionally, the training and inference of LLMs has been done on GPUs using full-precision 32-bit (FP32) or half-precision 16-bit (FP16) data type formats for the model parameter and weights. Recently, a new binary model format called GGUF was introduced by the `llama.cpp` team. This new GGUF model format uses compression and quantization techniques that remove the dependency on using FP32 and FP16 data type formats. For example, GGUF supports quantization where model weights that are generally stored as FP16 data types are scaled down to 4-bit integers. This significantly reduces the need for computational resources and the amount of RAM required. These advancements make Arm CPUs a strong fit for running LLM inference workloads.

## Install build dependencies on your Arm-based server

Install the following packages:

```
sudo apt update
sudo apt install make cmake -y
```

You also need to install `gcc` on your machine:

```
sudo apt install gcc g++ -y
sudo apt install build-essential -y
```

## Clone and build llama.cpp

You are now ready to start building `llama.cpp`.

Clone the source repository for llama.cpp:

```
git clone https://github.com/ggerganov/llama.cpp
```

By default, `llama.cpp` builds for CPU only. You don’t need to provide any extra switches to build it for the Arm CPU that you run it on.

Run `cmake` to build it:

```
cd llama.cpp
mkdir build
cd build
cmake .. -DCMAKE_CXX_FLAGS="-mcpu=native" -DCMAKE_C_FLAGS="-mcpu=native"
cmake --build . --config Release -v -j $(nproc)
```

`llama.cpp` is now built in the `bin` directory. Check that `llama.cpp` has built correctly by running the help command:

```
cd bin
./llama-cli -h
```

If `llama.cpp` has built correctly on your machine, you will see the help options being displayed. A snippet of the output is shown below:

```
__output__ usage: ./llama-cli [options]
__output__ 
__output__ general:
__output__ 
  -h,    --help, --usage          print usage and exit
         --version                show version and build info
  -v,    --verbose                print verbose information
         --verbosity N            set specific verbosity level (default: 0)
         --verbose-prompt         print a verbose prompt before generation (default: false)
         --no-display-prompt      don't print prompt at generation (default: false)
  -co,   --color                  colorise output to distinguish prompt and user input from generations (default: false)
  -s,    --seed SEED              RNG seed (default: -1, use random seed for < 0)
  -t,    --threads N              number of threads to use during generation (default: 4)
  -tb,   --threads-batch N        number of threads to use during batch and prompt processing (default: same as --threads)
  -td,   --threads-draft N        number of threads to use during generation (default: same as --threads)
  -tbd,  --threads-batch-draft N  number of threads to use during batch and prompt processing (default: same as --threads-draft)
         --draft N                number of tokens to draft for speculative decoding (default: 5)
  -ps,   --p-split N              speculative decoding split probability (default: 0.1)
  -lcs,  --lookup-cache-static FNAME
                                  path to static lookup cache to use for lookup decoding (not updated by generation)
  -lcd,  --lookup-cache-dynamic FNAME
                                  path to dynamic lookup cache to use for lookup decoding (updated by generation)
  -c,    --ctx-size N             size of the prompt context (default: 0, 0 = loaded from model)
  -n,    --predict N              number of tokens to predict (default: -1, -1 = infinity, -2 = until context filled)
  -b,    --batch-size N           logical maximum batch size (default: 2048)
```

## Set up Hugging Face and download the model

There are a few different ways you can download the DeepSeek-R1 model. In this Learning Path, you download the model from Hugging Face.

[Hugging Face](https://huggingface.co/) is an open source AI community where you can host your own AI models, train them and collaborate with others in the community. You can browse through the thousands of models that are available for a variety of use cases like NLP, audio, and computer vision.

The `huggingface_hub` library provides APIs and tools that let you easily download and fine-tune pre-trained models. You will use `huggingface-cli` to download the [DeepSeek-R1 model](https://huggingface.co/bartowski/DeepSeek-R1-GGUF).

Install the required Python packages:

```
sudo apt install python-is-python3 python3-pip python3-venv -y
```

Create and activate a Python virtual environment:

```
python -m venv venv
source venv/bin/activate
```

Your terminal prompt now has the `(venv)` prefix indicating the virtual environment is active. Use this virtual environment for the remaining commands.

Install the `huggingface_hub` python library using `pip`:

```
pip install huggingface_hub
```

You can now download the model using the huggingface cli:

```
huggingface-cli download bartowski/DeepSeek-R1-GGUF --include "*DeepSeek-R1-Q4_0*"  --local-dir DeepSeek-R1-Q4_0
```

Before you proceed and run this model, take a quick look at what `Q4_0` in the model name denotes.

## Understanding the Quantization format

`Q4_0` in the model name refers to the quantization method the model uses. The goal of quantization is to reduce the size of the model (to reduce the memory space required) and faster (to reduce memory bandwidth bottlenecks transferring large amounts of data from memory to a processor). The primary trade-off to keep in mind when reducing a model’s size is maintaining quality of performance. Ideally, a model is quantized to meet size and speed requirements while not having a negative impact on performance.

This model is `DeepSeek-R1-Q4_0-00001-of-00010.gguf`, so what does each component mean in relation to the quantization level? The main thing to note is the number of bits per parameter, which is denoted by ‘Q4’ in this case or 4-bit integer. As a result, by only using 4 bits per parameter for 671 billion parameters, the model drops to be 354 GB in size.

## Run the DeepSeek-R1 Chatbot on your Arm server

As of [llama.cpp commit 0f1a39f3](https://github.com/ggerganov/llama.cpp/commit/0f1a39f3), Arm has contributed code for performance optimization with three types of GEMV/GEMM kernels corresponding to three processor types:

- AWS Graviton2, where you only have Neon support (you will see less improvement for these GEMV/GEMM kernels).
- AWS Graviton3, where the GEMV/GEMM kernels exploit both SVE 256 and MATMUL INT8 support.
- AWS Graviton4, where the GEMV/GEMM kernels exploit Neon/SVE 128 and MATMUL_INT8 support.

With the latest commits in `llama.cpp` you will see improvements for these Arm optimized kernels directly on your Arm-based server. You can run the pre-quantized Q4_0 model as is and do not need to re-quantize the model.

Run the pre-quantized DeepSeek-R1 model exactly as the weights were downloaded from huggingface:

```
./llama-cli -m DeepSeek-R1-Q4_0/DeepSeek-R1-Q4_0/DeepSeek-R1-Q4_0-00001-of-00010.gguf -no-cnv --temp 0.6 -t 64 --prompt "<|User|>Building a visually appealing website can be done in ten simple steps:<|Assistant|>" -n 512
```

This command will use the downloaded model (`-m` flag), disable conversation mode explicitly (`-no-cnv` flag), adjust the randomness of the generated text (`--temp` flag), with the specified prompt (`-p` flag), and target a 512 token completion (`-n` flag), using 64 threads (`-t` flag).

You might notice there are many gguf files. Llama.cpp can load all series of files by passing the first one with `-m` flag.

## Analyze the output and performance statistics

You will see lots of interesting statistics being printed from `llama.cpp` about the model and the system, followed by the prompt and completion. The tail of the output from running this model on an AWS Graviton4 r8g.24xlarge instance is shown below:

```
__output__ build: 4963 (02082f15) with cc (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0 for aarch64-linux-gnu
__output__ main: llama backend init
__output__ main: load the model and apply lora adapter, if any
__output__ llama_model_loader: additional 9 GGUFs metadata loaded.
__output__ llama_model_loader: loaded meta data with 51 key-value pairs and 1025 tensors from /home/ubuntu/DeepSeek-R1-Q4_0/DeepSeek-R1-Q4_0/DeepSeek-R1-Q4_0-00001-of-00010.gguf (version GGUF V3 (latest))
...
```

The `system_info` printed from `llama.cpp` highlights important architectural features present on your hardware that improve the performance of the model execution. In the output shown above from running on an AWS Graviton4 instance, you will see:

- Neon = 1 This flag indicates support for Arm’s Neon technology which is an implementation of the Advanced SIMD instructions.
- ARM_FMA = 1 This flag indicates support for Arm Floating-point Multiply and Accumulate instructions.
- MATMUL_INT8 = 1 This flag indicates support for Arm int8 matrix multiplication instructions.
- SVE = 1 This flag indicates support for the Arm Scalable Vector Extension.

The end of the output shows several model timings:

- Load time refers to the time taken to load the model.
- Prompt eval time refers to the time taken to process the prompt before generating the new text.
- Eval time refers to the time taken to generate the output. Generally, anything above 10 tokens per second is faster than what humans can read.

## What’s next?

You’ve successfully run a large-scale LLM chatbot on an Arm server with KleidiAI optimizations. Continue experimenting with different prompts, quantization levels, or deployment methods.
