# Deploy a Large Language Model (LLM) chatbot with llama.cpp using KleidiAI on Arm servers

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/)
- [Demo](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/_demo/)
- [Run a large language model chatbot on Arm servers](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/llama-chatbot/)
- [Access the chatbot using the OpenAI-compatible API](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/llama-server/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/_next-steps/)

## About this Learning Path

| Skill level:        | Introductory        |
|---------------------|---------------------|
| Reading time:       | 30 min              |
| Last updated:       | 17 Sep 2026         |

### Authors:
- Pareena Verma, Arm [GitHub](https://github.com/pareenaverma) | [LinkedIn](https://linkedin.com/in/pareena-verma-7853607)
- Jason Andrews, Arm [GitHub](https://github.com/jasonrandrews) | [LinkedIn](https://linkedin.com/in/jason-andrews-7b05a8)
- Zach Lasiuk

### Arm IP:
[Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors)

### Tags:
- [ML](https://learn.arm.com/tag/ml)
- [AWS Graviton](https://learn.arm.com/tag/aws-graviton)
- [Linux](https://learn.arm.com/tag/linux)
- [LLM](https://learn.arm.com/tag/llm)
- [Generative AI](https://learn.arm.com/tag/generative-ai)
- [Python](https://learn.arm.com/tag/python)
- [Demo](https://learn.arm.com/tag/demo)
- [Hugging Face](https://learn.arm.com/tag/hugging-face)

### Who is this for?
This is an introductory topic for developers interested in running LLMs on Arm-based servers.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Download and build llama.cpp on your Arm server.
- Download a pre-quantized Llama 3.1 model from Hugging Face.
- Run the pre-quantized model on your Arm CPU and measure the performance.

### Prerequisites
Before starting, you will need the following:
- An AWS Graviton4 r8g.16xlarge instance to test Arm performance optimizations, or any [Arm based instance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/csp/) from a cloud service provider or an on-premise Arm server.

### Summary
You’ll deploy a persistent LLM chatbot on an Arm server with `llama.cpp` and a pre-quantized Llama 3.1 8B model from Hugging Face. First, you’ll build `llama.cpp`, obtain the model, launch its OpenAI-compatible server, and expose it on port `8080`. You’ll then access the chatbot using the OpenAI-compatible API.

### Frequently asked questions
<details>
<summary>What result should I expect when I start the llama.cpp server?</summary>
The server starts and listens on port `8080`. After the server starts running, you can send OpenAI-compatible requests without restarting the process between calls.
</details>

<details>
<summary>Do I need any extra tools to view API responses?</summary>
Yes. Install `jq` with `sudo apt install jq -y`. You’ll use `jq` to process JSON returned by the API.
</details>

<details>
<summary>Can I access the chatbot from another machine?</summary>
Yes. The server exposes an OpenAI-compatible API over the network, so a remote client can call the host running the LLM if it can reach port `8080`.
</details>

<details>
<summary>Which model should I download before launching the server?</summary>
Use a pre-quantized Llama 3.1 8B model from Hugging Face. Download the model to the Arm server before starting the server.
</details>

<details>
<summary>How do I send a request to the running llama.cpp server?</summary>
Send a `curl` request to `http://localhost:8080/v1/chat/completions` with a JSON prompt, then pipe the response to `jq -C`. Save the request in `curl-test.sh` and run it with `bash ./curl-test.sh`.
</details>
