# [Run distributed inference with llama.cpp on Arm-based AWS Graviton4 instances](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/)
- [Convert model to GGUF and quantize](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/how-to-1/)
- [Configure the worker nodes](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/how-to-2/)
- [Configure the master node](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/how-to-3/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/_next-steps/)

## About this Learning Path

| Skill level:          | Introductory          |
|-----------------------|----------------------|
| Reading time:         | 30 min               |
| Last updated:         | 31 Jul 2026          |

| Authors:              |
|-----------------------|
| Aryan Bhusari, Arm [LinkedIn](https://linkedin.com/in/https://www.linkedin.com/in/aryanbhusari)<br/>Joe Stech, Arm [GitHub](https://github.com/JoeStech) [LinkedIn](https://linkedin.com/in/joestech) |
| Arm IP:               | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:                 | [ML](https://learn.arm.com/tag/ml) [AWS](https://learn.arm.com/tag/aws) [Linux](https://learn.arm.com/tag/linux) [LLM](https://learn.arm.com/tag/llm) [Generative AI](https://learn.arm.com/tag/generative-ai) |

### Who is this for?
This introductory topic is for developers with some experience using llama.cpp who want to learn how to run distributed inference on Arm-based servers.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Set up a main host and worker nodes with llama.cpp
- Run a large quantized model (for example, Llama 3.1 405B) with distributed CPU inference on Arm machines

### Prerequisites
Before starting, you will need the following:
- Three AWS c8g.4xlarge instances with at least 500 GB of EBS storage
- Python 3 installed on each instance
- Access to Meta’s gated repository for the Llama 3.1 model family and a Hugging Face token to download models
- Familiarity with the Learning Path [Deploy a Large Language Model (LLM) chatbot with llama.cpp using KleidiAI on Arm servers](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-cpu/)
- Familiarity with AWS

### Summary
You’ll run distributed CPU inference with `llama.cpp` on AWS Graviton4 instances. You’ll convert Meta’s Llama 3.1 70B safetensors shards to one GGUF file, quantize the weights to 4-bit, configure worker RPC backends and a master through a comma-separated `worker_ips` list, and verify connectivity with `telnet`. You’ll then launch distributed inference across the Arm servers.

### Frequently asked questions

<details>
<summary>How do I pass the worker node addresses to the master?</summary>
Export the `worker_ips` environment variable with a comma-separated list of `host:port` entries, for example, `172.31.110.11:50052,172.31.110.12:50052`. Run this on the master node before starting inference.
</details>

<details>
<summary>How do I verify the master can reach a worker before running a distributed job?</summary>
From the master, run `telnet <worker_ip> 50052`. A successful connection confirms that the backend server on the worker is reachable.
</details>

<details>
<summary>Which model format should I use with llama.cpp in this path?</summary>
Convert Meta’s safetensors files for Llama 3.1 70B to a single GGUF file, then quantize to 4-bit GGUF. Use the quantized GGUF for inference.
</details>

<details>
<summary>How do I confirm the quantization step worked?</summary>
The process should create a new 4-bit GGUF weights file that is smaller than the 16-bit GGUF. Use this 4-bit file for the distributed run.
</details>

<details>
<summary>How many nodes are used in the example and what roles do they serve?</summary>
The example uses three AWS Graviton4 instances: one master node and two worker nodes. The master coordinates inference while the workers run the RPC backend.
</details>
