# Build and run vLLM on Arm servers

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/)
- [Build a vLLM from source code on Arm Linux](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-setup/)
- [Run batch inference using vLLM](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-run/)
- [Run an OpenAI-compatible vLLM server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-server/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/_next-steps/)

## About this Learning Path

| Skill level:      | Introductory       |
|-------------------|--------------------|
| Reading time:     | 45 min             |
| Last updated:     | 31 Jul 2026        |

| Author:           | Jason Andrews, Arm [GitHub](https://github.com/jasonrandrews) [LinkedIn](https://linkedin.com/in/jason-andrews-7b05a8) |
|-------------------|--------------------|
| Arm IP:           | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:             | [ML](https://learn.arm.com/tag/ml) [AWS](https://learn.arm.com/tag/aws) [Microsoft Azure](https://learn.arm.com/tag/microsoft-azure) [Google Cloud](https://learn.arm.com/tag/google-cloud) [Oracle](https://learn.arm.com/tag/oracle) [Linux](https://learn.arm.com/tag/linux) [vLLM](https://learn.arm.com/tag/vllm) [LLM](https://learn.arm.com/tag/llm) [Generative AI](https://learn.arm.com/tag/generative-ai) [Python](https://learn.arm.com/tag/python) [Hugging Face](https://learn.arm.com/tag/hugging-face) |

### Who is this for?
This is an introductory topic for software developers and AI engineers interested in learning how to use the vLLM library on Arm servers.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Build vLLM from source on an Arm server.
- Use a Qwen LLM from Hugging Face.
- Run local batch inference using vLLM.
- Create and interact with an OpenAI-compatible server provided by vLLM on your Arm server.

### Prerequisites
Before starting, you will need the following:
- An [Arm-based Linux instance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/csp/) from a cloud service provider, or a local Arm Linux computer running Ubuntu 24.04 with at least 8 CPUs, 16 GB RAM, and 50 GB of disk storage.
- A system that includes support for BFloat16.
