# Run batch inference using vLLM

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/)
- [Build a vLLM from source code on Arm Linux](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-setup/)
- [Run batch inference using vLLM](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-run/)
- [Run an OpenAI-compatible vLLM server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/vllm-server/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/vllm/_next-steps/)

## Use a model from Hugging Face
vLLM is designed to work with models from the Hugging Face Hub.

The first time you run vLLM, it downloads the required model. You don’t have to explicitly download any models.

To use a model that requires you to request access or accept terms and conditions, log in to Hugging Face using a token:

```bash
huggingface-cli login
```

Enter your Hugging Face token. You can generate a token from [Hugging Face Hub](https://huggingface.co/) by clicking your profile on the top right corner and selecting **Access Tokens**.

Visit the Hugging Face link printed in the login output and accept the terms and conditions. Click the **Agree and access repository** button or fill out the request-for-access form, depending on the model.

To run batched inference without logging in, use the `Qwen/Qwen2.5-0.5B-Instruct` model.

## Create a batch script
To run inference with multiple prompts, create a Python script to load a model and run the prompts.

Use a text editor to save the following Python script in a file called `batch.py`:

```python
import json
from vllm import LLM, SamplingParams

if __name__ == '__main__':
    # Sample prompts.
    prompts = [
        "Write a hello world program in C",
        "Write a hello world program in Java",
        "Write a hello world program in Rust",
    ]

    # Modify model here
    MODEL = "Qwen/Qwen2.5-0.5B-Instruct"

    # Create a sampling params object.
    sampling_params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=256)

    # Create an LLM.
    llm = LLM(model=MODEL, dtype="bfloat16", max_num_batched_tokens=32768)

    # Generate texts from the prompts. The output is a list of RequestOutput objects
    # that contain the prompt, generated text, and other information.
    outputs = llm.generate(prompts, sampling_params)

    # Print the outputs.
    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        result = {
            "Prompt": prompt,
            "Generated text": generated_text
        }
        print(json.dumps(result, indent=4))
```

The script uses `bfloat16` precision.

You can also change the length of the output using the `max_tokens` value.

Run the Python script:

```bash
python ./batch.py
```

The output shows vLLM starting, the model loading, and the batch processing of the three prompts:

```
__output__ INFO 10-23 18:38:40 [__init__.py:216] Automatically detected platform cpu.
__output__ INFO 10-23 18:38:42 [utils.py:233] non-default args: {'dtype': 'bfloat16', 'max_num_batched_tokens': 32768, 'disable_log_stats': True, 'model': 'Qwen/Qwen2.5-0.5B-Instruct'}
__output__ INFO 10-23 18:38:42 [model.py:547] Resolved architecture: Qwen2ForCausalLM
__output__ `torch_dtype` is deprecated! Use `dtype` instead!
__output__ INFO 10-23 18:38:42 [model.py:1510] Using max model len 32768
__output__ WARNING 10-23 18:38:42 [cpu.py:117] Environment variable VLLM_CPU_KVCACHE_SPACE (GiB) for CPU backend is not set, using 4 by default.
__output__ INFO 10-23 18:38:42 [arg_utils.py:1166] Chunked prefill is not supported for ARM and POWER and S390X CPUs; disabling it for V1 backend.
__output__ INFO 10-23 18:38:44 [__init__.py:216] Automatically detected platform cpu.
__output__ (EngineCore_DP0 pid=8933) INFO 10-23 18:38:46 [core.py:644] Waiting for init message from front-end.
__output__ (EngineCore_DP0 pid=8933) INFO 10-23 18:38:46 [core.py:77] Initializing a V1 LLM engine (v0.11.0) with config: model='Qwen/Qwen2.5-0.5B-Instruct', ...
```

## What you’ve accomplished and what’s next
You’ve now created a Python batch inference script that loads the `Qwen/Qwen2.5-0.5B-Instruct` model from Hugging Face, configures `bfloat16` precision, and sends multiple prompts to vLLM.

You ran the script and confirmed that vLLM starts on the CPU backend, loads the model, processes the prompts, and returns generated text.

Next, you’ll set up an OpenAI-compatible server so client applications can send requests to vLLM.
