# [Run a Large Language Model (LLM) chatbot with PyTorch using KleidiAI on Arm servers](https://learn.arm.com/learning-paths/servers-and-cloud-computing/pytorch-llama/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/pytorch-llama/)
- [Run a Large Language model (LLM) chatbot on Arm servers](https://learn.arm.com/learning-paths/servers-and-cloud-computing/pytorch-llama/pytorch-llama/)
- [Chatbot with Streamlit Frontend](https://learn.arm.com/learning-paths/servers-and-cloud-computing/pytorch-llama/pytorch-llama-frontend/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/pytorch-llama/_next-steps/)

## About this Learning Path

| Skill level:        | Introductory        |
|---------------------|---------------------|
| Reading time:       | 30 min              |
| Last updated:       | 31 Jul 2026         |

### Authors:
- Nikhil Gupta
- Pareena Verma, Arm [GitHub](https://github.com/pareenaverma) | [LinkedIn](https://linkedin.com/in/pareena-verma-7853607)
- Nobel Chowdary Mandepudi, Arm

**Arm IP:** [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors)

**Tags:** 
- [ML](https://learn.arm.com/tag/ml)
- [AWS](https://learn.arm.com/tag/aws)
- [Microsoft Azure](https://learn.arm.com/tag/microsoft-azure)
- [Google Cloud](https://learn.arm.com/tag/google-cloud)
- [Oracle](https://learn.arm.com/tag/oracle)
- [Linux](https://learn.arm.com/tag/linux)
- [LLM](https://learn.arm.com/tag/llm)
- [Generative AI](https://learn.arm.com/tag/generative-ai)
- [Python](https://learn.arm.com/tag/python)
- [PyTorch](https://learn.arm.com/tag/pytorch)
- [Hugging Face](https://learn.arm.com/tag/hugging-face)

### Who is this for?
This is an introductory topic for software developers interested in running LLMs using PyTorch on Arm-based servers.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Download the Meta Llama 3.1 model from the Meta Hugging Face repository.
- 4-bit quantize the model using optimized INT4 KleidiAI Kernels for PyTorch.
- Run an LLM inference using PyTorch on an Arm-based CPU.
- Expose an LLM inference as a browser application with Streamlit as the frontend and Torchchat framework in PyTorch as the LLM backend server.
- Measure performance metrics of the LLM inference running on an Arm-based CPU.

### Prerequisites
Before starting, you will need the following:
- An [Arm-based instance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/csp/) with at least 16 CPUs from a cloud service provider or an on-premise Arm server.
