# [Deploy a RAG-based Chatbot with llama-cpp-python using KleidiAI on Google Axion processors](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/)
- [Demo](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/_demo/)
- [Set up a RAG based LLM Chatbot](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/rag_llm/)
- [Deploy a RAG-based LLM backend server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/backend/)
- [Deploy RAG-based LLM frontend server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/frontend/)
- [The RAG Chatbot and its Performance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/chatbot/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/rag/_next-steps/)

## About this Learning Path

| Skill level:         | Advanced           |
|----------------------|--------------------|
| Reading time:        | 45 min             |
| Last updated:        | 31 Jul 2026        |

| Author:              | Nobel Chowdary Mandepudi, Arm   |
|----------------------|----------------------------------|
| Arm IP:              | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors)  |
| Tags:                | ML, [Google Cloud](https://learn.arm.com/tag/google-cloud), [Linux](https://learn.arm.com/tag/linux), [Python](https://learn.arm.com/tag/python), [Streamlit](https://learn.arm.com/tag/streamlit), [Google Axion](https://learn.arm.com/tag/google-axion), [Demo](https://learn.arm.com/tag/demo), [Hugging Face](https://learn.arm.com/tag/hugging-face) |

### Who is this for?

This Learning Path is for software developers, ML engineers, and those looking to deploy production-ready LLM chatbots with Retrieval Augmented Generation (RAG) capabilities, knowledge base integration, and performance optimization for Arm Architecture.

### What will you learn?

Upon completion of this Learning Path, you will be able to:

- Set up llama-cpp-python optimized for Arm servers.
- Implement RAG architecture using the Facebook AI Similarity Search (FAISS) vector database.
- Optimize model performance through 4-bit quantization.
- Build a web interface for document upload and chat.
- Monitor and analyze inference performance metrics.

### Prerequisites

Before starting, you will need the following:

- A Google Cloud Axion (or other Arm) compute instance with at least 16 cores, 8GB of RAM, and 32GB disk space.
- Basic understanding of Python and ML concepts.
- Familiarity with REST APIs and web services.
- Basic knowledge of vector databases.
- Understanding of LLM fundamentals.
