Who is this for?

This is an introductory topic for DevOps engineers, AI engineers, machine learning engineers, and software developers who want to build retrieval-augmented generation (RAG) applications using LlamaIndex on SUSE Linux Enterprise Server (SLES) Arm64, integrate vector databases, and query custom documents using local large language models (LLMs).

What will you learn?

Upon completion of this Learning Path, you will be able to:

  • Install and configure LlamaIndex on Google Cloud C4A Axion processors for Arm64.
  • Build indexing and retrieval pipelines using LlamaIndex.
  • Integrate ChromaDB vector databases with local LLMs using Ollama.
  • Build and test a browser-based RAG application using FastAPI.

Prerequisites

Before starting, you will need the following:

Summary

AI-assisted

This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
You’ll build a RAG application on a Google Axion C4A virtual machine (VM). First, you’ll configure SUSE Linux, Python 3.11, LlamaIndex, Ollama, and ChromaDB. Then, you’ll create a FastAPI backend and browser interface. You’ll add sample documents, connect the indexing and retrieval pipeline, and submit browser queries. Finally, you’ll observe how LlamaIndex retrieves context from ChromaDB and sends it to your local model for grounded responses.

Frequently asked questions

AI-assisted

These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
Which firewall port should be opened for the browser-based app?
Open TCP port 8000. After you create the rule, verify that it applies to the VM’s VPC network, then open http://<VM-EXTERNAL-IP>:8000 after the app starts.
Which C4A machine type should I use?
Select c4a-standard-4, which provides four vCPUs and 16 GB of memory.
Which Python version should I install?
Install Python 3.11 after you refresh and update the system packages.
How do I know that the RAG pipeline is wired correctly?
Submit a query from the browser interface and confirm that the response reflects your sample documents. Your FastAPI backend should call LlamaIndex, retrieve context from ChromaDB, and send that context to the local LLM through Ollama.
What should I check if the browser UI doesn't load?
Confirm that the FastAPI server is running, the VM’s external IP is reachable, and the firewall rule for port 8000 is active. After confirming, open http://<VM-EXTERNAL-IP>:8000 again.
Next