Build a RAG application using Zilliz Cloud on Arm servers
Introduction
Overview and Install dependencies
Offline Data Loading
Launch the LLM Server
Online RAG
Next Steps
Build a RAG application using Zilliz Cloud on Arm servers
Who is this for?
This is an introductory topic for software developers who want to create a retrieval-augmented generation (RAG) application on Arm servers.
What will you learn?
Upon completion of this Learning Path, you will be able to:
- Create a RAG application using Zilliz Cloud.
- Launch a large language model (LLM) service on Arm servers.
Prerequisites
Before starting, you will need the following:
- A basic understanding of a RAG pipeline
- An AWS Graviton3 C7g.2xlarge instance, or any Arm-based instance from a cloud service provider or an on-premise Arm server
- A Zilliz account , which you can sign up for with a free trial
Summary
This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
You’ll build a RAG workflow on Arm with Zilliz Cloud and llama.cpp. First, you’ll create a dedicated Zilliz Cloud cluster on Arm-based machines, then build a local llama.cpp server with an OpenAI-compatible API. You’ll use Python to create and inspect an embedding, connect to the local model endpoint, and issue test requests. Finally, you’ll validate vector search in Zilliz Cloud and local model inference on your Arm server.
Frequently asked questions
These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
Create a Dedicated cluster on AWS using Arm-based machines. You can also use self-hosted Milvus, although its setup is more involved.
After you select Create Cluster, check that the cluster appears in your Default Project with a running status. Continue after the cluster reports that it’s running.
No. You’ll run the llama.cpp server locally, so the OpenAI SDK can connect without a real API key.
Run the test code and confirm that it prints an embedding dimension and several floating-point values. The example reports a dimension of
384.Request access to Llama 3.1 through the Llama website. After you receive access, follow the instructions to host the model with llama.cpp on your Arm-based server.