Who is this for?

This is an introductory topic for software developers who want to create a retrieval-augmented generation (RAG) application on Arm servers.

What will you learn?

Upon completion of this Learning Path, you will be able to:

  • Create a RAG application using Zilliz Cloud.
  • Launch a large language model (LLM) service on Arm servers.

Prerequisites

Before starting, you will need the following:

  • A basic understanding of a RAG pipeline
  • An AWS Graviton3 C7g.2xlarge instance, or any Arm-based instance from a cloud service provider or an on-premise Arm server
  • A Zilliz account , which you can sign up for with a free trial

Summary

AI-assisted

This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
You’ll build a RAG workflow on Arm with Zilliz Cloud and llama.cpp. First, you’ll create a dedicated Zilliz Cloud cluster on Arm-based machines, then build a local llama.cpp server with an OpenAI-compatible API. You’ll use Python to create and inspect an embedding, connect to the local model endpoint, and issue test requests. Finally, you’ll validate vector search in Zilliz Cloud and local model inference on your Arm server.

Frequently asked questions

AI-assisted

These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
Which Zilliz Cloud cluster should I create?
Create a Dedicated cluster on AWS using Arm-based machines. You can also use self-hosted Milvus, although its setup is more involved.
How do I know my Zilliz Cloud cluster is ready before I continue?
After you select Create Cluster, check that the cluster appears in your Default Project with a running status. Continue after the cluster reports that it’s running.
Do I need an API key to call the LLM service from my script?
No. You’ll run the llama.cpp server locally, so the OpenAI SDK can connect without a real API key.
How do I confirm that the embedding model is working correctly?
Run the test code and confirm that it prints an embedding dimension and several floating-point values. The example reports a dimension of 384.
What do I need before I can use the Llama 3.1 model with llama.cpp?
Request access to Llama 3.1 through the Llama website. After you receive access, follow the instructions to host the model with llama.cpp on your Arm-based server.
Next