# [Orchestrate a persistent local AI agent with Hermes on NVIDIA DGX Spark](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/)

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/)
- [Explore persistent AI runtime architecture on NVIDIA DGX Spark](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/1_introduce/)
- [Build the DGX Spark AI runtime foundation](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/2_build/)
- [Deploy Hermes Agent as an orchestration runtime](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/3_deploy_orch_runtime/)
- [Add local LLM inference to Hermes Agent](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/4_local_llm/)
- [Build persistent semantic memory for Hermes Agent](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/5_persistent_memory/)
- [Add semantic retrieval and contextual reasoning to Hermes Agent](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/6_semantic_retrieval/)
- [Add autonomous workspace cognition to Hermes Agent](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/7_autonomous_workspace/)
- [Next Steps](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_persistent_agent/_next-steps/)

## About this Learning Path

| Skill level: | Advanced |
|--------------|----------|
| Reading time: | 1 hr 30 min |
| Last updated: | 29 Jul 2026 |

| Author: | Odin Shen, Arm [GitHub](https://github.com/odincodeshen) [LinkedIn](https://linkedin.com/in/odin-shen-lmshen) |
|----------|---------------------------------------------------------------|
| Arm IP: | [Cortex-A](https://support.arm.com/?tab=compute-ip&Product%20Type=Application%20Processors) |
| Tags: | [ML](https://learn.arm.com/tag/ml) [Linux](https://learn.arm.com/tag/linux) [Python](https://learn.arm.com/tag/python) [Docker](https://learn.arm.com/tag/docker) [Ollama](https://learn.arm.com/tag/ollama) |

### Who is this for?
This is an advanced topic for developers building persistent local AI agent systems on NVIDIA DGX Spark who want to use Arm Grace CPUs for orchestration and Blackwell GPUs for local LLM inference and embeddings.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Describe how persistent AI runtimes combine orchestration, semantic memory, and local inference
- Build a continuously running local AI agent using Hermes Agent, Ollama, and Qdrant
- Use Arm Grace CPUs to orchestrate event-driven AI workflows on NVIDIA DGX Spark
- Deploy semantic memory and contextual retrieval pipelines using vector embeddings and Qdrant

### Prerequisites
Before starting, you will need the following:
- An NVIDIA DGX Spark system with at least 15 GB of available disk space
- Familiarity with running Python scripts and basic Docker container workflows

### Summary
You’ll build a continuously running local AI agent on NVIDIA DGX Spark using Arm Grace CPUs for orchestration. You’ll configure Docker services with a persistent workspace, deploy Ollama, Qdrant, Open WebUI, and Hermes Agent. Then, you’ll connect Hermes to Ollama for document summaries and Qdrant embeddings. By the end, the services will monitor files, summarize documents, and provide contextual retrieval.

### Frequently asked questions
<details>
<summary>How do I know the base DGX Spark AI runtime is running correctly?</summary>
You’ll see containers for Ollama (inference), Qdrant (vector memory), and Open WebUI (browser access) running and healthy. Confirm that the persistent workspace exists and is mounted as expected.
</details>

<details>
<summary>Where should I put documents so Hermes picks them up automatically?</summary>
Place files in `workspace/inbox/`. Hermes watches that path and handles `on_created()` events to start the workflow.
</details>

<details>
<summary>I added Hermes but I don’t see summaries yet — what should I expect at this stage?</summary>
Not seeing summaries is expected. In its initial setup, Hermes acts as an orchestration and event layer, printing handling output but not invoking a language model until you connect it to Ollama.
</details>

<details>
<summary>After I connect Hermes to Ollama, what result should I expect to confirm inference is working?</summary>
When you add a new document to `workspace/inbox/`, Hermes sends the content to the local model and you see an AI-generated summary in the runtime output. This indicates the inference path is wired correctly.
</details>

<details>
<summary>How can I confirm persistent semantic memory is active in Qdrant?</summary>
Qdrant stores an embedding for each processed document. Check that new vector entries appear after Hermes handles files and are available for contextual retrieval in the workflow.
</details>
