# Deploy a mixed-placement AI shopping assistant on Google Kubernetes Engine with Axion-based compute

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/)
- [Understand mixed placement for a storefront AI assistant](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/background/)
- [Set up the source tree and cluster access](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/project-setup/)
- [Deploy and validate the storefront baseline on the Google N4A node pool](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/validate-n4a-baseline/)
- [Review the shopping assistant implementation](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/review-assistant/)
- [Build and push the assistant image to Artifact Registry](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/build-assistant-image/)
- [Deploy the assistant on the Google N4A node pool](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/deploy-assistant-n4a/)
- [Observe and benchmark the assistant on the Google N4A node pool](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/observe-benchmark-n4a/)
- [Move the assistant to the Google C4A node pool and compare results](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/move-assistant-c4a/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/storefront-ai-assistant-gke-axion/_next-steps/)

## About this Learning Path

| Skill level:         | Advanced        |
|----------------------|------------------|
| Reading time:        | 2 hrs            |
| Last updated:        | 04 Aug 2026      |

| Author:              | Rani Chowdary Mandepudi, Arm |
|----------------------|-------------------------------|
| Arm IP:              | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:                | [Containers and Virtualization](https://learn.arm.com/tag/containers-and-virtualization), [Google Cloud](https://learn.arm.com/tag/google-cloud), [Linux](https://learn.arm.com/tag/linux), [Kubernetes](https://learn.arm.com/tag/kubernetes), [GKE](https://learn.arm.com/tag/gke), [Docker](https://learn.arm.com/tag/docker), [Kustomize](https://learn.arm.com/tag/kustomize), [Python](https://learn.arm.com/tag/python), [Ollama](https://learn.arm.com/tag/ollama), [Gemma](https://learn.arm.com/tag/gemma) |

## Who is this for?
This is an advanced topic for cloud developers, platform engineers, and site reliability engineers who run applications on Google Kubernetes Engine (GKE) and want to place application tiers on the Axion-based machine series that fits each workload.

## What will you learn?
Upon completion of this Learning Path, you will be able to:
- Deploy and validate an Online Boutique storefront on an N4A node pool
- Build and push a `linux/arm64` container image, then add the AI shopping assistant to the storefront
- Use Kustomize overlays to run the assistant on N4A first, then move it to C4A
- Capture and compare benchmark summaries for the same assistant workload on N4A and C4A

## Prerequisites
Before starting, you will need the following:
- A [Google Cloud account](https://console.cloud.google.com/) with billing enabled
- Access to a [GKE Standard cluster with Arm node pools](https://cloud.google.com/kubernetes-engine/docs/how-to/create-arm-clusters-nodes), including N4A and C4A node pools, with the Kubernetes Metrics API enabled
- Permissions to get cluster credentials, deploy Kubernetes workloads and services, read pod logs and metrics, and create or use an Artifact Registry Docker repository
- Cloud Shell or a Linux or macOS administrative workstation with Docker Buildx, `gcloud`, `kubectl`, `git`, `curl`, Python 3.10 or later, and `jq`
- Basic familiarity with Docker, Kubernetes, Kustomize, and GKE

## Summary
You’ll deploy the Online Boutique storefront on Google Kubernetes Engine using Arm-based Axion nodes, validate a baseline on N4A, and add a gRPC-driven AI shopping assistant. You’ll build and push a single `linux/arm64` container image to Artifact Registry, then use Kustomize overlays to run the assistant on N4A before moving only that tier to C4A. After reviewing the assistant’s source code and runtime dependencies, you’ll confirm scheduling on the intended node pool and capture benchmark summaries to compare the same assistant workload across N4A and C4A. You’ll finish with a mixed-placement deployment where the steady storefront remains on N4A and the burstier assistant runs on the selected pool.

## Frequently asked questions
<details>
<summary>How do I verify the cluster has both N4A and C4A node pools before I start?</summary>
Use `kubectl` to list nodes and confirm that both pools are present. The workflow assumes an `arm64` GKE Standard cluster with separate N4A and C4A pools.
</details>

<details>
<summary>What result should I expect after I apply the baseline overlay?</summary>
The storefront runs on N4A, and `shoppingassistantservice` isn’t present. This is intentional because you’ll build and deploy the assistant in later steps.
</details>

<details>
<summary>I've run this path before. What should I remove before I recreate the baseline?</summary>
Delete any existing assistant deployment, related service, and service account so the baseline reflects a storefront without the assistant.
</details>

<details>
<summary>Do I need different container images for N4A and C4A when I deploy the assistant?</summary>
No. You’ll build one `linux/arm64` image targeted for Axion that’ll run in either placement.
</details>

<details>
<summary>How do I confirm the assistant is scheduled on the intended node pool when I switch from N4A to C4A?</summary>
Check the node assigned to the assistant pod and verify it matches the target pool after applying the appropriate Kustomize overlay. Inspect pod logs and service reachability to confirm the tier is healthy before capturing benchmarks.
</details>
