Deploy a mixed-placement AI shopping assistant on Google Kubernetes Engine with Axion-based compute
Introduction
Understand mixed placement for a storefront AI assistant
Set up the source tree and cluster access
Deploy and validate the storefront baseline on the Google N4A node pool
Review the shopping assistant implementation
Build and push the assistant image to Artifact Registry
Deploy the assistant on the Google N4A node pool
Observe and benchmark the assistant on the Google N4A node pool
Move the assistant to the Google C4A node pool and compare results
Next Steps
Deploy a mixed-placement AI shopping assistant on Google Kubernetes Engine with Axion-based compute
Introduction
Understand mixed placement for a storefront AI assistant
Set up the source tree and cluster access
Deploy and validate the storefront baseline on the Google N4A node pool
Review the shopping assistant implementation
Build and push the assistant image to Artifact Registry
Deploy the assistant on the Google N4A node pool
Observe and benchmark the assistant on the Google N4A node pool
Move the assistant to the Google C4A node pool and compare results
Next Steps
How the application is organized
Modern AI applications often contain more than one workload shape. For example, a storefront is steady, service-oriented, and always on. An AI assistant is burstier and more latency-sensitive because it runs only when users ask for help.
You’ll use this split in a live Online Boutique storefront on Google Kubernetes Engine (GKE). The storefront starts on N4A nodes powered by Google Axion processors and Arm Neoverse N3. You’ll add the missing shoppingassistantservice, run it on N4A first, and then move only that assistant tier to C4A nodes powered by Google Axion and Arm Neoverse V2.
The goal isn’t to prove that one Axion-based machine series replaces the other. You’ll use N4A for the steady storefront tier and evaluate C4A for the AI reasoning tier so you can decide which machine series fits each workload.
Mixed-placement agentic storefront on Axion
The diagram shows the final mixed-placement pattern you’ll build toward. It summarizes the multi-architecture image publishing process used for the existing storefront images; you don’t repeat that process here. The assistant image you’ll build later targets only linux/arm64 because both destination node pools are Arm-based. The diagram also shows the final placement: core services remain on N4A while the assistant, its Ollama sidecar, and the Gemma model move together to C4A.
How the storefront already runs on Arm
The storefront baseline uses container images that can run on Arm nodes. A common way to publish portable container images is to use a multi-architecture image, which is one image reference that contains variants for more than one CPU architecture, such as linux/amd64 and linux/arm64.
When Kubernetes schedules a pod that uses a multi-architecture image on an Arm node, the container runtime pulls the Arm-compatible variant from the same image reference. That’s why the storefront can already run on Axion nodes before you build the assistant image. You’ll confirm the storefront image reference from the source tree after you set up your tools and cluster access.
You’ll start with Arm-compatible storefront images already available. To learn the full build-and-publish workflow for multi-architecture images on GKE, see Migrate x86 workloads to Arm on Google Kubernetes Engine with Axion processors .
How assistant requests flow
The shoppingassistantservice is the AI layer for the application. It runs as its own Kubernetes service and handles requests sent through the storefront frontend.
When a shopper uses the assistant, the request follows this path:
- The browser sends the request to
frontend. frontendforwards the request toshoppingassistantservice.- The assistant calls
productcatalogserviceorcartservicethrough its application tools. - The assistant sends the live storefront context to an Ollama sidecar in the same pod. The sidecar runs the
gemma3:1b-it-qatmodel for the reasoning step. - The assistant returns the response through
frontend.
The assistant is agentic because it does more than generate text. It uses application tools, keeps short-lived session state, and asks for confirmation before it calls cartservice to update the cart.
The search_catalog and get_product_details tools query the live catalog. The get_cart tool reads the current cart, while add_to_cart updates it only after user confirmation.
This design keeps the placement comparison focused. You don’t add a separate vector database, retrieval pipeline, or hosted large language model API. When you move the assistant pod, you move the assistant logic and local reasoning runtime together as one AI tier.
What you’ve learned and what’s next
You’ve now learned about the architecture of the storefront application and its AI assistant.
Next, you’ll set up environment variables, connect to the GKE cluster, and clone the source tree that contains the assistant implementation.