Deploy a mixed-placement AI shopping assistant on Google Kubernetes Engine with Axion-based compute
Introduction
Understand mixed placement for a storefront AI assistant
Set up the source tree and cluster access
Deploy and validate the storefront baseline on the Google N4A node pool
Review the shopping assistant implementation
Build and push the assistant image to Artifact Registry
Deploy the assistant on the Google N4A node pool
Observe and benchmark the assistant on the Google N4A node pool
Move the assistant to the Google C4A node pool and compare results
Next Steps
Deploy a mixed-placement AI shopping assistant on Google Kubernetes Engine with Axion-based compute
Introduction
Understand mixed placement for a storefront AI assistant
Set up the source tree and cluster access
Deploy and validate the storefront baseline on the Google N4A node pool
Review the shopping assistant implementation
Build and push the assistant image to Artifact Registry
Deploy the assistant on the Google N4A node pool
Observe and benchmark the assistant on the Google N4A node pool
Move the assistant to the Google C4A node pool and compare results
Next Steps
Create the mixed-placement overlay
If you opened a new terminal, return to the source tree and restore the required variables:
source "${HOME}/.storefront-axion-env"
cd "${REPO}"
export FRONTEND_IP="$(kubectl get service frontend-external \
-o jsonpath='{.status.loadBalancer.ingress[0].ip}')"
export APP_URL="http://${FRONTEND_IP}"
Create a second overlay that keeps the storefront on N4A and moves only shoppingassistantservice to C4A:
mkdir -p kustomize/overlays/assistant-c4a
cat <<EOF > kustomize/overlays/assistant-c4a/kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ../../base
components:
- ../../components/shopping-assistant
images:
- name: shoppingassistantservice
newName: ${ASSISTANT_IMAGE_REPO}
newTag: ${ASSISTANT_IMAGE_TAG}
patches:
- path: keep-storefront-on-n4a.yaml
target:
kind: Deployment
- path: shoppingassistant-c4a.yaml
target:
kind: Deployment
name: shoppingassistantservice
EOF
cat <<EOF > kustomize/overlays/assistant-c4a/keep-storefront-on-n4a.yaml
- op: add
path: /spec/template/spec/nodeSelector
value:
cloud.google.com/gke-nodepool: ${N4A_NODE_POOL_NAME}
- op: add
path: /spec/template/spec/tolerations
value:
- key: kubernetes.io/arch
operator: Equal
value: arm64
effect: NoSchedule
EOF
cat <<EOF > kustomize/overlays/assistant-c4a/shoppingassistant-c4a.yaml
- op: replace
path: /spec/template/spec/nodeSelector
value:
cloud.google.com/gke-nodepool: ${C4A_NODE_POOL_NAME}
- op: replace
path: /spec/template/spec/tolerations
value:
- key: kubernetes.io/arch
operator: Equal
value: arm64
effect: NoSchedule
EOF
Render the mixed-placement overlay
Review and render the overlay before you apply it:
sed -n '1,140p' kustomize/overlays/assistant-c4a/kustomization.yaml
sed -n '1,140p' kustomize/overlays/assistant-c4a/keep-storefront-on-n4a.yaml
sed -n '1,140p' kustomize/overlays/assistant-c4a/shoppingassistant-c4a.yaml
kubectl kustomize kustomize/overlays/assistant-c4a | sed -n '1,240p'
Don’t continue if newName, newTag, or either node-pool value is blank.
The first patch keeps the storefront on N4A by default. The second patch overrides only shoppingassistantservice and moves that deployment to C4A.
Apply the mixed-placement overlay
Apply the overlay and wait for the rollout:
kubectl apply -k kustomize/overlays/assistant-c4a
kubectl rollout status deployment/frontend --timeout=600s
kubectl rollout status deployment/shoppingassistantservice --timeout=1200s
Confirm the final placement:
kubectl get pods -o wide | grep -E 'frontend|shoppingassistantservice'
kubectl get nodes -L cloud.google.com/gke-nodepool,node.kubernetes.io/instance-type,kubernetes.io/arch
The frontend remains on N4A, and shoppingassistantservice now runs on C4A.
Warm the assistant on C4A
Open the assistant endpoint to refresh the HTTP session cookie, then send one request through the storefront after the move. The cookie file preserves the same browser-like session across both calls:
curl --max-time 30 -s -c /tmp/assistant.cookies "${APP_URL}/assistant" >/dev/null
for i in {1..12}; do
if curl --max-time 120 -s \
-b /tmp/assistant.cookies \
-c /tmp/assistant.cookies \
-H 'Content-Type: application/json' \
-d '{"message":"Show me one desk item that looks minimal and practical."}' \
"${APP_URL}/bot" | tee /tmp/assistant-response.json | jq .; then
break
fi
sleep 10
done
The response contains a non-empty message.
Run the fixed C4A benchmark
Run the same fixed batch with the expected pool set to C4A:
python3 workshop/fixed_batch_trial.py \
--base-url "${APP_URL}" \
--expected-pool "${C4A_NODE_POOL_NAME}" \
--concurrency 2 \
--prompts-per-worker 8 \
--delay-seconds 0.2 \
--request-timeout-seconds 120 \
--telemetry-interval-seconds 0.5 \
--idle-threshold-cores 1.0 \
--idle-stable-samples 3 \
--idle-timeout-seconds 180 \
--request-output-file workshop/request-fixed-c4a.jsonl \
--telemetry-output-file workshop/telemetry-fixed-c4a.jsonl | \
tee workshop/fixed-c4a-summary.log
Print the C4A summary line:
grep 'SUMMARY_JSON' workshop/fixed-c4a-summary.log
The summary shows successful requests and the assistant running on the C4A node pool. Don’t compare the runs if the summary line is missing or reports a different node pool.
Compare the two runs
Compare the N4A and C4A summaries:
python3 workshop/compare_summaries.py \
workshop/fixed-n4a-summary.log \
workshop/fixed-c4a-summary.log | \
sed '/^Interpretation$/,$d'
This displays the measured comparison table without forcing a predetermined placement conclusion. Lower batch duration and request latency indicate the faster run for those measurements.
Focus on these signals:
- Successful request count
- Total benchmark duration
- Average request latency
- Maximum request latency
- Average assistant CPU cores during the measured batch
Your numbers can vary by prompt mix, model warmup state, node pressure, and cluster conditions. Treat the comparison as evidence for this assistant workflow, and repeat the same pattern with your own traffic before choosing a placement.
Assistant CPU values come from sampled Kubernetes telemetry. Short bursts can be underrepresented, especially on faster runs. Treat request success, benchmark duration, and latency as the primary comparison signals.
Understand the result
You changed one architectural variable: the placement of the assistant. The storefront tier, prompts, model, traffic shape, and benchmark script stayed the same.
The application now has three states:
- Storefront only on N4A
- Storefront plus assistant on N4A
- Storefront on N4A with only the assistant on C4A
That sequence is the core architectural point. A modern application can contain a steady application tier and a bursty AI reasoning tier. N4A supports the steady storefront and core services, while C4A gives you a placement to evaluate for concentrated, latency-sensitive assistant work.
Both Axion machine series run arm64, so the same assistant image moves between the Neoverse N3-based N4A pool and the Neoverse V2-based C4A pool. You can evaluate placement without maintaining a second image or build pipeline.
For this application, the useful pattern isn’t replacing one machine series with another. It’s placing each workload tier where its behavior fits best.
If your N4A and C4A results are close, or if N4A looks better for a particular run, that’s still useful information. The decision should come from the workload behavior you observe, not from assuming one Axion-based machine series is always the right answer.
What you’ve accomplished
You’ve built, observed, and compared a mixed-placement storefront application on Axion. You kept the storefront on N4A, moved only the assistant tier to C4A, and used the same benchmark to compare both placements.
You’re now ready to use the same overlay and benchmark pattern to evaluate which Axion-based machine series matches each tier’s workload behavior.