Build an ephemeral AI code sandbox with Firecracker on Arm
Introduction
Understand the sandbox architecture
Prepare the Arm KVM host
Run code in a disposable microVM
Validate the execution boundaries
Next Steps
Build an ephemeral AI code sandbox with Firecracker on Arm
Test sandbox controls
Earlier, you built the sandbox and ran a program inside a disposable microVM. Now, you’ll verify that the boundaries hold. Continuing in the ~/firecracker-ai-sandbox directory, you’ll run three independent checks, one for each control that the sandbox is meant to enforce:
- Filesystem disposability: confirm that files created by one job don’t appear in the next
- Execution timeout: confirm that the host stops a job that runs longer than its limit
- Network and runtime cleanup: confirm that the per-job TAP device and runtime directory are removed after the microVM exits
Each check runs an example program and inspects the result.
Verify filesystem disposability
Verify that a file created by one job is absent from the next job’s filesystem. This confirms that each execution starts from a fresh copy of the base image without retaining files from the previous run.
Run the marker program twice:
sudo ./sandbox/run-job.sh ./sandbox/examples/write-marker.sh
sudo ./sandbox/run-job.sh ./sandbox/examples/write-marker.sh
Each execution checks for /tmp/ai-sandbox-marker before creating it. Both jobs print:
Created /tmp/ai-sandbox-marker inside this microVM.
Run this example again: the marker will be absent because the disk is disposable.
The second job starts from another copy of the clean base image, so it can’t see the marker created by the first job.
You can also run the combined demonstration:
sudo ./sandbox/demo.sh
The demonstration runs the architecture inspection program followed by two marker jobs.
Verify the execution timeout
The timeout example sleeps longer than the configured job limit. Run the example with a three-second limit and capture the expected nonzero status:
set +e
sudo FC_JOB_TIMEOUT=3 ./sandbox/run-job.sh ./sandbox/examples/timeout.sh
STATUS=$?
set -e
echo "runner exit code=$STATUS"
The result includes:
outcome=timed_out
exit_code=124
runner exit code=124
The timeout runs on the host and terminates the SSH client if execution stops responding. After recording and printing the result, the cleanup trap terminates Firecracker and removes the job disk. Boot and file transfer happen before this timeout starts.
For this sleep example, exit code 124 demonstrates the timeout. The runner also labels exit code 137 as timed_out. A guest program can return either code itself, so this result label alone doesn’t prove that an arbitrary job exceeded its deadline.
Confirm network and runtime cleanup
Check that no sandbox TAP device remains:
if ip link show fc-ai0 >/dev/null 2>&1; then
echo "TAP cleanup failed"
else
echo "TAP cleanup succeeded"
fi
The expected output is:
TAP cleanup succeeded
Check that the runtime directory contains no job directories:
sudo find /opt/firecracker-ai/runtime -mindepth 1 -maxdepth 1 -type d -print
The command produces no output after cleanup. The runner.lock file remains and is reused to prevent concurrent jobs. These checks verify TAP and runtime-directory removal.
What you’ve accomplished
You’ve built an Arm-native execution sandbox where each shell program receives a dedicated Firecracker microVM, bounded resources, restricted networking, and a disposable filesystem. You also validated that filesystem state doesn’t cross job boundaries and that the host stops programs that exceed their time limit.
You can extend these examples into a service. Before extending the example, review the Firecracker production host setup recommendations and assess your deployment’s security requirements. Running untrusted code in production requires host hardening, resource controls, and network policies suited to your workload.