Who is this for?

This Learning Path is for developers and ML engineers running Arm-optimized Large Language Models (LLMs) on Arm Neoverse-based Linux machines. It provides a tested ONNX Runtime GenAI workflow and an optional AI coding agent workflow for compatible text-to-text packages that use alternative runtimes or formats.

What will you learn?

Upon completion of this Learning Path, you will be able to:

  • Prepare an Arm Neoverse Linux machine and download a model from the Arm AI Portal.
  • Generate text from the terminal and optionally serve the model through a local web application.
  • Identify how the shared application calls the supplied ONNX Runtime GenAI adapter.
  • Compare runtime and model-format choices, and optionally use a coding agent to replace the supplied adapter for another compatible text-to-text package.

Prerequisites

Before starting, you will need the following:

  • An Arm Neoverse-based Linux machine running Ubuntu 24.04 LTS, with Python 3.11 or later, for example an AWS m8g.xlarge instance
  • At least 16 GB of memory if you plan to use one of the 8B models
  • Basic familiarity with Linux command-line tools and Python
  • (Optional) Access to a coding agent if you want to generate an adapter for another runtime

Summary

AI-assisted

This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
You’ll run optimized LLMs from the Arm AI Portal on an Arm Neoverse-based Linux machine. First, you’ll create a Python environment, download the shared application files, and either use the supplied ONNX Runtime GenAI adapter or generate a compatible replacement. You’ll select and run a model and optionally start its web interface. Then, you’ll examine how the runner scripts call the adapter. You’ll optionally learn to compare alternative runtimes and formats for replacement if the supplied adapter doesn’t suit your needs.

Frequently asked questions

AI-assisted

These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.

Close
?
Which script should I run to generate text in the terminal or browser?
Use run_model.py for terminal generation and genai_web.py for an optional local browser interface. Both entry points call the shared adapter, which uses the supplied ONNX Runtime GenAI implementation.
How do I choose and set the correct ID for a model from the Arm AI Portal?
Select your model in the Arm AI Portal and copy its Hugging Face repository ID in the form Arm/<model-repository-name>. When you use the supplied adapter, choose an ID from the confirmed models table. If you generated an adapter for another runtime, keep your existing MODEL_ID and follow that package’s model-type and prompt guidance.
How do I verify my adapter before I run a model?
Use validate_adapter.py to check the adapter contract and required package files. If the check fails, align your implementation with the interface defined in adapter_contract.py.
What do I see after a successful terminal run?
You’ll see generated text in the terminal. run_model.py reports time to first token and decode throughput.
Which components change if I switch to another runtime or model format?
Implement the adapter defined in adapter_contract.py for the target package, replacing the supplied ONNX Runtime GenAI implementation in model_adapter.py. You can retain the terminal and web runners and follow the alternative package’s guidance for model type and prompting.
Next