Run optimized LLMs from the Arm AI Portal on Arm Neoverse-based instances
Introduction
Prepare the Arm Neoverse Linux environment
Download and run a model from the Arm AI Portal
Understand how the runner scripts work
(Optional) Explore other cloud LLM deployment options
Next Steps
Run optimized LLMs from the Arm AI Portal on Arm Neoverse-based instances
Who is this for?
This Learning Path is for developers and ML engineers running Arm-optimized Large Language Models (LLMs) on Arm Neoverse-based Linux machines. It provides a tested ONNX Runtime GenAI workflow and an optional AI coding agent workflow for compatible text-to-text packages that use alternative runtimes or formats.
What will you learn?
Upon completion of this Learning Path, you will be able to:
- Prepare an Arm Neoverse Linux machine and download a model from the Arm AI Portal.
- Generate text from the terminal and optionally serve the model through a local web application.
- Identify how the shared application calls the supplied ONNX Runtime GenAI adapter.
- Compare runtime and model-format choices, and optionally use a coding agent to replace the supplied adapter for another compatible text-to-text package.
Prerequisites
Before starting, you will need the following:
- An Arm Neoverse-based Linux machine running Ubuntu 24.04 LTS, with Python 3.11 or later, for example an AWS
m8g.xlargeinstance - At least 16 GB of memory if you plan to use one of the 8B models
- Basic familiarity with Linux command-line tools and Python
- (Optional) Access to a coding agent if you want to generate an adapter for another runtime
Summary
This summary was drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
Frequently asked questions
These FAQs were drafted with an approved AI-assisted workflow and reviewed by Arm contributors before publication. Human technical review remains part of the process so the final page reflects engineering rigor, accuracy, and Arm editorial standards.
run_model.py for terminal generation and genai_web.py for an optional local browser interface. Both entry points call the shared adapter, which uses the supplied ONNX Runtime GenAI implementation.Arm/<model-repository-name>. When you use the supplied adapter, choose an ID from the confirmed models table. If you generated an adapter for another runtime, keep your existing MODEL_ID and follow that package’s model-type and prompt guidance.validate_adapter.py to check the adapter contract and required package files. If the check fails, align your implementation with the interface defined in adapter_contract.py.run_model.py reports time to first token and decode throughput.adapter_contract.py for the target package, replacing the supplied ONNX Runtime GenAI implementation in model_adapter.py. You can retain the terminal and web runners and follow the alternative package’s guidance for model type and prompting.