Why use MLIA

Use Arm ML Inference Advisor (MLIA) to evaluate whether a machine learning model is suitable for a target inference platform.

MLIA is most useful before full deployment or runtime profiling, when you’re asking questions such as:

  • Will this model map cleanly to my target?
  • Which operators or layers are likely to matter most for performance?
  • Is the model compute-bound, memory-bound, or affected by low multiply-accumulate (MAC) utilization?
  • Is the model well-optimized for my Arm target hardware?
  • What should I investigate before building firmware or running on a board?

MLIA doesn’t make the final optimization decision for you. It gives target-aware evidence so that you can decide what to change, what to measure next, and which workflow stage deserves attention.

You’ll use the CLI to complete the following tasks:

  • Discover installed targets, target profiles, and backends
  • Run compatibility checks
  • Run performance analysis
  • Inspect advice and metrics

Arm Ethos-U is the example target in the Learning Path.

If you want to automate the same checks, you can optionally use the Python API . The API is useful when you want to embed MLIA results in another product, dashboard, CI job, or tool.

How you should use MLIA

MLIA isn’t a replacement for graph visualization or runtime profiling. It’s an advisory layer that helps earlier in the model preparation workflow. The following table shows which questions MLIA and related tools can answer:

Tool or backendUse it to answer
MLIAIs this model suitable for my target, and what should I change?
VelaWhich operators are supported, and which layers dominate compiler-estimated cycles?
Corstone FVPWhat NPU performance counters does a packaged .pte artifact produce for the whole model run on a virtual platform?
Model ExplorerWhat does the generated model artifact graph look like?
Runtime-specific profiling toolsWhat happened when the model ran? For example, use ETRecord, ETDump, and ExecuTorch Inspector for ExecuTorch deployments. Use LiteRT benchmark and profiling tools for LiteRT deployments.

Vela-backed MLIA checks use compiler estimates. The checks can include operator-level breakdowns, such as which layers dominate estimated cycles or have low MAC utilization.

Corstone-backed MLIA checks run a packaged .pte file on an FVP and report NPU performance counters for the whole model run. The checks don’t provide per-layer estimates or operator breakdowns.

Model Explorer can show how an ExecuTorch .pte artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment.

Use the following tools together:

  • Use MLIA before or during model preparation.
  • Use Model Explorer to inspect generated artifacts and delegation structure.
  • Use runtime profiling tools after you can execute the model.

Supported model formats

MLIA can analyze different kinds of model artifacts depending on what workflow you are using and the stage that you want to analyze:

FormatWhere it fits
.pteSerialized ExecuTorch program. Ethos-U .pte performance analysis uses Corstone backends.
.tfliteLiteRT model format used in many Ethos-U and embedded ML workflows.
.tosaIntermediate representation consumed by compiler or backend flows such as Ethos-U Vela.

What you’ve learned and what’s next

You’ve learned what MLIA does, what you use it for, and how MLIA fits alongside other tools.

Next, you’ll install MLIA and inspect the capabilities available in your environment.

Back
Next