Prepare a pet image

Keep your Python environment active and examples/arm/arm-scratch/setup_path.sh sourced. Use the downloaded helper to prepare the first image in the dataset’s test split:

    

        
        
python arm_test/deit_vgf/deit_vgf_helper.py prepare

    

The output is similar to:

    

        
        Input image: arm_test/deit_vgf/input.jpg
Input tensor ready: (1, 3, 224, 224)
Expected breed: newfoundland

        
    

Expected breed is the dataset’s reference label. It’s not a model prediction because you haven’t run inference yet.

The helper uses the same pinned image processor as the exporter. It saves input.bin, input.jpg, and reference.json under arm_test/deit_vgf/. The reference file stores the breed-label mapping. It also stores the selected image’s reference label for inspection.

input.bin contains the normalized float32 tensor in [1, 3, 224, 224] batch, channel, height, width order. The runner reads this tensor, not the JPEG. Open input.jpg in an image viewer to inspect the pet that you’ll classify.

Run inference through VGF

Inference produces one score for each of the 37 breeds. The inspection step then maps the highest score to a breed name. Execute the exported program with your input file and save the scores:

    

        
        
set -o pipefail
./cmake-out-deit-vgf/executor_runner \
  --model_path=arm_test/deit_vgf/deit_quantized_vgf.pte \
  --inputs=arm_test/deit_vgf/input.bin \
  --output_file=arm_test/deit_vgf/prediction \
  2>&1 | tee arm_test/deit_vgf/runtime.log

    

A successful run reports Model executed successfully and writes arm_test/deit_vgf/prediction-0.bin. The runner appends -0.bin for the first output tensor.

Keep --inputs. Without the flag, the generic runner fills the input tensor with ones instead of classifying your pet image.

Inspect the breed prediction

Decode the saved scores and verify VGF execution:

    

        
        
python arm_test/deit_vgf/deit_vgf_helper.py inspect

    

The output is similar to:

    

        
        Expected breed: newfoundland
VGF prediction: newfoundland
Matches dataset label: True
Output scores: 37 finite values
VGF execution: confirmed

        
    

Output scores: 37 finite values means there’s one valid numeric score per breed, with no NaN or infinite values. The helper maps the largest score to the breed shown as VGF prediction. In this example, Matches dataset label: True confirms that the predicted Newfoundland breed matches the image’s reference label.

The helper also checks the runtime log for Entered VGF init and Model executed successfully before reporting VGF execution: confirmed.

A valid prediction and confirmed VGF execution complete the deployment workflow. A matching dataset label means the model recognizes this image. A mismatch doesn’t, by itself, indicate a deployment failure.

Note

One image doesn’t measure dataset accuracy. The export log reports quantized PyTorch accuracy, not VGF accuracy across the test set. This host emulation run also doesn’t establish performance on an Arm GPU.

(Optional) Compare with the floating-point model

Run the original fine-tuned model on the same input and compare its winning class with the VGF result:

    

        
        
python arm_test/deit_vgf/deit_vgf_helper.py inspect --compare-fp32

    

The output is similar to:

    

        
        Expected breed: newfoundland
VGF prediction: newfoundland
Matches dataset label: True
Output scores: 37 finite values
VGF execution: confirmed
FP32 prediction: newfoundland
Matches FP32 prediction: True

        
    

The helper adds the floating-point prediction and whether the two predictions match. Quantization can change the winning class. Compare more images before drawing conclusions about accuracy or numerical equivalence.

What you’ve accomplished

You’ve fine-tuned DeiT-Tiny, exported a VGF-backed program, and classified a pet image with the host runtime. Your model, input, prediction, and logs are in arm_test/deit_vgf/.

You can now reuse the .pte to classify other test images. To classify another test image, repeat prepare with --sample-index 1, then rerun inference and inspection. Each preparation replaces the previous input artifacts. The helper rejects predictions and logs that predate the prepared image. Rerun executor_runner before inspecting a new image.

Back
Next