# Deploy a LLM-based vision chatbot with PyTorch and Hugging Face transformers on Google Axion processors

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/)
- [Set up an LLM based-Vision Chatbot](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/vision_chatbot/)
- [Deploy Vision Chatbot LLM backend server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/backend/)
- [Deploy Vision Chatbot LLM frontend server](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/frontend/)
- [Inference with Vision Chatbot](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/conclusion/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama-vision/_next-steps/)

## About this Learning Path

| Skill level:          | Advanced            |
|-----------------------|---------------------|
| Reading time:         | 45 min              |
| Last updated:         | 17 Sep 2026         |

| Author:                             | Nobel Chowdary Mandepudi, Arm                       |
|-------------------------------------|----------------------------------------------------|
| Arm IP:                             | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags:                               | ML, Google Axion, Linux, Python, PyTorch, Streamlit |

### Who is this for?
This Learning Path is for software developers and ML engineers who are interested in deploying a production-ready vision chatbot for their application with optimized performance on the Arm Architecture.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Build a frontend with Streamlit to input images and prompts.
- Build the backend to download a Llama 3.2-Vision model, quantize it, and run it using PyTorch and Hugging Face Transformers.
- Monitor and analyze inference on Arm CPUs.

### Prerequisites
Before starting, you will need the following:
- A Google Cloud Axion compute instance or [any Arm-based instance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/csp/) from a cloud service provider with at least 32 cores
- Familiarity with REST APIs and web services
- A basic understanding of Python and ML concepts
- A basic understanding of Streamlit
- A basic understanding of LLM fundamentals

### Summary
You’ll build and deploy a vision-enabled chatbot with PyTorch, Transformers, and Streamlit on an Arm-based instance powered by Google Axion. First, you’ll run a Flask backend that downloads and serves a quantized Llama 3.2-Vision model, then create a Streamlit frontend for image uploads and prompts. You’ll configure firewall access, start both services on Ubuntu, and verify image-plus-text responses in the web app.

### Frequently asked questions
#### What result should I expect when both the backend and frontend are running?
Open the browser to the app. You’ll see the title **LLM Vision Chatbot on Arm** with controls to upload an image and enter a prompt. After submitting, the page will display a generated text response that uses the image as context.

#### Which address should I use to open the web app?
Use `http://[your instance ip]:8501` in your browser. If the page doesn’t load, allow inbound TCP traffic to port `8501` in your instance’s security rules.

#### How do I run the backend and frontend at the same time?
Start the backend script in one terminal with the virtual environment activated. Then, open a new terminal, activate the same environment, and start the Streamlit frontend.

#### What should I check if the frontend can't reach the backend?
Confirm that `backend.py` is running on port `5000` and that `frontend.py` uses `http://localhost:5000/v1/chat/completions`. For remote browser access, open the Streamlit frontend on port `8501`. The frontend connects to the backend locally.

#### How do I know that the model download and 4-bit quantization completed?
Start `backend.py` and wait for the Flask output showing that the server is running on port `5000`. The `backend.py` startup loads the Llama 3.2 Vision model and performs quantization before serving requests.
