# Build a real-time offline voice chatbot using STT and vLLM

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/)
- [Build an offline voice assistant with whisper and vLLM](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/1_offline_voice_assistant/)
- [Install faster-whisper for local speech recognition](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/2_setup/)
- [Build a real-time STT pipeline on CPU](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/3_fasterwhisper/)
- [Fine-tune segmentation parameters](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/3a_segmentation/)
- [Build a real-time offline voice chatbot using STT and vLLM](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/4_vllm/)
- [Connect speech recognition to vLLM for real-time voice interaction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/4a_integration/)
- [Specialize offline voice assistants for customer service](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/5_chatbot_prompt/)
- [Enable context-aware dialogue with short-term memory](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/6_chatbot_contextaware/)
- [Next Steps](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/_next-steps/)

## Deploy vLLM for local language generation
In the previous section, you built a complete Speech-to-Text (STT) engine using faster-whisper, running efficiently on Arm-based CPUs. Now it’s time to add the next building block: a local large language model (LLM) that can generate intelligent responses from user input.

You’ll integrate [vLLM](https://vllm.ai/), a high-performance LLM inference engine that runs on GPU and supports advanced features such as continuous batching, OpenAI-compatible APIs, and quantized models.

### Why vLLM?
When building a real-time AI assistant, low latency and high throughput are critical. vLLM is designed for modern GPUs, including CUDA and TensorRT backends. It provides an OpenAI-compatible API, allowing you to use existing prompt structures and client code. It supports GPTQ, AWQ, FP16, and other formats for flexible deployment of open-source models, and enables serving large models such as LLaMA 2 and Mistral with efficient memory usage, even on constrained devices.

vLLM is especially effective in hybrid systems like the DGX Spark, where CPU cores handle STT and preprocessing, while the GPU focuses on fast, scalable text generation.

### Install and launch vLLM with GPU acceleration
In this section, you’ll install and launch vLLM - an optimized large language model (LLM) inference engine that runs efficiently on GPU. This component will complete your local speech-to-response pipeline by transforming transcribed text into intelligent replies.

#### Install Docker and pull vLLM image
The most efficient way to install vLLM on DGX Spark is using the NVIDIA official Docker image.

Before you pull the image, ensure [Docker](https://docs.nvidia.com/dgx/dgx-spark/nvidia-container-runtime-for-docker.html) is installed and functioning on DGX Spark. Then enable Docker GPU access and pull the latest NVIDIA vLLM container:

```
export LATEST_VLLM_VERSION=25.11-py3
docker pull nvcr.io/nvidia/vllm:${LATEST_VLLM_VERSION}
```

Confirm the image was downloaded:

```
docker images
```

The image is shown in the output:

```
__output__
nvcr.io/nvidia/vllm      25.11-py3                 d33d4cadbe0f   2 months ago   14.1GB
```

#### Download a quantized model (GPTQ)
Use Hugging Face CLI to download a pre-quantized LLM such as Mistral-7B-Instruct-GPTQ and Meta-Llama-3-70B-Instruct-GPTQ models for real-time AI conversations.

```
pip install huggingface_hub
hf auth login  # log in to Hugging Face with your token
```

After logging in successfully, download the specific models:

```
mkdir -p ~/models
# Mistral 7B GPTQ
hf download TheBloke/Mistral-7B-Instruct-v0.2-GPTQ --local-dir ~/models/mistral-7b
# (Optional) LLaMA3 70B GPTQ (requires more GPU memory)
hf download TechxGenus/Meta-Llama-3-70B-Instruct-GPTQ --local-dir ~/models/llama3-70b
```

Check model contents to ensure download success:

```
tree ~/models/mistral-7b -L 1
```

The files should include config.json, tokenizer.model, model.safetensors, etc.

```
__output__
├── config.json
├── generation_config.json
├── model.safetensors
├── quantize_config.json
├── README.md
├── special_tokens_map.json
├── tokenizer_config.json
├── tokenizer.json
└── tokenizer.model
__output__
1 directory, 9 files
```

#### Run the vLLM server with GPU
Mount your local ~/models directory and start the vLLM inference server with your downloaded model:

```
docker run -it --gpus all -p 8000:8000 \
  -v ~/models:/models \
  nvcr.io/nvidia/vllm:${LATEST_VLLM_VERSION} \
  vllm serve /models/mistral-7b \
  --quantization gptq \
  --gpu-memory-utilization 0.9 \
  --dtype float16
```

> **Note**: The first launch compiles and caches the model. To reduce startup time in future runs, consider creating a Docker snapshot with docker commit.

You can also check your NVIDIA driver and CUDA compatibility during the vLLM launch by looking at the output.

```
__output__
NVIDIA Release 25.11 (build 231063344)
vLLM Version 0.11.0+582e4e37
Container image Copyright (c) 2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
...
```

#### Verify the server is running
Once you see the message “Application startup complete.” vLLM is ready to run the model.

Send a test request with curl on other terminal:

```
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "/models/mistral-7b",
    "messages": [{"role": "user", "content": "Explain RISC architecture"}],
    "max_tokens": 256
  }'
```

If successful, the response includes a text reply from the model.

```
__output__
{"id":"chatcmpl-19aee139aabc474c93a3d211ee89d2c8","object":"chat.completion","created":1769183473,"model":"/models/mistral-7b","choices":[{"index":0,"message":{"role":"assistant","content":" RISC (Reduced Instruction Set Computing) is a computer architecture design where the processor has a simpler design and a smaller instruction set compared to CISC (Complex Instruction Set Computing) processors. RISC processors execute a larger number of simpler, more fundamental instructions. Here are some key features of RISC architecture:\n\n1. **Reduced Instruction Set:** RISC processors use a small set of basic instructions that can be combined in various ways to perform complex tasks. This is in contrast to CISC processors, which have a larger instruction set that includes instructions for performing complex tasks directly.\n2. **Register-based:** RISC processors often make extensive use of registers to store data instead of memory. They have a larger number of registers compared to CISC processors, and instructions typically operate directly on these registers. This reduces the number of memory accesses, resulting in faster execution.\n3. **Immediate addressing:** RISC instruction format includes immediate addressing, meaning some instruction operands are directly encoded within the instruction itself, like an add instruction with a constant value. This eliminates the need for additional memory fetch operations, which can save clock cycles.\n4.","refusal":null,"annotations":null,"audio":null,"function_call":null,"tool_calls":[],"reasoning_content":null},"logprobs":null,"finish_reason":"length","stop_reason":null,"token_ids":null}],"service_tier":null,"system_fingerprint":null,"usage":{"prompt_tokens":14,"total_tokens":270,"completion_tokens":256,"prompt_tokens_details":null},"prompt_logprobs":null,"prompt_token_ids":null,"kv_transfer_params":null}
```

## What you’ve accomplished and what’s next
You’ve successfully installed and launched vLLM on DGX Spark using Docker, downloaded a quantized LLM model (GPTQ format), and verified the server responds to API requests. Your GPU-accelerated language model is now ready to generate intelligent responses.

In the next section, you’ll connect this vLLM server to the STT pipeline you built earlier, creating a complete voice-to-response system.
