# Build an offline voice chatbot with faster-whisper and vLLM on DGX Spark

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/)
- [Build an offline voice assistant with whisper and vLLM](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/1_offline_voice_assistant/)
- [Install faster-whisper for local speech recognition](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/2_setup/)
- [Build a real-time STT pipeline on CPU](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/3_fasterwhisper/)
- [Fine-tune segmentation parameters](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/3a_segmentation/)
- [Build a real-time offline voice chatbot using STT and vLLM](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/4_vllm/)
- [Connect speech recognition to vLLM for real-time voice interaction](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/4a_integration/)
- [Specialize offline voice assistants for customer service](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/5_chatbot_prompt/)
- [Enable context-aware dialogue with short-term memory](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/6_chatbot_contextaware/)
- [Next Steps](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_voicechatbot/_next-steps/)

## About this Learning Path

| Skill level: | Advanced |
|--------------|----------|
| Reading time: | 1 hr |
| Last updated: | 29 Jul 2026 |

| Author: | Odin Shen, Arm [GitHub](https://github.com/odincodeshen) [LinkedIn](https://linkedin.com/in/odin-shen-lmshen) |
|----------|--------------------------------------------------|
| Arm IP: | [Neoverse](https://support.arm.com/?tab=compute-ip&Product%20Type=Infrastructure%20Processors) |
| Tags: | [Performance and Architecture](https://learn.arm.com/tag/performance-and-architecture), [Linux](https://learn.arm.com/tag/linux), [Docker](https://learn.arm.com/tag/docker), [Python](https://learn.arm.com/tag/python) |

### Who is this for?
This is an advanced topic for developers and ML engineers who want to build private, offline voice assistant systems on Arm-based servers such as DGX Spark.

### What will you learn?
Upon completion of this Learning Path, you will be able to:
- Explain the architecture of an offline voice chatbot pipeline combining speech-to-text (STT) and vLLM
- Capture and segment real-time audio using PyAudio and Voice Activity Detection (VAD)
- Transcribe speech using faster-whisper and generate replies using vLLM
- Tune segmentation and prompt strategies to improve latency and response quality
- Deploy and run the full pipeline on Arm-based systems such as DGX Spark

### Prerequisites
Before starting, you will need the following:
- An NVIDIA DGX Spark system with at least 15 GB of available disk space
- A USB microphone for audio input

### Summary
You’ll build a local voice chatbot on an Arm-based NVIDIA DGX Spark with `faster-whisper` for speech-to-text and vLLM for response generation. You’ll capture microphone audio with PyAudio, add voice and turn detection, tune segmentation for stable low-latency transcription, and connect the CPU STT pipeline to GPU-backed vLLM. You’ll validate segmented transcripts followed by local replies.

### Frequently asked questions

<details>
<summary>What result should I expect after installing `faster-whisper`?</summary>
After a successful installation, you can transcribe a short audio sample or live microphone input with readable text and no runtime errors. Use this to confirm the installation before moving on to pipeline changes.
</details>

<details>
<summary>When should I upgrade the speech model in the CPU STT pipeline?</summary>
Upgrade after you confirm baseline transcription works. The build step adds a more accurate model and VAD. If latency increases, proceed to segmentation tuning.
</details>

<details>
<summary>How do I know VAD and turn detection are working correctly?</summary>
Transcriptions should arrive as sentence-like chunks, and pauses should start new segments. If long monologues merge into one block or speech is cut mid-sentence, adjust the segmentation parameters.
</details>

<details>
<summary>What should I verify before integrating vLLM with the STT engine?</summary>
Ensure the CPU-based STT runs in real time on your DGX Spark and produces stable, segmented text. A clean, timely text stream simplifies downstream integration with vLLM.
</details>

<details>
<summary>What behavior confirms the end-to-end offline chatbot is running?</summary>
Speak into the microphone and watch for segmented transcriptions from `faster-whisper`, followed by a locally generated reply from vLLM. Seeing this sequence consistently indicates the pipeline is integrated and running on the system.
</details>
