Run vLLM inference with INT4 quantization on Arm servers
Introduction
Build and validate vLLM for inference
Quantize an LLM to INT4
Serve high throughput inference with vLLM
Evaluate accuracy with LM Evaluation Harness
Next Steps
Run vLLM inference with INT4 quantization on Arm servers
Share
Bring your insights to the conversation.
Give Feedback
How would you rate this Learning Path?
What is the primary reason for your feedback ?
Thank you! We're grateful for your feedback.
- Have more feedback? Log an issue on GitHub.
- Want to collaborate? Join our Discord server.
Continue Learning
Read related resources
Find more information about the topics in this Learning Path:
Join the Arm Developer Program
Connect, upskill, and build with the Arm Developer Community. Join today for hands-on technical resources and education materials, along with the support of Arm engineers and the broader ecosystem.