Skip to main content
Microsoft
separator
https://catalogartifact.azureedge.net/publicartifacts/cloudimg1647283583153.vllm-ubuntu-24-04-cd9dbfa2-684a-4399-a66f-936f646dc4e8/image5_logolarge.png

vLLM on Ubuntu 24.04 LTS

by cloudimg

Private self-hosted OpenAI-compatible LLM inference on Ubuntu 24.04 (vLLM): per-VM bearer key, TLS, GPU-ready

vLLM is a high-throughput, memory-efficient inference and serving engine for large language models. Its PagedAttention scheduler delivers leading serving throughput, and it exposes an OpenAI-compatible REST API so existing OpenAI SDK code works unchanged. This cloudimg image installs the pinned vLLM release into a dedicated Python virtual environment on Ubuntu 24.04 LTS and runs it under systemd, with a small open-weights starter model pre downloaded so the API serves a model out of the box on a GPU VM.

Security is built in and there is no anonymous access and no default login. vLLM ships open, so the server binds to loopback only and nginx is the sole network facing surface. On first boot each VM generates its own random bearer API key and regenerates a fresh self signed TLS certificate whose subject alternative names include the VM public IP and hostname. Anonymous requests to the inference API return HTTP 401 with a bearer challenge; only the per VM key is accepted, and it is stored in a root only file. An open health endpoint remains available for load balancer probes.

The image is GPU-ready. Launch it on an NVIDIA Ampere or newer NC or ND series GPU VM and add the Azure NVIDIA GPU Driver Extension, and vLLM serves the model on the GPU. nginx terminates TLS on port 443 and proxies to vLLM with streaming friendly settings; port 80 redirects to HTTPS with an unauthenticated health endpoint. Call the OpenAI-compatible endpoints from LangChain, LlamaIndex or any OpenAI SDK.

vLLM is Apache-2.0 licensed, free and open source with no per CPU or per deployment fee. cloudimg provides packaging, credential and TLS automation, security patching, and 24/7 support with a guaranteed 24 hour response SLA.

At a glance

https://catalogartifact.azureedge.net/publicartifacts/cloudimg1647283583153.vllm-ubuntu-24-04-cd9dbfa2-684a-4399-a66f-936f646dc4e8/image4_screenshot01.png
https://catalogartifact.azureedge.net/publicartifacts/cloudimg1647283583153.vllm-ubuntu-24-04-cd9dbfa2-684a-4399-a66f-936f646dc4e8/image7_screenshot02.png
https://catalogartifact.azureedge.net/publicartifacts/cloudimg1647283583153.vllm-ubuntu-24-04-cd9dbfa2-684a-4399-a66f-936f646dc4e8/image0_screenshot03.png
https://catalogartifact.azureedge.net/publicartifacts/cloudimg1647283583153.vllm-ubuntu-24-04-cd9dbfa2-684a-4399-a66f-936f646dc4e8/image1_screenshot04.png
English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Us Privacy & Cookies Terms of Use About our ads Manage cookies