Local LLM server: Ollama on Ubuntu 24.04, hardened, run open models on CPU or GPU sizes.
Ollama on Ubuntu 24.04 LTS
A production-ready Ollama 0.35 local large language model (LLM) runtime on Ubuntu 24.04 LTS, hardened for Azure. Ollama lets you pull and run open-weight models - Llama, Mistral, Gemma, Phi, Qwen and more - behind a simple REST API, with your prompts and data never leaving your subscription. Installing it by hand means runtime setup, service wiring, and storage planning; a manual, time-consuming setup. With this image the Ollama API answers within minutes of deployment: pull a model and run your first inference immediately.
Who this is for: Developers and data teams who want private LLM inference - prototyping, retrieval-augmented generation (RAG), or batch processing - on Azure VMs they control, without per-token API fees.
What's pre-configured in this image:
- Ollama 0.35.0 installed from the official upstream release, SHA-256 verified at build time
- systemd service enabled, serving the Ollama REST API on port 11434
- Dedicated model storage directory owned by an unprivileged service user
- Runs on CPU out of the box; attach an Azure GPU SKU and install NVIDIA drivers to accelerate inference
What happens at first boot (once, on your VM):
- Starts the Ollama service and health-checks the API before marking first boot complete
- Writes a customer README on the VM with quick-start commands (ollama pull, ollama run)
- No credentials are issued - the Ollama API has no built-in authentication and serves on port 11434; you choose who can reach it with your network security group or an authenticating reverse proxy
Use cases:
- Private chat and completion API - keep prompts and outputs inside your Azure tenant
- Retrieval-augmented generation (RAG) - pair Ollama with your vector database for grounded answers
- Model evaluation - compare open-weight models side by side before committing
- Batch summarization and extraction - process documents overnight on Azure Spot VMs
- Air-gapped inference - regulated workloads with no external AI dependency
Azure integration: Ships as a Generation 2 Azure virtual machine image with Trusted Launch support (Secure Boot and virtual Trusted Platform Module) so you deploy on a verified chain of trust. Azure Monitor Agent, Microsoft Defender for Cloud, and Azure Update Manager install cleanly on first boot, and the Azure Linux Agent plus cloud-init are pre-validated, so your existing Azure automation - Custom Script Extension, Run Command, VM Applications - works without modification. Keep traffic private with Azure Private Link or a network security group scoped to your own virtual network.
Security posture: The image is built from a fully patched Ubuntu 24.04 LTS base at build time, and the application is installed from its official upstream release channel with cryptographic verification during the build. No default or shared credentials are ever baked into the image: anything secret is generated per-VM on first boot and stored root-only on your VM, so no two deployments share a secret and there is no claimable open instance waiting on the network.
System requirements: Minimum Standard_D4as_v5 (4 vCPU, 16 GB RAM) for 7B-parameter models on CPU. Recommended: Standard_D8as_v5 (8 vCPU, 32 GB RAM) for 13B models, or an NC-series GPU SKU for low-latency inference. Attach a Premium SSD data disk for model storage - models range from 2 GB to 40+ GB.
Deploy Ollama on Azure today and run your first local model inference in under 15 minutes. Get started by clicking the button above.