Skip to content

AI

How to self-host an LLM with Ollama, and when it makes sense

Run an open-weight language model on your own machine or server with Ollama, call it from code, size the hardware, and decide whether self-hosting beats an API.

By · Published · 3 min read

Short answer: install Ollama, pull an open-weight model that fits your memory, and call its local HTTP API from your code. It takes about ten minutes. Whether it is worth it depends on your privacy needs, your volume and your hardware. For many small projects a hosted API is cheaper and better. For private data or steady high volume, self-hosting can win.

How do you run a model locally?

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# download and chat with a model
ollama pull llama3.2
ollama run llama3.2

Ollama starts a server on localhost:11434. Model names change as new ones release, so check the library page for what is current.

How do you call it from code?

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "stream": false,
  "messages": [{ "role": "user", "content": "Summarise: invoice 118 is 30 days overdue." }]
}'

Ollama also offers an OpenAI-compatible endpoint at /v1, so many libraries work by changing the base URL.

How much hardware do you need?

Memory is the constraint. A model's weights must fit in GPU memory (VRAM), or in system RAM if you accept slower speed. As a rough guide, a model quantised to 4 bits needs about half a gigabyte per billion parameters, plus overhead for the context.

  • Around 3B parameters: about 2 to 3 GB. Runs on most laptops.
  • Around 8B parameters: about 5 to 6 GB. A recent consumer GPU or a Mac with 16 GB.
  • Around 30B parameters: about 18 to 20 GB. A 24 GB GPU or a Mac with 32 GB or more.
  • 70B and above: 40 GB or more. Multiple GPUs or a large unified-memory machine.

Those numbers are estimates. Context length adds memory, and different quantisation levels change quality and size. Start small and move up only if the output is not good enough.

When does self-hosting make sense?

  • Privacy. Data must not leave your network, such as patient records or legal documents.
  • Steady volume. You process a large, predictable amount of text every day, so a fixed machine beats per-token pricing.
  • Offline use. The site has unreliable internet.
  • Experiments. You want to try many models cheaply or fine-tune.

When does it not?

  • You need the best available quality. The top hosted models are generally stronger than the open models you can run on one machine.
  • Your usage is low or bursty. A GPU you pay for sits idle.
  • You have no one to operate it. Updates, monitoring and capacity are your job now.

How do you expose it safely?

Ollama has no authentication. Keep it bound to localhost or a private network. If other machines need it, put a reverse proxy in front with authentication and TLS, and do not publish port 11434 to the internet. People have found open instances by scanning. The same hygiene as the self-hosting checklist applies.

How do you compare against an API fairly?

Take fifty real inputs, run them through the local model and the hosted one, and judge the outputs blind. Then calculate cost: hardware purchase or rental, electricity, your time, versus API spend from your cost logs. Decide with numbers from your own workload.

References

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello