Infrastructure for Independence 7 min read

Picking a Local AI Runtime Without Picking a Fight

A practical comparison of Ollama, LM Studio, and LocalAI: which local AI runtime fits which job, what they share under the hood, and how to choose without picking a fight.

Picking a Local AI Runtime Without Picking a Fight
Gemini_Generated_Image_ogkbozogkbozogkb.jpg

In the first post of this series I made the case for inspectable AI: models you can run, audit, and reason about on hardware you control. The obvious next question is also the most common one I get asked: what do I actually install?

The answer depends less on which tool is "best" and more on the specific role you're hiring it to fill. Three runtimes dominate the local/self-hosted conversation right now: Ollama, LM Studio, and LocalAI. They can load many of the same models. They all speak roughly the same OpenAI-shaped API. But they are built for different people—and different degrees of autonomy.

This is a map, not a verdict.

Ollama: the fastest path from zero to curl

Ollama is a CLI-first runtime and daemon. Install it, run ollama pull qwen2.5:14b, and you have a model serving an OpenAI-compatible API at localhost:11434/v1. That’s the value proposition in a nutshell, and it’s why Ollama reports 9M+ monthly installs and over a billion model downloads.

It is the obvious choice if you are:

  • Writing scripts or small automations that need a local LLM.
  • Plugging a chat frontend, IDE extension, or Raycast workflow into a local model.
  • Building on a laptop or small server and want something that "just runs."

The ecosystem is Ollama's real advantage. Modelfiles let you version-control a model configuration the same way a Dockerfile version-controls an environment:

FROM qwen2.5:14b
PARAMETER temperature 0.3
PARAMETER num_ctx 16384
SYSTEM "You are a precise technical editor."

That reproducibility matters when you move from "I got it working once" to "I need this to keep working."

The trade-offs are subtle but real. Ollama defaults to a 2,048-token context window unless you override it, and it unloads models after five minutes of inactivity to keep RAM free. Both defaults are sensible for casual use, but they can surprise you in automation where the first call after idle pays a reload cost. Quantization choices are also more opaque than in a GUI tool.

Use Ollama if: you want a local model as a service and you want it today.

→ Get started at ollama.com

LM Studio: the best evaluation bench

LM Studio is a desktop GUI built around model discovery, side-by-side chat, and a local server toggle. Under the hood it uses llama.cpp, and on Apple Silicon it can also use MLX.

Where it shines is evaluation. You can search Hugging Face from inside the app, see per-quantization RAM-fit indicators, download a model, and immediately run it against the same prompt in two different windows. If you are trying to answer "does this 32B model actually run well on my 16 GB card?" LM Studio gives you the answer faster than any other tool.

On Apple Silicon, the MLX engine is meaningfully faster than the same GGUF through llama.cpp — roughly 20–30% in generation, and more in prompt processing. The catch is that the win comes from MLX, not from LM Studio being magic. Run GGUF in LM Studio and the numbers look very similar to Ollama, because the engine underneath is the same.

The server mode is useful but secondary. It defaults to localhost:1234/v1 and works with the same OpenAI-compatible clients as Ollama, but it only runs while the app is open. LM Studio is not pretending to be infrastructure; it is a cockpit.

Use LM Studio if: you are evaluating models, comparing quantizations, or you want the easiest GUI for experimentation.

→ Download at lmstudio.ai

LocalAI: the swappable production runtime

LocalAI is the odd one out in this comparison because it is not trying to be a desktop app or a simple daemon. It is a server-oriented inference layer with a swappable backend engine. One OpenAI-compatible API can route models through llama.cpp, vLLM, vllm.cpp, transformers, MLX, and many others — including speech, vision, image generation, embeddings, and newer model types like decision models.

That breadth is exactly why I run it for my own stack. CloudHerder's tools — including the agent that helps produce this site — talk to LocalAI. When I want to test whether a new model even runs on 16 GB, I use LM Studio. When I need that model to become a reliable endpoint for a production-ish agent, I migrate it to LocalAI.

The trade-off is configuration. LocalAI has more knobs because it exposes more of the engine. You choose backends, quantizations, context lengths, and gallery sources explicitly. The payoff is flexibility: you can switch from a small CPU-friendly model to a large GPU model without changing the client code, and you can add audio, vision, or agent endpoints behind the same API.

As of the latest release, LocalAI supports more than 30 backends across text, speech, vision, image and video generation, audio processing, and utilities. That number is impressive, but it is also a warning: breadth does not mean every backend is the right choice for every workload.

Use LocalAI if: you are building pipelines or agents around a stable local API, and you need one runtime to handle multiple model types.

→ Learn more at localai.io

The hard truths

All three tools sit on top of the same engines. Ollama and LM Studio both use llama.cpp. LocalAI can use llama.cpp, vLLM, transformers, and MLX. The differentiation is packaging and audience, not some proprietary magic the others lack.

Ollama and LM Studio also converge more than their fanbases admit. Both serve OpenAI-compatible APIs. Both can run the same GGUF files. A developer can point the same client at localhost:11434 or localhost:1234 with one line changed.

"Local" does not mean "no setup." It means the setup is under your control. You still have to think about context length, quantization, VRAM budgets, and whether your GPU drivers are cooperating.

Which one should you use?

If your constraint is... Likely fit
"I want to chat with a model in under 5 minutes" LM Studio
"I want my scripts to call a local API" Ollama
"I want one API for text, vision, speech, and agents" LocalAI
"I want to evaluate quantizations before committing" LM Studio
"I want to run headless on a server or in Docker" Ollama or LocalAI
"I want to switch backends without changing clients" LocalAI

The right question is not "which runtime is best?" It is "what is the job, and which tool is built for it?"

What we actually run at CloudHerder

Our setup is deliberately split:

  • LM Studio for first-pass model evaluation. Does this model run? Which quantization fits? Is the output quality worth the VRAM?
  • LocalAI as the runtime our own tools talk to. It is the production-ish layer for a stack that is fully local, not cloud-hosted.
  • Ollama as the simpler path we did not need, but regularly recommend to anyone who just wants a local API without the extra configuration.

That split is not indecision. It is using the right tool for the right phase.

What's in your control, what's not, and what's next

In your control: which runtime you install, which model you load, which quantization you choose, and which API client you point at it.

Not in your control: whether a given model fits in your VRAM, whether an upstream project changes a license or terms of service, and whether your GPU drivers decide to cooperate on a Tuesday morning.

Next in this series: what "runs on 16 GB" actually means across quantizations and model sizes — and how to avoid the classic mistake of downloading a model that immediately crashes your drivers.

D

Dallum Brown

Writer and curator exploring the impact of technology on everyday life.

View All Articles

Subscribe to
The Brief

Our curated selection of tech news and other discoveries, delivered every month.

No spam. Unsubscribe anytime.

Comments (0)

Please sign in to leave a comment.

No comments yet. Be the first to share your thoughts!

Privacy Notice

We use essential cookies for site functionality (session management, CSRF protection) and do not track you across the web. By using this site, you acknowledge our Privacy Policy and New Zealand Privacy Act 2020 compliance.

Learn More