Local models: a complete overview

Open language models on your own server: seven families, memory for each size, local versus cloud, three examples with Ollama, mistakes and rules for production.

Stack and technologies Updated

In short

Local models are open language models that run on your own server or computer instead of a provider’s cloud: Qwen, Llama, Mistral, Gemma, gpt-oss and others. Data never leaves your perimeter, there is no bill per request, and the model works without the internet. The price is hardware and quality: a model that fits on an ordinary server is weaker than the best cloud ones, so local models are chosen for clear, repeatable tasks — classification, extraction of data from documents, search by meaning, short answers from a knowledge base. Tools such as Ollama and llama.cpp start a model with one command and give it the same API as the cloud.

Local models at a glance

The main facts in one table — what is meant, what is needed and how it is run.

What it is
Language models with open weights that run on your own hardware
Families
Qwen, Llama, Mistral, Gemma, gpt-oss, DeepSeek, Phi
Sizes
From 1 to hundreds of billions of parameters; for a server usually 4–32 billion
Quantisation
Weights compressed to 4–8 bits: several times less memory, a small loss of quality
How to run
Ollama, llama.cpp, vLLM, LM Studio
API
OpenAI-compatible — code moves between cloud and local by changing the address
Licences
From Apache 2.0 and MIT to licences with conditions — read before commercial use

Open model families

Seven families most often run locally. New versions come out every few months, so the choice is made on your own examples, not by rankings.

FamilyAuthorLicenceStrong side
Qwen Alibaba mostly Apache 2.0 many sizes, many languages
Llama Meta own, with conditions the largest ecosystem
Mistral Mistral AI Apache 2.0 for many models compact and fast
Gemma Google own, with conditions small models, pictures as input
gpt-oss OpenAI Apache 2.0 reasoning and tools
DeepSeek DeepSeek MIT for many models reasoning, code
Phi Microsoft MIT very small models

How much memory a model needs

Approximate memory for models compressed to 4 bits, plus a margin for the context. On a graphics card it is video memory; without one — ordinary memory, and the answer is several times slower.

Model sizeMemorySuits
1–4 billion about 2–4 GB classification, short extraction, phones
7–9 billion about 6–8 GB answers from a knowledge base, summaries
12–14 billion about 10–12 GB more complex texts and instructions
27–32 billion about 20–24 GB close to cloud quality on many tasks
70 billion and more from about 40 GB a dedicated server with graphics cards
Embedding models about 0.5–2 GB search by meaning, even without a graphics card

Local model or cloud API

The choice is made per task: one project often uses both.

CriterionLocal modelCloud API
Where the data goes nowhere to the provider
Quality on hard tasks lower the best available
Cost hardware, then almost free per request
Large volumes profitable the bill grows with the volume
Speed depends on your hardware high and stable
Without the internet works does not
Maintenance on you: updates, monitoring on the provider

What running a local model looks like: 3 examples

Starting a model with Ollama, one client for cloud and local models, and settings for a production server.

Start with Ollama

Two commands to a working model, and the same model over HTTP for applications.

start.sh
# Download a model and ask it a question — all on your own server
ollama pull qwen3:8b
ollama run qwen3:8b "Summarise the delivery terms in two sentences"

# The same model over HTTP: an OpenAI-compatible API on port 11434
curl http://localhost:11434/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "qwen3:8b", "messages": [{"role": "user", "content": "Hello"}]}'

One client for both

The code does not know where the model is: the address and the name come from settings.

llm.ts
// One client for a cloud and a local model: only the address, the key and the model change
const LLM_URL = process.env.LLM_URL ?? 'http://localhost:11434/v1';
const LLM_MODEL = process.env.LLM_MODEL ?? 'qwen3:8b';

export async function complete(prompt: string): Promise<string> {
  const res = await fetch(`${LLM_URL}/chat/completions`, {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      Authorization: `Bearer ${process.env.LLM_KEY ?? 'local'}`,
    },
    body: JSON.stringify({
      model: LLM_MODEL,
      messages: [{ role: 'user', content: prompt }],
      temperature: 0.2, // fewer surprises in business answers
    }),
  });
  if (!res.ok) throw new Error(`LLM: HTTP ${res.status}`);
  const data = (await res.json()) as { choices: { message: { content: string } }[] };
  return data.choices[0].message.content;
}

Settings for a server

Closed to the outside, warm between requests and limited in memory so as not to affect the sites next to it.

override.conf
# /etc/systemd/system/ollama.service.d/override.conf — Ollama on a production server
[Service]
# Listen only for applications on this server, not for the internet
Environment="OLLAMA_HOST=127.0.0.1:11434"
# Keep the model in memory between requests — no cold start
Environment="OLLAMA_KEEP_ALIVE=24h"
# How many requests one model serves at the same time
Environment="OLLAMA_NUM_PARALLEL=4"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
# A ceiling: the model cannot take the memory of the whole server
MemoryMax=24G

Common mistakes with local models

  1. Expecting cloud quality

    A model that fits on one server will not replace the best cloud model on complex tasks.

  2. Choosing by rankings

    Public benchmarks say little about your documents and your language.

  3. An API open to the internet

    A model server without authentication becomes free computing for strangers.

  4. No memory limit

    A large model loaded next to the sites takes their memory, and they start to fail.

  5. Ignoring the licence

    Some open models have conditions for commercial use and large audiences.

  6. One model for everything

    A small model for sorting and a strong one for complex answers are cheaper and better than one compromise.

6 rules for local models in production

  1. 01

    A test set first

    30–50 real tasks with expected answers — every model is compared on them.

  2. 02

    The smallest model that copes

    Faster answers, less memory, more requests in parallel.

  3. 03

    The same API as the cloud

    The model can be swapped either way without rewriting the code.

  4. 04

    Closed to the outside

    Only applications on the server or the internal network can reach it.

  5. 05

    Limits on memory and processor

    The model shares the server with other work and must not take everything.

  6. 06

    Versions are pinned

    An exact model version, and a new one only after a run on the test set.

Questions about local models

What is a local language model?

An open model that runs on your own server or computer instead of a provider’s cloud.

Do I need a graphics card?

For fast answers from models of 7 billion parameters and more — yes; small models and embeddings work on the processor.

Are local models as good as Claude or GPT?

On narrow, clear tasks they come close; on complex reasoning and long texts the cloud is stronger.

Is it really free?

The model is free; the server, electricity and maintenance are not. It pays off on large volumes.

Ollama or vLLM?

Ollama — simple start and moderate load; vLLM — many parallel users on graphics cards.

Can a local model be fine-tuned?

Yes, open weights allow it; but for knowledge of your documents RAG is usually enough.

Can it run on a laptop?

Yes, models up to about 8 billion parameters work on a modern laptop with 16 GB of memory.

Online form

AI inside
your perimeter

I choose the model for the task: Claude or GPT where quality matters most, local models where data cannot leave your servers. Tell me about the task — I answer within one working day.

Or write to [email protected]