Running a Local LLM: What Your Machine Can Actually Handle

A laptop showing code on a wooden desk in a dim room, the kind of everyday machine that can run a local LLM
Photo by Daniil Komov on Pexels

A local LLM is a large language model that runs entirely on your own computer, with no API key and no request leaving the machine. On a laptop with 16 GB of memory you can have one answering questions in about ten minutes. The hard part is choosing a model that fits in memory and still runs at a speed you can live with.

Almost every decision in this guide comes back to memory. Know how much of it the model gets, and the rest of the choices get easier.

Why run an LLM locally at all

The honest reasons are privacy and control. Hosted models are smarter than anything you will run at home, so the case for local rests on what they cannot give you.

Privacy is the headline. When the model runs on your machine, a contract or a patient note never crosses the network. For teams in regulated work, that can turn a months-long vendor review into a conversation about one laptop.

Control is the quieter benefit. A hosted model can be updated overnight, and a prompt you tuned for weeks starts behaving differently on a Tuesday. A model file on your disk answers the same way next year.

Cost matters less than people expect at small scale. Once the hardware is paid for, each extra token costs electricity. Before buying a graphics card for the purpose, check it against what each Claude plan really costs to run.

Then there is working offline. A local model answers on a plane, and during the outage that always seems to arrive the afternoon before a deadline.

What hardware you need, and does an LLM use Apple unified memory?

Memory decides what you can run, and compute decides how fast. A model that does not fit in memory runs badly at any speed.

The rule of thumb is parameters multiplied by bits per weight, divided by eight, which gives bytes. An 8-billion-parameter model stored at 16 bits needs about 16 GB for its weights alone. At 4 bits the same model needs about 4 GB, plus a gigabyte or two for the KV cache, the working memory that holds the conversation so far.

On a PC, the number that matters is VRAM, the memory on the graphics card itself. A 12 GB card runs 8B models comfortably, and a 24 GB card reaches models in the low 30-billion range at 4 bits. When a model spills past VRAM, llama.cpp can split it across CPU and GPU, and speed drops with every layer left on the CPU.

Close-up of a graphics card with metal cooling fans, whose VRAM sets the size of local LLM it can run
Photo by Sergei Starostin on Pexels

Apple Silicon changes the maths. The CPU and GPU share one pool of unified memory, so a Mac with 64 GB can hand most of it to the model. The llama.cpp project calls Apple silicon a first-class citizen and runs models on its GPU through Metal. macOS keeps a share back for itself, so plan on most of the pool and never all of it.

Speed comes from memory bandwidth. Generating each token means reading every active weight once, so 4 GB of weights on a machine that moves 100 GB per second tops out near 25 tokens per second. That is why a big Mac runs large models slowly and a gaming card runs small ones fast. Upgrading the chip without the memory buys you a sports car for a road that is still jammed.

How to run an LLM locally in about ten minutes

The shortest path is Ollama, a free tool that downloads models and serves them from your own machine. Install it from the official site, then run two commands in a terminal:

ollama pull llama3.1:8b
ollama run llama3.1:8b "Explain a context window in two sentences."

The first command downloads about 5 GB of weights, already compressed to 4 bits. The second loads them and answers. Ollama also serves a local API on port 11434, which is how other software talks to the model:

curl http://localhost:11434/api/generate \
  -d '{"model": "llama3.1:8b", "prompt": "Why is the sky blue?", "stream": false}'

The default that catches people is the context window, meaning how much text the model can hold at once. The Ollama FAQ puts the default at 4,096 tokens. Paste in a 40-page document and you get a confident summary of the part that fit, with no error to say the rest never made it in. Raise the limit with the OLLAMA_CONTEXT_LENGTH environment variable, and budget for the extra memory it costs.

A silver Mac mini on a white desk, a small machine that can run a local LLM from unified memory
Photo by Pavel Danilyuk on Pexels

LM Studio offers the same workflow in a desktop app, if you prefer clicking to typing. Either way, each model expects its own chat template, the special tokens marking where the user stops and the assistant starts.

Load the wrong template and the model starts writing both sides of the conversation, which is impressive for about four seconds. The tools handle this for their own libraries; a raw file from elsewhere is where it bites, and how special chat tokens get trained into an LLM explains why.

LLM optimization for small machines: quantization

Quantization stores each weight in fewer bits, trading a little accuracy for a much smaller model. It is the one optimization that makes local models practical.

Quantizing a 16-bit model to 4 bits cuts it to a quarter of the size, and the loss is smaller than you would expect. Across more than 35,000 experiments, Dettmers and Zettlemoyer found 4-bit precision almost universally optimal for a fixed memory budget. Put plainly, a bigger model at 4 bits usually beats a smaller model at 16 bits that takes the same space.

Below 4 bits the curve bends. The llama.cpp project supports formats down to 1.5 bits, and a 2-bit model will run, but it loses reasoning quality in ways a quick chat hides and a real task exposes. Quantization is packing for a fortnight in a carry-on. At 4 bits you leave behind the fourth pair of shoes; at 2 bits you start leaving behind the passport.

GGUF, the file format llama.cpp uses, labels quantized files with tags such as Q4_K_M or Q8_0, where the number is roughly the bits per weight. Start at Q4_K_M, and move up to Q5 or Q6 when you have spare memory and notice mistakes.

Quantization changes how a model is stored. Changing what it knows is fine-tuning, covered in depth in training and adapting large language models yourself, and worth reading before deciding whether to build an LLM of your own.

Free LLM models worth downloading first

Nearly every open-weight model is a free LLM in the sense that matters: nothing to pay to download it and nothing per token. What varies is size and licence.

Pick by memory first. This table assumes 4-bit quantization and leaves headroom for a longer context window.

Memory for the modelSize that fits at 4 bitsWhere to start
8 GB3B to 8B, tight at the topSmall Llama or Qwen models
16 GB8B to 14B, plus gpt-oss-20bLlama 3.1 8B, gpt-oss-20b
24 to 32 GBUp to about 30BGemma or Mistral Small in the mid-20B range
64 GB and up70B classLlama 70B-class models

One model in that table works differently. The gpt-oss-20b model is a mixture-of-experts design, meaning only a slice of its weights works on each token. OpenAI's model card lists 21 billion parameters with 3.6 billion active, and says it runs within 16 GB of memory. You pay for the memory of a large model and get the speed of a small one.

Free also comes with terms. OpenAI releases gpt-oss under the Apache 2.0 licence, while other families attach their own licences, some with conditions on commercial use. Read the licence file before a model goes anywhere near a product.

When a private LLM beats the API, and when it loses

A private LLM is a model running on hardware you control, and a local LLM on a laptop is the smallest version of one. It wins when the data cannot leave and the task is narrow enough for an open model to do well.

The privacy claim holds at the inference layer. The Ollama FAQ states that it runs locally and that Ollama does not see your prompts or data when you do. Watch the model names, though: Ollama also offers cloud-hosted models, tagged as cloud, and those run on remote servers.

Ollama listens only on your own machine by default. If you change that to share a model across an office, anyone who can reach the port can query it without a password. Put an authenticating proxy in front before you open it up.

Local loses on raw capability. The strongest hosted models still outperform anything that fits in 64 GB on hard reasoning and long documents. It also loses on upkeep, since you now own the updates and the hardware that fails at 2 a.m.

A sensible first week looks like this:

  • Check how much memory the model can use: VRAM on a PC, most of unified memory on a Mac.
  • Install Ollama and pull an 8B model at Q4_K_M.
  • Raise the context length before feeding it real documents.
  • Run your actual task against the local model and a hosted one, side by side.
  • Keep the local model for the work where it holds up, and send the rest to the API.

Hosted models will keep improving. So will the ones on your desk, and those have the courtesy to stay exactly the same until you decide otherwise.

Frequently asked questions

Is ChatGPT an LLM?

ChatGPT is a chat product built on OpenAI's GPT family of large language models, so the engine underneath is an LLM. You cannot download the models behind ChatGPT, but OpenAI has released open-weight gpt-oss models under Apache 2.0, and gpt-oss-20b runs within 16 GB of memory.

Is a local LLM free to run?

The software and most open weights cost nothing, so the running cost is electricity plus hardware you already own. Check each model's licence, since some attach conditions on commercial use. If you would need to buy a GPU, price it against a year of API usage first.

Does an LLM use Apple unified memory?

Yes. On Apple Silicon the CPU and GPU share one memory pool, and runtimes such as llama.cpp run models on the GPU through Metal. A Mac with 64 GB can therefore hold models too large for any 24 GB graphics card, minus the share macOS keeps for itself.

How do you use a different LLM with Claude Code?

Ollama exposes an Anthropic-compatible API on your machine, so Claude Code can send its requests to a local model. Ollama's documentation points ANTHROPIC_BASE_URL at http://localhost:11434 and recommends a context length of 64k or more for real codebases. Expect weaker coding results than Claude itself.

Can you train an LLM on your own computer?

Training from scratch is out of reach, since it takes trillions of tokens and a cluster of accelerators. Fine-tuning a small open model with a parameter-efficient method such as LoRA is realistic on one consumer GPU with enough memory, and it is where most useful customisation happens.

What is a private LLM?

A private LLM is a model you run on hardware you control, so prompts and outputs stay inside your network. A local LLM on a laptop is the smallest version. How private it really is depends on who can reach the model's API port.

Sources