Services AI Audit Guides Marketplace Blog Contact
Journal

Local AI, in five minutes. What actually runs on the machine you already own.

An owner asking this question has no interest in vocabulary, so here is the short version — with one thing worth knowing before any hardware quote gets signed. Since the spring of 2026, "will it run on my machine?" no longer has one answer. It has two, and confusing them has become the most common way to buy the wrong computer.

What it is, without the jargon

A model is a file. A very large file of numbers, downloaded once, which a small program reads in order to produce text. That is the whole of it. No subscription, no account, no permanent connection: once the file is on the disk you can unplug the network cable and it keeps answering.

The consequence fits in a sentence. Whatever you hand it to read goes nowhere — not the contract, not the client list, not the accounts, not an employee file. For a good number of trades that is not a refinement. It is the only arrangement under which anyone is willing to show a machine the real documents, and showing it the real documents is where nearly all the value sits.

The arithmetic that changed this year

For three years the rule of thumb was convenient: a model wants roughly as many gigabytes of memory as it has billions of parameters, once compressed. We wrote it that way on our own hardware page. It has become incomplete, and it is better said here than discovered after the purchase order.

What broke it has a name not worth memorising — mixture of experts — and an effect that shows up on an invoice. Take Mistral Small 4, released on 16 March 2026 by the Paris company with downloadable weights under an Apache 2.0 licence: 119 billion parameters in total, of which only 6.5 billion do any work on a given response. It therefore has the speed of a small model and the footprint of a very large one. Compressed to 4 bits, the file wants something in the order of sixty gigabytes of memory to load. The word "Small" in the name is not describing your computer.

Hence the two answers, and this is the only technical sentence in the piece: memory follows the total parameter count, speed follows the active one. Anyone praising a model for being fast has told you nothing about the machine required to hold it. That gap is precisely where hardware budgets go wrong.

Memory in the machineWhat fits, compressed to 4 bitsWhat it honestly does
8GBModels of 3 to 4 billion parametersSummarising, sorting, extracting. Decent short drafting.
16GBModels of 7 to 8 billionThe realistic entry point. Covers most desk work.
32GB12 to 14 billion, with comfortable contextMarkedly better on long documents and formal register.
64GB and upLarge open-weight models, mixture-of-experts includedThe closest thing to cloud quality you can keep in the building.

The middle column is arithmetic rather than opinion: a parameter compressed to 4 bits weighs half a byte, and the remainder is room for the text being worked on. The right-hand column is not arithmetic — and it is the only one that decides anything.

What an eight-billion model actually does

On a 16GB machine you already own, a model that size reads a forty-page contract and returns the dates and the penalty clauses. It sorts two hundred emails by urgency. It writes the first version of a reply, a set of minutes, a translation. It finds, in your own past quotes, what was charged on a comparable job. These are dull, frequent tasks, and that is where free-per-run matters: locally, the two-hundredth answer costs exactly what the first one did, which is nothing.

What it does not do deserves the same directness. It gets long chains of reasoning wrong, and arithmetic wrong more often than that. It knows nothing of what happened after it was trained. It invents citations with the same composure it uses for real ones. A local model removes the blank page. It does not remove the read-through, and it certainly does not remove responsibility for what leaves your business with your name on it.

Language is not a comfort feature

For a firm in Kanaky (New Caledonia), Vanuatu or Fiji, a model's ranking on English-language tests says little about what it is worth on the correspondence you actually send. Where the work is bilingual, European open-weight models are the serious option: Mistral publishes its weights and aims its Les Ministraux family explicitly at edge devices, personal computers and phones included. Speaking to TechCrunch on 4 July 2026, chief executive Arthur Mensch flagged a further open-weight model for the summer with early access opening in July. Artificial Analysis scores Mistral Small 4 at 20 on its intelligence index, against a median of 9 for open-weight models in the same class.

The test that counts appears in no league table. Take three documents your business genuinely produced and see what the model makes of them. A benchmark run in a laboratory will never tell you whether the thing understands how a purchase order is worded here.

The five minutes, in practice

  1. Install Ollama on the machine you already have. One command, free, and reversible — uninstalling leaves nothing behind.
  2. Download exactly one model in the 7 to 8 billion range. One. The urge to try six of them costs a week and concludes nothing.
  3. Give it real work for five days: your documents, your correspondence, your quotes. Not questions designed to test it.
  4. Keep two columns, "good enough" and "not good enough". That list is the only deliverable of the week, and it is worth more than any comparison article online.
  5. Only then decide whether hardware is warranted. An hour of scoping usually settles it, and the answer is often no.

When the answer is no

Three cases where building this makes no sense and a cloud subscription remains the right call. If the volume is low — a few dozen requests a month — the hardware will never pay itself back, and the arithmetic is not close. If your data is not particularly sensitive, you are buying privacy you have no use for. And if nobody in the business wants to look after a machine, the system will be switched off in a cupboard within two months. That is not a technology problem, and no amount of hardware fixes it.

Common questions

How much memory do I need to run a model locally?

Budget half a byte per parameter for a model compressed to 4 bits, plus headroom for the text being processed. An 8-billion-parameter model therefore occupies about 5 to 6GB and sits comfortably in a 16GB machine, which describes a great many computers already in service. Below 8GB you can still run very small models, but the drop in quality is obvious within a day of real work.

Why does a model called "Small" need sixty gigabytes?

Because "small" now refers to how many parameters work on each response, not to the size of the file you have to load. Mistral Small 4, released on 16 March 2026, holds 119 billion parameters of which only 6.5 billion are active per request: it answers with the speed of a small model, but the full set of weights must fit in memory, which is roughly sixty gigabytes once compressed to 4 bits. That distinction is the thing to check before buying any hardware.

Does a local model really work without an internet connection?

Yes, entirely, once the file is downloaded. It is the property that separates local from cloud most sharply, and it matters more in a territory whose connectivity depends on a single submarine cable. What still needs a connection is updating the model and anything you wire around it — mail, calendars, third-party tools. The model itself does not.

Which open-weight models handle languages other than English well?

European releases are currently the strongest argument for local AI in a bilingual business. Mistral publishes downloadable weights, Apache 2.0 in the case of Mistral Small 4, and its Les Ministraux family targets personal devices. That said, a model's position in an English-language ranking tells you very little about your case. The only worthwhile check is to feed it three of your own documents in the language you actually work in and judge the output yourself.

Keep reading
Hardware requirements for local models Local versus cloud, compared honestly Book a free AI opportunity audit

Find out what AI
can do for you.

Kanaky Tech is an AI automation agency working across New Zealand and the Pacific. Start with a free AI opportunity audit: we map how your business actually runs, rank what is worth automating, and give you a clear scope before anything is built — no obligation.

Book a free AI audit AI automation in New Zealand