Blog · AI

Local LLM for companies: what runs, on what and when it pays off

A local LLM is a language model that runs on your own computer or server, without the text reaching an outside provider. Sizing starts from memory: roughly 0.5 GB for every billion parameters at 4-bit quantization, plus room for context. It pays off when the data is sensitive and the volume is steady, not as a free replacement for ChatGPT.

8minute read
2026-09-17published
AIcategory
Tower servers standing side by side, with their status lights on
AI
01

What a local LLM is and how it differs from ChatGPT

A local LLM is a language model that runs on your own hardware, a laptop, a PC or a server in the company, instead of being accessed over the internet from a provider. The text you enter and the answers stay on that machine. This is possible because more and more developers publish open-weight models that anyone can download and run: OpenAI with gpt-oss, Google with Gemma or Alibaba with Qwen, all under the permissive Apache 2.0 license.

The practical difference from ChatGPT or Claude used in the browser is how responsibilities are split. In the cloud, the provider takes care of servers, updates and capacity, and you pay the subscription or the usage. Locally, you have full control over the data and the chosen model, but also the burden of the hardware, the speed and the maintenance. Open models have moved closer to some commercial ones: OpenAI says gpt-oss-120b comes close to o4-mini on core reasoning benchmarks, but for the hardest tasks the top cloud models usually keep an edge.

02

How much memory you need for a local LLM

The basic sizing rule is memory, not the processor. Every parameter of the model takes up space: at FP16 precision it is two bytes, roughly 2 GB of memory for every billion parameters, and at 4-bit quantization, the variant usually used locally, roughly 0.5 GB. A model with 20 billion parameters therefore needs around 10 GB for the weights alone. The memory can be the graphics card's or, on computers with unified memory, the system's shared memory.

On top of the weights comes the memory for context, called the KV cache, which grows with every token in the conversation and with every user working at the same time. That's why a model that fits comfortably when one person tests it can run out of memory when a whole team uses it. Sizing guides recommend a reserve of 15-20% for running it. As concrete reference points, gpt-oss-20b is a 14 GB download and runs with 16 GB of memory, while gpt-oss-120b is 65 GB and fits on a single 80 GB GPU.

Model sizeMemory for weights, at 4 bitsWhere it runs, roughly
3-4 billion parametersabout 2 GBan ordinary laptop, for tests
12-14 billionabout 6-7 GBa PC with a 12-16 GB graphics card
20-32 billionabout 10-16 GBa 24 GB graphics card or a Mac with at least 32 GB of unified memory
over 100 billionover 50 GBa server with an 80 GB GPU or several cards
03

Which open-source models you can run locally in 2026

The landscape has changed a lot in the past year, and the license has become as important as performance. Apache 2.0, the license used by gpt-oss, Gemma 4 and Qwen3, allows commercial use, modification and integration into products without a separate agreement with the developer. Other model families come with their own licenses, with conditions you need to read before using them in a product or in a service for clients. An “open-source” label in a presentation doesn't automatically mean you can do anything with the model.

For a company in Romania, language matters too. Gemma 4 declares support for more than 140 languages, and Qwen3 lists 119 languages and dialects, Romanian among them. That doesn't guarantee the same quality as in English, especially for legal, tax or technical terminology. The only test that counts is on your own documents: the same questions, put to several models, with the answers compared by someone who knows what a correct answer looks like in that field.

  • 01gpt-oss OpenAI, since August 2025, in the 20b and 120b variants, Apache 2.0 license.
  • 02Gemma 4 Google, since April 2026, from small variants for devices up to a model with 31 billion parameters, Apache 2.0 license.
  • 03Qwen3 Alibaba, since April 2025, in several sizes, with declared support for Romanian, Apache 2.0 license.
04

Ollama, LM Studio, llama.cpp or vLLM: what to run the model with

For the first tests you don't need anything complicated. Ollama, an open-source project under the MIT license, downloads and runs a model with a single command and offers an API compatible with OpenAI's, so many existing applications can connect to the local model without major changes. LM Studio does the same through a graphical interface and, since July 2025, it is free for use at work as well, with no separate license. For anyone who doesn't want a command line, it is usually the easiest starting point.

Under many of these applications runs llama.cpp, the open-source engine that made it possible to run quantized models on ordinary hardware. When you move from one person to a team, the problem becomes capacity: several users at the same time, fast answers for everyone and monitoring. That's where vLLM is usually used, an inference server under the Apache 2.0 license, built for many requests in parallel. The difference between a test on a laptop and a system a team uses every day has more to do with infrastructure than with the model.

ToolWhat it's good forLicense
OllamaSimple command-line use and a local OpenAI-compatible APIMIT
LM StudioGraphical interface, individual tests, no command lineFree app, including at work
llama.cppThe core engine for quantized models on ordinary hardwareMIT
vLLMServer for many simultaneous usersApache 2.0
05

Local LLM vs cloud: what you gain and what you lose

The fair comparison is not between a free model and a subscription, but between two ways of paying. In the cloud you pay per user or by the volume of text processed, and the cost grows with use. Locally you pay for the hardware up front, then for the time of the people who install, update and secure it. At low or irregular volume, the cloud almost always comes out cheaper. At high and steady volume, with sensitive data, the balance can tip toward the local option.

There are also differences that don't show up in the numbers. Locally, the model doesn't change overnight without you knowing and it works without internet, but it doesn't get improvements on its own either. In the cloud you get immediate access to the newest models, but the data passes through someone else's infrastructure, even if under contract. For many companies, the realistic answer is a combination: tasks with sensitive data run locally, while the ones without risk go to a carefully chosen cloud service.

CriterionLocal LLMCloud
Where the data goesStays on the company's hardwarePasses through the provider's servers
Top qualityGood, below the top modelsThe newest and largest models
CostHardware paid up front, plus maintenanceSubscription or pay-as-you-go
MaintenanceUp to you or a partnerUp to the provider
Without internetWorksDoesn't work
06

When a local LLM pays off for a company and when it doesn't

A local LLM makes sense when at least one of the conditions below is clearly met, not just in principle. Most often it is about data that, by contract or by its nature, must not reach an outside provider: client files, medical data, technical or financial documentation. For a company with many users and steady volume, the team version becomes a private AI infrastructure, with a server sized for the number of people working at the same time.

It doesn't pay off for a one-month experiment, for a team of a few people who use AI occasionally or when nobody in the company and no partner takes care of maintenance. An unsupervised server, with an outdated model and access configured in a hurry, creates more risks than it solves. For a broader analysis of the self-hosted option, with the GDPR implications and the current limits, we've written separately about self-hosted AI for companies.

  • 01Sensitive data documents that are not allowed to reach an outside provider.
  • 02Steady volume daily use, large enough for the cloud cost to matter.
  • 03Offline operation locations or workflows where internet access isn't guaranteed.
  • 04Someone owns the system a person or partner who updates and secures it.
07

How to test a local LLM in your company, step by step

A first test can be done in a few days and starts from a real task, not from a demo. Pick something the team does often, for example summarizing internal documents or answering questions about procedures, and for the first round prepare a set of documents without sensitive data. Run the same questions on a small model, a medium one and a cloud service, then compare the answers, the speed and the memory used.

If the results are good, the next step is not a more powerful laptop but a server designed for the team, with access control, logging and a safe way to search the company's documents. That usually means an AI assistant on the company's documents that answers by citing its source, built on the model chosen in the test. If the results aren't good enough in Romanian for your task, you've found that out cheaply, before any investment in hardware.

  • 01Pick a real task and a set of documents without sensitive data.
  • 02Install Ollama or LM Studio on a test computer.
  • 03Run the same questions on a small model and a medium one.
  • 04Compare with a cloud service: quality in Romanian, speed, memory.
  • 05Decide from the results whether a server for the team is worth it.
08

Sources and further reading.

FAQ

Frequently asked questions

What is a local LLM?

A local LLM is a language model downloaded and run on your own hardware, a laptop, a PC or a server in the company. The text you enter and the answers don't reach an outside provider, but maintenance, updates and security are up to you.

What computer do I need for a local LLM?

It depends on the size of the model. At 4-bit quantization, a model takes roughly 0.5 GB of memory for every billion parameters, plus room for context. A model with 20 billion parameters runs, roughly, on a 16-24 GB graphics card or on a Mac with enough unified memory.

Ollama or LM Studio: which should I choose?

LM Studio is simpler if you want a graphical interface and are testing on your own computer. Ollama is a better fit when you want to connect the model to other applications through an API or run it on a server. Both can be used for free at work.

Is a local LLM as good as ChatGPT?

For many everyday tasks, such as summaries, classification or answers based on documents, today's open models give good results. For complex reasoning and very hard tasks, the top cloud models usually keep an edge. The difference is best seen by testing on your own tasks.

Can I use an open-source model commercially?

It depends on the license. Models published under Apache 2.0, such as gpt-oss, Gemma 4 and Qwen3, allow commercial use and modification. Other models have their own licenses with conditions, so read the license before using the model in a product or service.

Does a local LLM understand Romanian?

Several open models declare support for Romanian: Qwen3 includes it among its 119 languages and dialects, and Gemma 4 declares more than 140 languages. Quality still differs from English, especially for specialist terms, so it's worth testing on the company's documents before any decision.

The Niche Society
The Niche Society TeamAI and software engineers from Bucharest · LinkedIn
published 2026-09-17

Let's see what can be automated in your business.

A free 30-minute session: we'll tell you what can be automated, how long it takes and what it costs, with a fixed price after discovery.

Book a free sessionoffice@thenichesociety.ro

We reply the same business day.

+40 733 045 833