# Local LLM for companies: what runs, on what and when it pays off

> What a local LLM is, how much memory open-source models need, Ollama or LM Studio, and when it pays off more than a cloud service for a company.

URL: https://thenichesociety.ro/en/blog-llm-local-pentru-firme

A local LLM is a language model that runs on your own computer or server, without the text reaching an outside provider. **Sizing starts from memory: roughly 0.5 GB for every billion parameters at 4-bit quantization, plus room for context.** It pays off when the data is sensitive and the volume is steady, not as a free replacement for ChatGPT.

## What a local LLM is and how it differs from ChatGPT

A local LLM is a language model that runs on your own hardware, a laptop, a PC or a server in the company, instead of being accessed over the internet from a provider. The text you enter and the answers stay on that machine. This is possible because more and more developers publish open-weight models that anyone can download and run: OpenAI with gpt-oss, Google with Gemma or Alibaba with Qwen, all under the permissive Apache 2.0 license.

The practical difference from ChatGPT or Claude used in the browser is how responsibilities are split. In the cloud, the provider takes care of servers, updates and capacity, and you pay the subscription or the usage. Locally, you have full control over the data and the chosen model, but also the burden of the hardware, the speed and the maintenance. Open models have moved closer to some commercial ones: OpenAI says gpt-oss-120b comes close to o4-mini on core reasoning benchmarks, but for the hardest tasks the top cloud models usually keep an edge.

## How much memory you need for a local LLM

The basic sizing rule is memory, not the processor. Every parameter of the model takes up space: at FP16 precision it is two bytes, roughly 2 GB of memory for every billion parameters, and at 4-bit quantization, the variant usually used locally, roughly 0.5 GB. A model with 20 billion parameters therefore needs around 10 GB for the weights alone. The memory can be the graphics card's or, on computers with unified memory, the system's shared memory.

On top of the weights comes the memory for context, called the KV cache, which grows with every token in the conversation and with every user working at the same time. That's why a model that fits comfortably when one person tests it can run out of memory when a whole team uses it. Sizing guides recommend a reserve of 15-20% for running it. As concrete reference points, gpt-oss-20b is a 14 GB download and runs with 16 GB of memory, while gpt-oss-120b is 65 GB and fits on a single 80 GB GPU.

| Model size | Memory for weights, at 4 bits | Where it runs, roughly |
|---|---|---|
| 3-4 billion parameters | about 2 GB | an ordinary laptop, for tests |
| 12-14 billion | about 6-7 GB | a PC with a 12-16 GB graphics card |
| 20-32 billion | about 10-16 GB | a 24 GB graphics card or a Mac with at least 32 GB of unified memory |
| over 100 billion | over 50 GB | a server with an 80 GB GPU or several cards |

## Which open-source models you can run locally in 2026

The landscape has changed a lot in the past year, and the license has become as important as performance. Apache 2.0, the license used by gpt-oss, Gemma 4 and Qwen3, allows commercial use, modification and integration into products without a separate agreement with the developer. Other model families come with their own licenses, with conditions you need to read before using them in a product or in a service for clients. An “open-source” label in a presentation doesn't automatically mean you can do anything with the model.

For a company in Romania, language matters too. Gemma 4 declares support for more than 140 languages, and Qwen3 lists 119 languages and dialects, Romanian among them. That doesn't guarantee the same quality as in English, especially for legal, tax or technical terminology. The only test that counts is on your own documents: the same questions, put to several models, with the answers compared by someone who knows what a correct answer looks like in that field.

- 01**gpt-oss** OpenAI, since August 2025, in the 20b and 120b variants, Apache 2.0 license.
- 02**Gemma 4** Google, since April 2026, from small variants for devices up to a model with 31 billion parameters, Apache 2.0 license.
- 03**Qwen3** Alibaba, since April 2025, in several sizes, with declared support for Romanian, Apache 2.0 license.

## Ollama, LM Studio, llama.cpp or vLLM: what to run the model with

For the first tests you don't need anything complicated. Ollama, an open-source project under the MIT license, downloads and runs a model with a single command and offers an API compatible with OpenAI's, so many existing applications can connect to the local model without major changes. LM Studio does the same through a graphical interface and, since July 2025, it is free for use at work as well, with no separate license. For anyone who doesn't want a command line, it is usually the easiest starting point.

Under many of these applications runs llama.cpp, the open-source engine that made it possible to run quantized models on ordinary hardware. When you move from one person to a team, the problem becomes capacity: several users at the same time, fast answers for everyone and monitoring. That's where vLLM is usually used, an inference server under the Apache 2.0 license, built for many requests in parallel. The difference between a test on a laptop and a system a team uses every day has more to do with infrastructure than with the model.

| Tool | What it's good for | License |
|---|---|---|
| Ollama | Simple command-line use and a local OpenAI-compatible API | MIT |
| LM Studio | Graphical interface, individual tests, no command line | Free app, including at work |
| llama.cpp | The core engine for quantized models on ordinary hardware | MIT |
| vLLM | Server for many simultaneous users | Apache 2.0 |

## Local LLM vs cloud: what you gain and what you lose

The fair comparison is not between a free model and a subscription, but between two ways of paying. In the cloud you pay per user or by the volume of text processed, and the cost grows with use. Locally you pay for the hardware up front, then for the time of the people who install, update and secure it. At low or irregular volume, the cloud almost always comes out cheaper. At high and steady volume, with sensitive data, the balance can tip toward the local option.

There are also differences that don't show up in the numbers. Locally, the model doesn't change overnight without you knowing and it works without internet, but it doesn't get improvements on its own either. In the cloud you get immediate access to the newest models, but the data passes through someone else's infrastructure, even if under contract. For many companies, the realistic answer is a combination: tasks with sensitive data run locally, while the ones without risk go to a carefully chosen cloud service.

| Criterion | Local LLM | Cloud |
|---|---|---|
| Where the data goes | Stays on the company's hardware | Passes through the provider's servers |
| Top quality | Good, below the top models | The newest and largest models |
| Cost | Hardware paid up front, plus maintenance | Subscription or pay-as-you-go |
| Maintenance | Up to you or a partner | Up to the provider |
| Without internet | Works | Doesn't work |

## When a local LLM pays off for a company and when it doesn't

A local LLM makes sense when at least one of the conditions below is clearly met, not just in principle. Most often it is about data that, by contract or by its nature, must not reach an outside provider: client files, medical data, technical or financial documentation. For a company with many users and steady volume, the team version becomes a [private AI infrastructure](https://thenichesociety.ro/en/ai-engineering/ai-infrastructure), with a server sized for the number of people working at the same time.

It doesn't pay off for a one-month experiment, for a team of a few people who use AI occasionally or when nobody in the company and no partner takes care of maintenance. An unsupervised server, with an outdated model and access configured in a hurry, creates more risks than it solves. For a broader analysis of the self-hosted option, with the GDPR implications and the current limits, we've written separately about [self-hosted AI for companies](https://thenichesociety.ro/en/blog-ai-suveran-self-hosted-ue).

- 01**Sensitive data** documents that are not allowed to reach an outside provider.
- 02**Steady volume** daily use, large enough for the cloud cost to matter.
- 03**Offline operation** locations or workflows where internet access isn't guaranteed.
- 04**Someone owns the system** a person or partner who updates and secures it.

## How to test a local LLM in your company, step by step

A first test can be done in a few days and starts from a real task, not from a demo. Pick something the team does often, for example summarizing internal documents or answering questions about procedures, and for the first round prepare a set of documents without sensitive data. Run the same questions on a small model, a medium one and a cloud service, then compare the answers, the speed and the memory used.

If the results are good, the next step is not a more powerful laptop but a server designed for the team, with access control, logging and a safe way to search the company's documents. That usually means an [AI assistant on the company's documents](https://thenichesociety.ro/en/ai-engineering/document-assistant) that answers by citing its source, built on the model chosen in the test. If the results aren't good enough in Romanian for your task, you've found that out cheaply, before any investment in hardware.

- 01Pick a real task and a set of documents without sensitive data.
- 02Install Ollama or LM Studio on a test computer.
- 03Run the same questions on a small model and a medium one.
- 04Compare with a cloud service: quality in Romanian, speed, memory.
- 05Decide from the results whether a server for the team is worth it.

## Sources and further reading.

- 01[OpenAI — Introducing gpt-oss](https://openai.com/index/introducing-gpt-oss/)
- 02[Google — Gemma 4 model card](https://ai.google.dev/gemma/docs/core/model_card_4)
- 03[Qwen — Qwen3: Think Deeper, Act Faster](https://qwenlm.github.io/blog/qwen3/)
- 04[Ollama — official repository on GitHub](https://github.com/ollama/ollama)
- 05[LM Studio — LM Studio is free for use at work](https://lmstudio.ai/blog/free-for-work)
- 06[Runpod — GPU memory sizing guide for LLM inference](https://www.runpod.io/articles/guides/gpu-memory-sizing-guide-for-llm-inference)

## Frequently asked questions

### What is a local LLM?

A local LLM is a language model downloaded and run on your own hardware, a laptop, a PC or a server in the company. The text you enter and the answers don't reach an outside provider, but maintenance, updates and security are up to you.

### What computer do I need for a local LLM?

It depends on the size of the model. At 4-bit quantization, a model takes roughly 0.5 GB of memory for every billion parameters, plus room for context. A model with 20 billion parameters runs, roughly, on a 16-24 GB graphics card or on a Mac with enough unified memory.

### Ollama or LM Studio: which should I choose?

LM Studio is simpler if you want a graphical interface and are testing on your own computer. Ollama is a better fit when you want to connect the model to other applications through an API or run it on a server. Both can be used for free at work.

### Is a local LLM as good as ChatGPT?

For many everyday tasks, such as summaries, classification or answers based on documents, today's open models give good results. For complex reasoning and very hard tasks, the top cloud models usually keep an edge. The difference is best seen by testing on your own tasks.

### Can I use an open-source model commercially?

It depends on the license. Models published under Apache 2.0, such as gpt-oss, Gemma 4 and Qwen3, allow commercial use and modification. Other models have their own licenses with conditions, so read the license before using the model in a product or service.

### Does a local LLM understand Romanian?

Several open models declare support for Romanian: Qwen3 includes it among its 119 languages and dialects, and Gemma 4 declares more than 140 languages. Quality still differs from English, especially for specialist terms, so it's worth testing on the company's documents before any decision.

### Let's see what can be automated in your business.

A free 30-minute session: we'll tell you what can be automated, how long it takes and what it costs, with a fixed price after discovery.

We reply the same business day.
