# Private AI infrastructure: on‑premise or in the EU cloud

> Private AI infrastructure for companies: an LLM run on-premise or on a dedicated GPU in the EU, with RAG, agents on your ERP, and data that never leaves…

URL: https://thenichesociety.ro/en/ai-engineering/ai-infrastructure

An external API is simple to get started with, but contracts, access codes and customer data end up on servers outside the company. **We build AI platforms that run on-premise or on a dedicated GPU in the EU, with data kept inside the perimeter the company sets.**

## What self-hosted AI infrastructure means for a company

Self-hosted AI infrastructure means the language model runs on servers the company controls, or on a dedicated GPU at an EU provider — not on a public American API that you send contracts, access codes or customer data to, request by request. The difference isn't cosmetic: the data never leaves the agreed perimeter again.

We're not just talking about “running a model”. A private AI platform means an interface for employees, role-based permissions, connections to internal documents and systems, audit logs and monitoring — everything a generic AI chat subscription lacks.

## The platform isn't just a model: interface, RAG, agents, audit

A private LLM, on its own, doesn't solve anything for a company. The value appears when the model is connected to an interface employees actually use, to the company's documents and, where it makes sense, to existing operational systems.

We build each component as part of the same platform, not as separate tools manually stitched together. The part that makes the platform useful every day, [connecting agents to your ERP through MCP servers](https://thenichesociety.ro/en/ai-engineering/erp-integration-mcp), has its own separate page.

**Interface with roles, permissions and SSO**
Every employee signs in with the company account; what they can see and request depends on their role, not on good judgment.

**Search across the company's documents**
Answers come from internal procedures, contracts and reports, not from the model's general knowledge.

**Connectors via API or custom-built MCP**
Agents that read and, with explicit approval, write into existing systems — ERP, CRM, SAGA, SAP.

**Usage and access logs**
Who asked what, and what the model answered — visible internally, not hidden behind an external API.

**Evaluation sets for each use case**
We measure accuracy on the company's real data, not on a generic public benchmark.

## How big does an AI server actually need to be

Sizing isn't based on “the biggest model available”, but on the use case: how many employees use the system at the same time, how long the processed context is (whole documents or short questions), and the size and quantization of the chosen model — these variables decide the video memory needed.

A very large MoE (mixture-of-experts) model can need hundreds of GB of video memory even when quantized; a 20-30 billion-parameter model, by contrast, runs comfortably on a single high-memory GPU. For most internal processes, the right option is the second one — not the biggest model, but the one that fits the task. For a first test on a single computer, we've described separately how to run a [local LLM](https://thenichesociety.ro/en/blog-llm-local-pentru-firme) and how much memory it needs.

## Private LLM: on-premise at your site or on a dedicated GPU in the EU

There are two ways to run a private LLM, not just one. On your own servers (on-premise), the company physically controls the infrastructure — the right option where internal or contractual policy explicitly requires it. On a dedicated GPU at a provider in the European Union, the company doesn't buy hardware, but the data still doesn't leave the EU or reach an American provider.

The choice also depends on the data residency required by GDPR, the transparency obligations under the AI Act and, for regulated sectors, NIS2 requirements — factual topics to raise with the company's legal team, not an area where we provide legal advice. If the terms get mixed up in internal discussions, we've explained separately the [difference between self-hosted AI and sovereign AI](https://thenichesociety.ro/en/blog-ai-suveran-self-hosted-ue). The same infrastructure can run an [AI assistant on the company's documents](https://thenichesociety.ro/en/ai-engineering/document-assistant), with answers that cite their source.

## Self-hosted or API: where's the break-even point

There's no universal answer. For few users and bursty traffic, pay-per-token on a hosted API is almost certainly cheaper than running your own server — the infrastructure would sit idle outside usage peaks.

For several dozen daily users, constant traffic and sensitive data, self-hosted starts to win out — both on cost at high volume and on control over the data. The exact threshold depends on the number of users and the real volume of requests; we calculate it during discovery, using the company's own numbers, not a generic table.

## From discovery to infrastructure handed over to the company

We start with discovery: use cases, what data exists and how clean it is, security requirements, a sizing estimate. A pilot in a single department follows, with acceptance criteria written in advance — exactly as with any serious AI implementation.

After the pilot, we scale into production with active monitoring and, at the end, hand over the code and infrastructure to the company. They stay yours — we don't operate a closed platform you depend on indefinitely.

## A fixed price, after a short discovery call.

- use case and data audit
- security and compliance requirements
- sizing recommendation
- model + runtime selection (vLLM/SGLang)
- RAG on the department's documents
- written acceptance criteria
- multi-department rollout
- monitoring and audit logs
- code + infrastructure handover
Every project has a different context, workflows and infrastructure, so the price is set after a short, paid discovery and does not change along the way.

## Frequently asked questions

### How big does the server need to be?

It depends on three things: how many employees use it at the same time, how long the processed context is, and which model you choose. A very large MoE-type model can need hundreds of GB of video memory even when quantized; a 20-30 billion-parameter model runs on a single high-memory GPU — the right fit for most internal processes. The exact sizing is done during discovery, based on your use case.

### Can I use local AI with data from my ERP?

Yes — agents connect to your ERP/CRM via API or custom-built MCP connectors, read the data they need and, with explicit approval, can write back into the system. The whole flow stays on your chosen self-hosted or EU infrastructure, without ERP data ever passing through an external API.

### Is it cheaper than ChatGPT Team/Enterprise?

There's no universal answer. For few users and bursty traffic, a subscription or pay-per-token usually stays cheaper. For several dozen daily users, constant traffic and sensitive data, self-hosted starts to win out — the exact threshold depends on the number of users and the volume of requests, and it's calculated during discovery, not from a generic price table.

### Which open-source model do we choose?

It depends on the task, the language and the latency you can accept — there's no single best universal model. We work with open-weight model families such as Qwen, DeepSeek, GLM, Gemma or Mistral, served through production engines such as vLLM or SGLang, and we pick the right option during discovery, based on tests against your real use case.

### Does the data stay in Romania/the EU?

Yes — the platform runs either on the company's own servers (on-premise, physically where the company is) or on a dedicated GPU at a provider in the European Union. Data doesn't leave the agreed perimeter, either to an external API or outside the EU — relevant for the data residency GDPR requires and for NIS2 in regulated sectors.

### Where can an AI server for a company in Romania run?

Either on the company's own servers, on-premise, or on dedicated GPUs rented from a provider in the European Union. The right option depends on how many employees use it at the same time, how sensitive the data is and who manages the infrastructure, and the sizing is done in discovery, on the company's own numbers.

### Let's see what can be automated in your business.

A free 30-minute session: we'll tell you what can be automated, how long it takes and what it costs, with a fixed price after discovery.

We reply the same business day.
