AI INFRASTRUCTURE

Private AI infrastructure: on‑premise or in the EU cloud

An external API is simple to get started with, but contracts, access codes and customer data end up on servers outside the company. We build AI platforms that run on-premise or on a dedicated GPU in the EU, with data kept inside the perimeter the company sets.

EUthe infrastructure runs within the company's perimeter or on a dedicated GPU in the European Union, not on a US API
open-weightopen-weight model families (Qwen, DeepSeek, GLM, Gemma, Mistral), not a model locked inside an API
the company'sthe code and infrastructure remain the client's property at handover, not the provider's
AI INFRASTRUCTURE
AI INFRASTRUCTURE
01

What self-hosted AI infrastructure means for a company

Self-hosted AI infrastructure means the language model runs on servers the company controls, or on a dedicated GPU at an EU provider — not on a public American API that you send contracts, access codes or customer data to, request by request. The difference isn't cosmetic: the data never leaves the agreed perimeter again.

We're not just talking about “running a model”. A private AI platform means an interface for employees, role-based permissions, connections to internal documents and systems, audit logs and monitoring — everything a generic AI chat subscription lacks.

Private AI infrastructure, explained in two and a half minutes
From dozens of isolated chats on US servers to agents that work inside the company's systems, on an open-source model running on GPUs in the European Union.
02

The platform isn't just a model: interface, RAG, agents, audit

A private LLM, on its own, doesn't solve anything for a company. The value appears when the model is connected to an interface employees actually use, to the company's documents and, where it makes sense, to existing operational systems.

We build each component as part of the same platform, not as separate tools manually stitched together. The part that makes the platform useful every day, connecting agents to your ERP through MCP servers, has its own separate page.

INTERNAL CHAT

Interface with roles, permissions and SSO

Every employee signs in with the company account; what they can see and request depends on their role, not on good judgment.

RAG

Search across the company's documents

Answers come from internal procedures, contracts and reports, not from the model's general knowledge.

AGENTS + ERP/CRM

Connectors via API or custom-built MCP

Agents that read and, with explicit approval, write into existing systems — ERP, CRM, SAGA, SAP.

AUDIT & MONITORING

Usage and access logs

Who asked what, and what the model answered — visible internally, not hidden behind an external API.

EVALUATION

Evaluation sets for each use case

We measure accuracy on the company's real data, not on a generic public benchmark.

03

How big does an AI server actually need to be

Sizing isn't based on “the biggest model available”, but on the use case: how many employees use the system at the same time, how long the processed context is (whole documents or short questions), and the size and quantization of the chosen model — these variables decide the video memory needed.

A very large MoE (mixture-of-experts) model can need hundreds of GB of video memory even when quantized; a 20-30 billion-parameter model, by contrast, runs comfortably on a single high-memory GPU. For most internal processes, the right option is the second one — not the biggest model, but the one that fits the task. For a first test on a single computer, we've described separately how to run a local LLM and how much memory it needs.

04

Private LLM: on-premise at your site or on a dedicated GPU in the EU

There are two ways to run a private LLM, not just one. On your own servers (on-premise), the company physically controls the infrastructure — the right option where internal or contractual policy explicitly requires it. On a dedicated GPU at a provider in the European Union, the company doesn't buy hardware, but the data still doesn't leave the EU or reach an American provider.

The choice also depends on the data residency required by GDPR, the transparency obligations under the AI Act and, for regulated sectors, NIS2 requirements — factual topics to raise with the company's legal team, not an area where we provide legal advice. If the terms get mixed up in internal discussions, we've explained separately the difference between self-hosted AI and sovereign AI. The same infrastructure can run an AI assistant on the company's documents, with answers that cite their source.

05

Self-hosted or API: where's the break-even point

There's no universal answer. For few users and bursty traffic, pay-per-token on a hosted API is almost certainly cheaper than running your own server — the infrastructure would sit idle outside usage peaks.

For several dozen daily users, constant traffic and sensitive data, self-hosted starts to win out — both on cost at high volume and on control over the data. The exact threshold depends on the number of users and the real volume of requests; we calculate it during discovery, using the company's own numbers, not a generic table.

Self-hosted or API? 5 questions before you decide

Check what's true for you.

0 of 5

Check the boxes above.

06

From discovery to infrastructure handed over to the company

We start with discovery: use cases, what data exists and how clean it is, security requirements, a sizing estimate. A pilot in a single department follows, with acceptance criteria written in advance — exactly as with any serious AI implementation.

After the pilot, we scale into production with active monitoring and, at the end, hand over the code and infrastructure to the company. They stay yours — we don't operate a closed platform you depend on indefinitely.

07

A fixed price, after a short discovery call.

AI infrastructure discovery
on request · after discoveryOne day of analysis: use cases, available data, security requirements and a sizing estimate
  • use case and data audit
  • security and compliance requirements
  • sizing recommendation
Pilot in one department
on request · after discoveryThe chosen model runs in a single department, on on-premise or EU infrastructure, with acceptance criteria written in advance
  • model + runtime selection (vLLM/SGLang)
  • RAG on the department's documents
  • written acceptance criteria
Rollout + handover
on request · after discoveryWe roll out to more departments with usage monitoring, then hand over the code and infrastructure to the company
  • multi-department rollout
  • monitoring and audit logs
  • code + infrastructure handover

Every project has a different context, workflows and infrastructure, so the price is set after a short, paid discovery and does not change along the way.

FAQ

Frequently asked questions

How big does the server need to be?

It depends on three things: how many employees use it at the same time, how long the processed context is, and which model you choose. A very large MoE-type model can need hundreds of GB of video memory even when quantized; a 20-30 billion-parameter model runs on a single high-memory GPU — the right fit for most internal processes. The exact sizing is done during discovery, based on your use case.

Can I use local AI with data from my ERP?

Yes — agents connect to your ERP/CRM via API or custom-built MCP connectors, read the data they need and, with explicit approval, can write back into the system. The whole flow stays on your chosen self-hosted or EU infrastructure, without ERP data ever passing through an external API.

Is it cheaper than ChatGPT Team/Enterprise?

There's no universal answer. For few users and bursty traffic, a subscription or pay-per-token usually stays cheaper. For several dozen daily users, constant traffic and sensitive data, self-hosted starts to win out — the exact threshold depends on the number of users and the volume of requests, and it's calculated during discovery, not from a generic price table.

Which open-source model do we choose?

It depends on the task, the language and the latency you can accept — there's no single best universal model. We work with open-weight model families such as Qwen, DeepSeek, GLM, Gemma or Mistral, served through production engines such as vLLM or SGLang, and we pick the right option during discovery, based on tests against your real use case.

Does the data stay in Romania/the EU?

Yes — the platform runs either on the company's own servers (on-premise, physically where the company is) or on a dedicated GPU at a provider in the European Union. Data doesn't leave the agreed perimeter, either to an external API or outside the EU — relevant for the data residency GDPR requires and for NIS2 in regulated sectors.

Where can an AI server for a company in Romania run?

Either on the company's own servers, on-premise, or on dedicated GPUs rented from a provider in the European Union. The right option depends on how many employees use it at the same time, how sensitive the data is and who manages the infrastructure, and the sizing is done in discovery, on the company's own numbers.

The Niche Society
The Niche Society TeamAI and software engineers from Bucharest · LinkedIn
updated 17 Sep 2026

Let's see what can be automated in your business.

A free 30-minute session: we'll tell you what can be automated, how long it takes and what it costs, with a fixed price after discovery.

Book a free sessionoffice@thenichesociety.ro

We reply the same business day.

+40 733 045 833