Infrastructure · Live

Local AI Server

AI in your own building. Not a byte leaves your network.

A preconfigured device runs the models locally and serves the applications inside your own network. Access happens from the building or through an encrypted tunnel, without a single port opened on the router. Setup, maintenance and updates are handled by us.

scroll
Models
Local, kept warm
Access
From the network only
Maintenance
Via secure tunnel
Port forwarding
None required
Where it breaks down

The data AI would help with most is exactly the data that may not go to the cloud.

Patient records, client files, design data, personnel matters. That is where the biggest leverage sits, and that is where the hurdle is highest.

01

Approval takes months

Processing agreements, impact assessments, works council, professional confidentiality. By the time that is settled, the project has lost its momentum.

02

Shadow AI appears anyway

When the official route takes too long, people use private accounts. Then exactly the data that was meant to be protected flows out, only without any control at all.

03

Running it yourself looks too complex

Models, containers, networking, hardening, updates. Without a team of your own that looks like a project you would rather not start. So it never starts.

The solution

A local AI server that arrives, gets connected and runs.

The local AI server is not a set of build instructions, it is a finished appliance. It stands in your building, it belongs to your network, and it is looked after by us. As demand grows, the setup grows with it.

On the local AI server a model engine runs the language models locally and keeps the frequently used ones permanently in memory, so the answer comes immediately. In front of it sits an access gateway through which every application is reachable under a single address instead of a list of ports.

We rely on open-source models with licences that expressly permit commercial use. That is not a compromise: for the tasks that come up day to day (summarising, structuring, rephrasing, classifying, dictating) they are entirely sufficient, and on simple tasks small models are noticeably faster than large ones. Where a task genuinely needs a large model, we say so in advance instead of selling local as equivalent everywhere.

Every application runs in its own container, alongside a local database for users, roles and content. New applications get added without the existing ones being touched. Speech recognition for dictation runs locally too, so even the spoken word does not leave the building.

The setup scales in both directions. If one device is no longer enough, a second one joins it and shares the load. For sites with high availability requirements a backup server stands ready and steps in without anyone intervening. And several sites can be linked into a network in which each keeps its own data.

Support runs through an encrypted point-to-point tunnel. It needs no port forwarding on the router and grants no access to other devices in the network. Added to that is the hardening: firewall, password login switched off, network isolation, separate roles for operation and application. Setup runs through a repeatable script that only repairs what is missing, and can therefore run as often as you like.

Open-source modelsModel engineAccess gatewayContainer per applicationLocal databaseSpeech recognitionBackup serverSecure tunnel
In practice

Three tasks that otherwise fail at the approval stage.

Dictate instead of typing

Findings, notes and minutes are spoken and turned into text locally. Neither the recording nor the result leaves the building.

Analyse your own file holdings

Documents are read, structured and made searchable, on your own device. Including files protected by professional confidentiality.

Questions to your own knowledge

The specialist question goes to your own knowledge base. Not only the answer stays in the building, the question does too, and the question often reveals more than the answer.

How it works

Size it, set it up, configure it, look after it.

01

Size it

The number of workplaces, the planned applications and the model sizes determine the hardware. Too small slows you down, too large costs without benefit.

02

Set it up

The device arrives prepared, gets connected to the network and receives a fixed address. No port forwarding, no change to the outward-facing firewall.

03

Configure it

A repeatable script installs every component: model engine, gateway, database, containers, speech recognition and the hardening.

04

Look after it

Updates, model care and monitoring run through the encrypted tunnel. You can see the state at any time in the interface.

Let us work out in a conversation which of your use cases hold up locally and which do not.

Arrange a conversation
Core capabilities

What the module brings.

Preconfigured and expandable

Hardware, operating system and every component tuned to each other. You plug in, we have prepared. If one device is no longer enough, a second joins it and shares the load.

Local model engine

Open-source language models run on the device graphics unit. Frequently used ones stay in memory so the first answer does not wait for loading. Small models for quick tasks, larger ones for the demanding ones.

Central access gateway

Every application under one address instead of many ports. At the same time the place for logging and later access control.

Container per application

Every application isolated. New ones get added without touching the existing ones and can be rolled back individually.

Local database

Users, roles, configuration and application data in one place in the building, with a secured backup copy.

Local speech recognition

Dictation is turned into text on the device. Neither the recording nor the transcript leaves the network.

Encrypted remote access

A point-to-point tunnel for maintenance and mobile work. Without port forwarding on the router, without access to other devices in the network.

Network isolation and hardening

Firewall, password login disabled, telemetry switched off, separate roles for operation and application. The server is not reachable from outside.

Repeatable setup

One script installs everything and repairs only what is missing. That makes the state reproducible at any time, even years later.

Monitoring interface

Load, memory, running applications and loaded models at a glance, reachable inside your own network.

Connecting the modules

Corporate Memory, Content Creator and the other modules run on the same device, with the same roles as in the cloud.

File share in the network

One storage location in the building through which documents reach local processing, without a detour via an external service.

Where it works

Three environments where it otherwise does not work at all.

Practice and clinic

Patient data is subject to confidentiality. On the local device findings can be dictated and records analysed without a processing agreement becoming necessary.

Law firm and consultancy

Client files must not leave the building. Research and summarising run locally, and the question itself stays confidential.

Industry and development

Design data and formulations are the capital of the company. Analysis inside your own network rules out leakage technically rather than contractually.

A look inside

Two views of the same device.

On the left the administration, which is responsible for the state. On the right the user, who simply wants to work and should not notice the server in the basement at all. The tabs are switchable.

Load38 %
Memory41 / 96 GB
Uptime84 days
Load over 7 days · threshold of 85 per cent not reached
Last update applied 6 days ago, without interrupting operation.
Warm General purpose model, medium in memory
Warm Dictation and transcription in memory
On demand Document understanding, large loads in 9 s
On demand Image understanding rarely used
Models kept warm answer immediately. Which ones those are is decided from actual usage.
Backup today at 03:00 · complete
Backup server ready · last sync 04:12
Restore last tested 21 days ago
Restoring is rehearsed regularly. A backup that has never been restored is not one.
Access From the local network only active
Maintenance Encrypted tunnel active
Router No inbound port forwarding verified
Login Password login disabled keys only

The state of the server at a glance

Load, loaded models, running applications and the backup. Reachable only from your own network or through the encrypted tunnel.

Knowledge Questions to your own holdings open
Dictation Speak a finding or a note open
Documents Analyse a file open
Inbox Drafts for replies new
No account with a provider, no agreeing to terms of use. The login is the one from the building.
Recording 2:14 min
Recognition local
Structure done
Recording and text stay on the device in the building. There is no upload you could forget to revoke.
Response time0.9 s
Sources3
Data leavingnone
Question Which rule applies to short-time work?
Answer With a source from your own holdings 0.9 s
Evidence Works agreement 2024 · page 4
Note The question does not leave the building
Not only the answer stays in the building, the question does too. And it often reveals more about a plan than the answer.
Pages today1,240
Recognitionlocal
Storagein house
Incoming Scanned file folder, 340 pages in progress
Recognised Text, tables and stamps done
Assigned Case, date, parties
Open 6 pages illegible for visual check
Illegible pages get reported rather than guessed. What was not recognised does not appear in any analysis either.
New today38
Drafts31
Approved24
Request Appointment change, standard case draft ready SendAmend
Request Documents requested draft ready SendAmend
Request Complaint, personal case no draft SendAmend
Delicate cases deliberately get no draft. The server recognises them by tone and puts them in front of a person.
Practice Dictate a finding and file it structured
Law firm Check briefs against the file holdings
Industry Analyse and compare test protocols
People Pre-sort applications, with no data leaving
Engineering Turn maintenance reports into metrics
Each of these applications runs in its own container. A new one gets added without the existing ones being touched.

For the user it is simply an address in the building

Every application under one address, no port knowledge needed, no cloud sign-in. Whoever is in the building is in.

Abstracted representation with invented values. Which applications, models and access paths run in your installation is something we define together.

Where it sits

The foundation everything else can run on.

On the home page this module sits under “Operation and sovereignty”, and that is exactly its role. It is not another tool, it is the place where the other tools can run when the cloud is not an option.

Culture
Enable people, do not replace them.
Agents
Automate processes, create room to work.
Knowledge
Secure experience and make it usable.
Content
Scale communication, on brand and on quality.
Modules that can run here
Corporate MemoryKnowledge with a source, locally
Content CreatorContent with no data leaving
Agent-as-a-ServiceAgents inside your own network

Not every task belongs on a local device. Where larger models are needed and the data permits it, hybrid operation is the more honest route. In the conversation we say which of your use cases hold up locally and which do not.

Operation and data sovereignty

Data sovereignty as an architectural principle, not an add-on.

For this module data sovereignty is not one operating variant among several. It is the entire purpose.

One site

One device in the building, every application on it. The normal case for practices, law firms and single premises.

Several sites

One device per site, central support through the encrypted tunnel. Data stays where it arises.

Hybrid

Sensitive processing locally, compute-heavy tasks in European data centres. You draw the line, and it stays visible in operation.

One quote says more than many promises
“The decisive question in the data protection review was not which model we use, but where the data goes. The answer was nowhere, and with that the topic was settled in twenty minutes instead of four months.”

Medical Director, Specialist practice with four sites

Pricing

Setup, support and hardware kept transparently apart.

Tiered by workplaces, model sizes and the number of applications. Setup covers sizing, installation, hardening and instruction.

PackageFor whomScope and hardware *Setup *Operation *
PracticePractice, law firm, small business1 device · up to 15 workplaces · 1 to 3 applications · plus hardware €5,000 to €8,000from €4,900€190 / mo.
BusinessMid-sized siteup to 50 workplaces · larger models · up to 3 applications · plus hardware €15,000 to €20,000from €9,900€390 / mo.
ProSeveral sites, high availabilitysecond device for failover · custom applications · plus hardware €30,000 to €50,000from €18,500€690 / mo.
CorporateRollout across many sites, connection to an existing directory, or particular requirements for evidence and certification: on request.

Also available for rent

Anyone who does not want to buy the hardware can rent it. Device, setup, maintenance and model care then run in one monthly rate, the device stays in your building and so does the data. Sensible when an investment would break the budget or when the technology should be renewed on a predictable cycle. We will put together an individual offer for this at any time.

Hardware is passed on transparently at the purchase price plus a procurement fee, or supplied by you. The ranges above are experience values and depend on memory and compute. Operation includes maintenance, updates, model care and remote monitoring. We earn on setup and support, not on a margin on the device.

* Guide values. Scope, hardware and prices are adapted to the individual case in the quotation, because the number of workplaces, the model sizes and the availability requirement drive the effort.

All prices are net, plus statutory VAT.

Frequent questions

What decision makers ask first.

How capable is this compared to the cloud?

For summarising, structuring, dictation and analysing your own holdings, locally run models are clearly sufficient today. With very large models and very long contexts the cloud stays ahead. We say in advance which of your use cases fall into which category, instead of selling local as equivalent everywhere.

What happens if the device fails?

Then the local applications stand still, as with any server in the building. That is why a backup and an agreed restore path are part of operation. Where an outage is not acceptable, the Pro package provides a second device.

Who maintains the device?

We do, through the encrypted tunnel. Updates, model care and monitoring are part of operation. Nobody is needed on site except in the case of a hardware defect.

Does this require an opening in the firewall?

No. The maintenance tunnel is established by the device outwards, not from outside inwards. No inbound port forwarding is needed, and the tunnel grants no access to other devices in your network.

Can we move to the cloud later, or the other way round?

Yes. The modules are the same regardless of where they run. A move is a migration task for data and configuration, not a new project.

The first step

Your use cases. An honest sizing.

Tell us what you want to do with AI and which data is involved. We will tell you which cases hold up locally, which hardware that needs and where the cloud stays the better route.