How it works

What actually gets installed

The non-technical version first, then the detail your IT team will want.

In plain language

One computer, in your office, doing the whole job

I install a private AI assistant on a computer inside your building. Your team opens a browser, logs in with the account they already have, and types — the same experience as the AI tools they're used to.

The difference is invisible to them and decisive for you: the model runs on that machine. The question never gets sent anywhere. The answer is generated in the room. There is no subscription, no per-message charge, and no third party with a copy of what your team asked.

The test that settles it: disconnect the building from the internet and ask it a question. It answers. Nothing was ever going out.
The stack

Every layer, and what it does

All open source, all running on your hardware. No component requires an account with anyone, including me.

LayerWhat it doesWhat I typically use
Inference engineRuns the AI model on your hardwareOllama, llama.cpp, vLLM
Language modelsThe actual intelligenceLlama 3.x, Mistral, Phi-4, Gemma 2, Qwen 2.5
Chat interfaceWhat your team sees in the browserOpen WebUI, AnythingLLM, LibreChat
Document searchAnswers grounded in your own filesOpen WebUI RAG, Chroma, Qdrant
Access controlSingle sign-on and per-user permissionsAuthelia, Keycloak, LLDAP
Reverse proxyRouting, TLS, and the auth gateNginx, Traefik, Caddy
AdministrationFile management and system healthFile Browser, Cockpit, Portainer

The specific choices depend on your environment. Nothing here is a product I resell — I don't take vendor commissions.

The hardware question

You provide the machine — that's what makes it yours. Many organizations already own something suitable and don't realize it.

  • Small team (5–15 people): a workstation with a modern NVIDIA GPU handles it comfortably
  • Mid-size (15–50): a dedicated server, ideally with 24GB or more of GPU memory
  • Compact models in the 7–8 billion parameter range run well on about 8GB of GPU memory
  • CPU-only is possible for light use — slower, but it works and costs nothing extra

The audit tells you specifically what you need, including whether hardware you already own will do the job.

The timeline

Call — 30 minutes

Same week, usually within a couple of days.

Audit — about a week

A site visit or remote session, then a written assessment.

Deployment — days

Most builds are a matter of days on site, not months. If hardware needs ordering, that's usually the long pole.

Training — one session

A working session with your team once it's live.

Being precise

What "nothing leaves your building" does and doesn't mean

I'd rather over-explain this than have you discover a caveat later.

What's genuinely true

  • Prompts and documents are processed only on your hardware
  • No AI vendor receives, stores, or trains on your data
  • The system works with the internet fully disconnected
  • No per-user or per-message fees, ever
  • You own the hardware, the models, and the data outright

The honest caveats

  • Model updates require an internet connection to download — on your schedule, not automatically
  • Remote support, if you want it, means giving me access when you ask
  • This secures the AI layer; it doesn't fix unrelated gaps elsewhere in your environment
  • Open models are strong but not identical to the frontier ones
  • Someone has to maintain it — you, your IT team, or me on retainer
Running proof, not a pitch deck. I maintain this exact architecture on my own infrastructure — local models, single sign-on, reverse proxy, document search, the whole stack, running continuously. When I describe a failure mode on a call, it's because I've already hit it and fixed it on my own hardware first.

Start with a conversation, not a contract

Thirty minutes, no cost, no obligation. It's entirely possible the answer is that you don't need anything from me yet — and I'll tell you that.

Book a Free 30-Minute Call