THE MACHINE IN DETAIL

One appliance. Three layers. One vendor.

The zenpAI appliance combines hardware, a local model, and agent orchestration into a controllable AI system on your own infrastructure. On this page we show the components in detail: security architecture, stack, hardware classes, and operations.

§01

Security as architecture.

NO ENDPOINT · NO FOREIGN JURISDICTION

Security here is not an add-on feature, but part of the architecture: no shared cloud hardware, no external inference, no third-party administrators on your data. Your data stays where it belongs — with you.

GDPR DSGVO Meets the strict GDPR & EU AI Act requirements — no US CLOUD Act, no third country.
01

0% cloud exposure

All inference, all model weights, all training pipelines remain on your hardware. No sub-processor, no third country.

02

Local audit log

Every login, every tool call, and every relevant response is logged locally. This creates traceability for internal controls, compliance processes, and AI Act documentation.

§02

The stack.

FROM APPLICATION TO HARDWARE

Seven layers, one responsibility. Proven datacenter software, not a hobby rig: a hardened Linux with optimized drivers, Kubernetes on GitOps principles above it — versioned and auditable — and at the top the agent layer that orchestrates — delegating tasks, keeping state, controlling flow — while setting the governance.

§02 · B — ARCHITECTURE

What actually runs in the rack — from the application down to the hardware.

APPLICATIONS User interfaces people work with
OpenWebUIStreamlitJetBrains JunieMCP-Clients
AGENT LAYER Orchestrates processes and keeps context
LangGraphModel Context ProtocolOAuth / LDAP / SSOlocal audit log
INFERENCE & MODEL LAYER Runs local AI models efficiently on the GPU
vLLMllama.cppTensorRT-LLMSGLangLlamaQwenDeepSeekGLMKimiMistralLoRA-Pipeline
PLATFORM Deployment, updates, and operation
KubernetesArgoCDHelmGitOpsHardened LinuxNVIDIA drivers / CUDAcontainerd
HARDWARE GPU server in your rack
NVIDIA Enterprise-GPUDELL PowerEdgeAMD EPYCredundant PSU
§03

Hardware classes.

ONE SIZING MODEL · VRAM · TOPS · USERS · kW

Three build sizes, matched to your load. Standard configuration based on the current Blackwell generation; the final GPU, VRAM, and PSU redundancy are set jointly after your load and latency profile.

SFOUNDATION / PILOT
GPU
1× NVIDIA RTX PRO 6000 · Blackwell
VRAM
96 GB
AI performance
~4,000 TOPS
Power max
~1 kW
Users
1–10
MPROFESSIONAL TEAM
GPU
4× NVIDIA RTX PRO 6000 · Blackwell
VRAM
384 GB
AI performance
~16,000 TOPS
Power max
~3 kW
Users
10–60
LCUSTOM ENTERPRISE
GPU
8× NVIDIA RTX PRO 6000 · Blackwell
VRAM
768 GB
AI performance
~32,000 TOPS
Power max
~5.5 kW
Users
60–200
IndividualMAXIMUM SCALE
GPU
Individually configured
VRAM
as required
AI performance
as required
Power max
individual
Users
200+

User counts depend on model and load. The appliance grows with you: it can be scaled up at any time — we add GPU capacity as your needs grow. On request.

§04

How we deliver.

FROM FIRST CALL TO PRODUCTION · IN PHASES

Anyone can buy hardware — the real work is delivery: integration, data remediation, agentic workflows. We take it off your hands, from process assessment to ongoing operations — in clearly defined phases, billed per phase — the timeline varies by customer and scope. If the first phase shows it is too early, we say so, before any hardware is ordered.

BEFORE KICK-OFF

Consulting & pilot

Assessment: where does it hurt, which data, which tools? We pick the first use case and show a concrete result fast — with the honest option to stop afterwards.

PHASE 1

Analysis & data

Map data paths, define the goals. Most of it is data remediation — making scattered records production-ready.

PHASE 2

Build & fine-tuning

The appliance is racked and connected. The model is fine-tuned on your domain and quantized for your GPU.

PHASE 3

Orchestration & handover

Agent layer set onto your process, pilot operation, training for your IT, full handover. You take over, we stay on call — no lock-in.

Sounds like your project? Book an intro call

§05

Operations after go-live.

OPERATIONS · OPTIONAL

An on-premise AI has to be operated — monitored, updated, maintained. By default we hand everything over to you so your own IT can run it. But if you prefer, we take operations on entirely — remotely, for a fixed monthly price. No token fees, no usage billing, no surprises.

INCLUDED IN THE SERVICE
  • Continuous 24/7 monitoring — GPU, memory, latencies, system status
  • Security and system updates for the hardened Linux stack
  • Current, curated models — new generations are continuously evaluated and rolled out
  • Remote support and error analysis
  • Hardware early warning and upgrade recommendations
ADD-ONS · PROJECT-BASED
  • Re-training and RAG optimization on your own data
  • Custom development and integration
  • On-site visits by arrangement

Not part of the monthly price: the hardware as a one-time purchase (see §04) and — for larger configurations — cooling. Monthly price on request — depending on hardware class and the response time you need.

One appliance, one vendor — let us talk.

Specific hardware class, model family, and integration depth are not set by a configurator but together with you — once we understand your data, systems, and first use case.

Book intro call →
Book an intro call →