
Services · AI & Machine Learning
Private AI Deployment
Open-weight models inside your VPC or on-premise
What this is
What private ai deployment covers
Private AI deployment for regulated organisations whose data, prompts, and outputs cannot leave a defined boundary. We size and stand up an inference stack on your GPUs or in your cloud account, put a model gateway in front of it with logging and quotas, and connect the agents and retrieval systems that need it. The work includes the math on when on-premise beats API pricing and when it does not.
Typical timeline: 4–10 weeks
Who hires us for this
- Banks, insurers, and healthcare providers with residency or BAA constraints
- Government and defence programmes with air-gapped networks
- Manufacturers and utilities with OT data that never touches the public internet
Problems we get hired for
What usually goes wrong before we are called
Legal blocked the API vendor
We deploy Llama 3, Qwen, or Mistral class models inside your boundary with the same gateway interface, so application teams do not rewrite integrations.
Nobody knows if on-premise is cheaper
We model tokens per day, latency targets, and GPU utilisation against API pricing. Below a threshold the answer is a private cloud endpoint, not a rack.
Shadow AI usage is untracked
A model gateway with per-team keys, quotas, prompt logging, and PII redaction gives security a single place to see and control usage.
How the engagement runs
Steps and the artifact each one produces
- 01
Sizing
Workload model: tokens, concurrency, latency, and growth. Compare on-premise, private cloud, and API options.
Artifact: TCO memo with break-even
- 02
Stand-up
Inference servers, model gateway, monitoring, and access controls in your environment.
Artifact: Running stack with runbooks
- 03
Connect
Agents and RAG systems pointed at the gateway. Evals re-run on the private models.
Artifact: Eval comparison and migration notes
- 04
Operate
Capacity alerts, model update process, and quarterly re-check of the TCO assumptions.
Artifact: Ops handbook and refresh schedule
Stack
Tools we reach for
Chosen per engagement against your existing platform. Nothing here is mandatory; everything here has shipped.
| Inference | vLLM · TGI · llama.cpp for CPU edge · TensorRT-LLM |
|---|---|
| Models | Llama 3 · Qwen 2.5 · Mistral · BGE embeddings |
| Gateway | LiteLLM · Kong or Envoy · Per-team API keys and quotas |
| Platform | Kubernetes · NVIDIA GPU Operator · Terraform · Prometheus and Grafana |
| Controls | Presidio PII redaction · Prompt and output logging · Network policies for air-gap |
What you receive
Deliverables
- TCO memo with break-even against API pricing
- Inference stack and model gateway in your environment
- Access controls, quotas, and logging
- Eval comparison between API and private models
- Ops handbook and model update process
Proof
Where this has run
Banking
Durable AI Ops Layer for Mid-Market Banking
KYC exceptions and payment repairs on Temporal + Vercel AI SDK + MCP: maker-checker Signals, PII redaction, zero double-posts on core retries.
Payment exception handle time: 45–90 min median → 12–25 min median
Legal
Case Law Research Assistant for Mid-Market Law Firm
Hybrid RAG with citations: research time per matter from 4–6 hours down to ~45–90 minutes.
Research time per matter: 4-6 hours → ~45-90 minutes
Products that ship inside this service
- Xenith Private AI · Open-weight LLM stack inside your VPC or air-gapped network.
- Xenith LLM Platform · Model routing, caching, and cost controls between your apps and every LLM provider.
Engagement models
How we can be hired for this
| Model | Shape | Typical duration |
|---|---|---|
| Fixed-scope pilot | Defined workflow, defined integration points, defined exit metric. Priced as a project. | 4–12 weeks |
| Production retainer | Monthly capacity for evals, monitoring, model refreshes, and new workflows. Month-to-month. | 3+ months |
Pricing and phase gates are described on How we work.
Questions buyers ask
Frequently asked questions
- When does on-premise beat API pricing?
- In our sizing work the break-even has landed around sustained high daily token volumes with steady load; below that, a private endpoint in your cloud account wins. The on-premise math post shows the model with the assumptions you can change.
- Are open-weight models good enough?
- For classification, extraction, drafting, and retrieval-grounded answers, Llama 3 and Qwen class models pass most of the eval sets we have run. For hard reasoning tasks we keep a frontier model available through the same gateway under a stricter data policy.
- Can this run air-gapped?
- Yes. Model weights, containers, and dependencies are delivered as an offline bundle; updates follow your change window. The stack has no outbound calls.
- Who owns the certifications?
- You do. We design the data flows, access controls, and audit logs to fit HIPAA, SOC 2, or GDPR requirements and hand over the evidence; your auditors own the certification.
When this is the wrong fit
- Teams with no data-residency or compliance driver; API models with a data-processing agreement are cheaper.
- Organisations without anyone to operate GPUs or Kubernetes after handover, unless a retainer is in place.
Next step
Send the workflow. We reply with an architecture sketch.
Describe the queue, the system of record, and who signs off today. First brief is free; NDA first if you prefer.