The most expensive AI failure we see in industrial and fintech programs is not a model hallucination on a slide. It is a language model that quietly owns a number someone will later treat as truth: a billed weight, a demurrage hour count, a risk grade, a valve trip recommendation.
We call the corrective pattern Zero-LLM-Trust. Numbers that move money or safety stay in rules engines, sensors, and signed human approval. LLMs summarize, extract, draft, and explain. They do not invent figures or execute irreversible actions alone.
That sounds strict. It is. After enough control-room and compliance reviews, we stopped treating "the model is usually right" as an architecture.
What Zero-LLM-Trust Actually Means
Zero-LLM-Trust is not "no LLMs." It is a hard split between conversation and commitment.
On the conversation side, models are useful: they turn a chute jam into a readable diagnosis, draft a demurrage dispute narrative, or explain why a spatial near-miss fired. On the commitment side—rate-card corrections, rake free-time clocks, forklift routes, publish to air, credit decisions—execution paths never call the model as the source of the figure or the sole authorizer of the write.
In practice that means three constraints we build into every industrial and regulated agent design:
- Deterministic owners for facts. Weight comes from DWS. Free time comes from tariff rules and timestamped events. Policy version comes from a registry. The model may cite those facts; it may not mint them.
- Tool-grounded answers only for ops Q&A. If the copilot cannot point at a signal, a row, or a rule ID, it says so. We prefer a short "I don't have that reading" over a fluent guess.
- Human approval before money, safety, or publish. Assist mode proposes. Operators and supervisors commit. Audit trails record who approved what, with what evidence.
This is the same doctrine we document under governed AI and ship across industrial AI and media workflows: autonomy is earned, not assumed.
Why Industrial and Fintech Fail the Same Way
Sortation hubs and payment rails look unrelated until you watch the failure mode.
A CEP hub loses 2–4% of revenue when volumetric DWS weight and WMS billed weight diverge and nobody closes the loop with a governed correction. A plant rail yard accrues demurrage while status still lives in radio and logbooks. A warehouse CCTV stack lights up dashboards while nobody owns the next physical action. A lending agent drafts a decision memo that embeds a score the policy engine never computed.
In each case the organization bought "AI" as a brain and skipped the boring question: who owns the number, and who owns the write?
Fintech teams already know this instinct for core banking and AML. Industrial teams learn it the first time a model-suggested chute change hits a live sorter without a supervisor gate. The domains converge on the same architecture: systems of record stay authoritative; AI sits above them; irreversible paths are gated.
If you are mapping this onto logistics or manufacturing stacks, start from the industry pages for logistics and supply chain and manufacturing—then insist every vendor diagram show where the LLM stops.
The Boundary We Enforce in Production Patterns
When we design with Xenith Agents, we treat the agent as a coordinator of tools and drafts, not as a ledger. When we wrap high-risk decisions with Xenith Seal, we want the decision UUID, policy version, and hash chain to survive a regulator's "prove this one." Neither product assumes the model is trustworthy with money.
Concrete split we use in workshops:
| Decision class | Owner | LLM role |
|---|---|---|
| Parcel weight / billable dimensions | DWS + rate card engine | Explain mismatch; draft correction proposal |
| Demurrage clock / free time | Rules + immutable event timestamps | Draft dispute dossier language |
| Safety trip / intrusion siren | Sensors + HITL policy | Summarize evidence; never auto-siren alone |
| Credit / AML gate | Policy engine + human escalation | Narrative for reviewers |
| Newsroom publish / CMS push | Editorial SoR + approve | Draft packages; wait for desk |
That table is the product. Everything else—model choice, prompt style, agent framework—is secondary.
How Zero-LLM-Trust Shows Up in Xenith Wedges
We did not invent the doctrine as a whitepaper and then hunt for products. The wedges force it:
- Xenith Sort sits above WCS and DWS. Revenue leakage proposals compare sensor and billing facts. The LLM never calls execute APIs directly; Temporal or an inline approval core owns writes after RBAC.
- Xenith Yard keeps detention clocks in deterministic rules. Models help commercial teams narrate dossiers; they do not own the clock.
- Xenith Twin and Xenith Spatial keep OPC-UA/MQTT state and CV detections as the live truth. Operator writes are signed. Optimizers propose; they do not silently route equipment.
- Xenith Desk keeps the broadcast desk as system of record. Multi-format drafts accelerate output; air and CMS still require human approve.
Across all of them the phased path is the same: Connect → Shadow → Assist → Scale. Shadow mode is where trust is measured before Assist is allowed to touch a write path. We expand that pattern in a companion post on shadow mode before assist.
Failure Modes We Refuse to Ship
Fluent billing corrections without evidence. If the proposal cannot show DWS reading, WMS row, and rate-card line, it is not a correction—it is a story.
Safety "automation" that skips HITL. Edge CV that pages a supervisor is useful. Edge CV that sounds a plant siren because a model "is confident" is a liability.
Agent frameworks with direct execute tools on money paths. Giving an LLM a post_billing_adjustment tool and hoping temperature-zero saves you is not governance.
Proof theater. Screenshots of chat logs are not audit. If you need non-repudiation, you need sealed decision records—Seal-shaped—not a Slack export.
Rip-and-replace as the sales motion. Replacing WCS, MES, or the newsroom desk to "put AI in the middle" fails politically and operationally. Integrate-don't-replace is the durable path; we unpack it in integrate, don't replace the system of record.
A Practical Checklist for Buyers
Use this in RFPs and architecture reviews. If a vendor cannot answer clearly, treat the gap as intentional risk.
- Name the system of record for every money and safety fact. If the answer is "the model," stop.
- Show the approval inbox for physical and financial actions—roles, not vibes.
- Require append-only audit with who approved, what evidence, which policy version.
- Demand Shadow Mode with logged recommendations and measured false-positive rates before Assist.
- Separate summarize tools from execute tools at the API boundary. No shared "do_anything" agent toolbelt on irreversible paths.
- Label synthetic demo metrics. Demo leakage euros and throughput figures are workshop labels until your site produces them.
- Ask what happens when the LLM is wrong. The correct answer is "nothing irreversible commits."
Teams that pass this checklist still use LLMs heavily. They just refuse to let a probabilistic system own deterministic commitments.
What Changes When You Adopt the Doctrine
Operators stop treating the copilot as a mysterious oracle and start treating it as a grounded assistant. Supervisors get proposals with citations instead of dashboards that only describe the past. Commercial and compliance teams get dossiers and decision UUIDs instead of "the AI said so." Engineering stops arguing about which frontier model is smartest and starts arguing about signal quality, tariff rules, and approval SLAs—which is where industrial and fintech value actually lives.
For media groups the same split applies under media AI: drafts are cheap; publish authority is not.
We will keep shipping model upgrades. We will not quietly move the trust boundary. If your 2026 roadmap still has an LLM writing rate cards, detention hours, or safety trips without a deterministic owner and a human gate, rewrite the architecture before you scale the pilot.
Xenqube builds governed industrial and enterprise AI where sensors, rules, and human approval own the decisions that matter.
Request an architecture brief → · Explore governed AI → · Industrial AI solutions →
