Skip to content
← Projects

2026 · Solo · agentic systems platform

Governed multi-agent enterprise platform

A multi-agent platform for enterprise data with role-scoped tools, read-only SQL access, human approvals that survive restarts, and end-to-end audit traces, evaluated offline on real documents and 541,909 real transactions.

Real transactions
541,909
Docs faithfulness
0.94
Approval gate
1.00
  • Python
  • LangGraph
  • FastAPI
  • MCP
  • PostgreSQL
  • Ollama
  • Docker
  • Kubernetes

The problem

Enterprise workflows can't hand an LLM a database connection and hope for the best. Airlock Agents is a multi-agent platform that answers questions over internal docs, queries a real analytics database, and proposes actions — with role-based permissions, human approval gates, and a full audit trail sitting between every agent and the systems it touches.

Approach & tradeoffs

A supervisor agent in LangGraph routes each request to a specialist — docs_qa, analyst, or actions — but the specialist only exists with the tools its caller's role grants. Every external system — docs, database, tickets — sits behind its own custom MCP server, so agents never hold credentials directly:

  • Defense in depth, not prompt engineering. Role→tool grants decide what the agent is even built with; the MCP server re-checks on every call; the SQL server connects with a read-only DB role. Three independent layers.
  • Approval as durable state. A critical tool call interrupts the LangGraph run via interrupt(); the graph checkpoints to Postgres and can resume hours later from an approval queue — not a callback that dies with the process.
  • Fully local model serving. Ollama probes host VRAM/RAM at startup and picks a model tier (qwen3:32b down to qwen3:4b on CPU) — no prompt, document, or query ever leaves the machine.
  • One trace ID end to end, linking the HTTP request, agent steps, MCP calls, Langfuse LLM trace, and audit-log rows — so a bad answer is debuggable, not just loggable.
  • Packaged to deploy, not just to demo. Docker Compose for local runs and kustomize manifests for Kubernetes: API and MCP deployments, autoscaling, and a migrations job.

The datasets are real, not synthetic: an 8-page slice of the public GitLab Handbook for docs_qa, and 541,909 real UK retail transactions (UCI Online Retail, cancellations and guest checkouts included) for analyst. Eval ground truths are computed by SQL against the loaded data, not hand-typed.

Results

Offline eval gate on qwen3:8b (RTX 2070 8 GB) — passed:

SuiteMetricScoreThreshold
docs_qacorrect source cited0.86≥ 0.80
docs_qafaithfulness (LLM judge)0.94≥ 0.75
analystfigure matches SQL ground truth over 541k rows1.00≥ 0.70
routingsupervisor picked the right specialist0.92≥ 0.85
actionscritical tool correctly interrupted for approval1.00= 1.00

On the same GPU, a docs answer takes about 6 s (~2.1k tokens) and an analyst run 8–20 s of multi-step SQL over the 541k rows.

The gate isn't decorative — it caught the supervisor→subagent handoff hallucinating ticket confirmations on a small model, and an analyst run answering a 2024 question with all-time figures.

What I'd flag

The eval sets are small (4–14 cases per agent) and judged by a same-family local model, so treat the scores as a regression gate, not a benchmark claim. The next real step is a load-tested vLLM serving path for multi-GPU deployment — Ollama is fine for one resident model, not for concurrent load.