Capability · Epcotec

AI / LLM enablement

RAG, assistants, productivity suite connectors, eval harnesses, cost controls. Delivered by named senior engineers, fixed scope, staging from week one.

Start a project All capabilities
What we deliver

Scope, in plain terms.

Grounded retrieval assistants, deployed on-tenancy

RAG assistants deployed inside your tenancy with citation-grade answers, permission inheritance from your existing access control, and complete audit trails for every answer served.

Enterprise connectors with delta sync and ACL mapping

Go and TypeScript connectors for document stores, CRM and intranet systems, with delta sync, ACL mapping and resumable backfills that keep the index faithful to source permissions.

Evaluation harnesses as release gates

Golden sets, groundedness scoring and regression suites wired into CI, so every prompt, model or index change is measured against a fixed quality bar before release.

Cost controls and model routing

Routing tiers, caching and token budgets that cut inference spend without cutting answer quality, with per-team cost attribution so budgets are visible to engineering and finance.

Document intelligence pipelines

Python classification, extraction and normalisation pipelines at millions of pages scale, with human-in-the-loop gates where confidence is low and provenance retained for every downstream field.

AI receptionist agents

Front-of-house agents that answer routine enquiries, qualify requests and hand off cleanly to people, with transcripts, escalation rules and quality metrics reviewed weekly.

LLM-era security review

Prompt injection testing, data egress checks and tenant isolation review, producing a signed findings report your security team and auditors can accept.

Rough time estimate

Window: 4-10 weeks

Team: 1 AI engineer + 1 platform engineer

Complexity: M-L


Indicative only. A fixed price follows a free 30-minute scope review.

Get a fixed quote
Delivery process

Fixed scope. Weekly demos. Staging from week one.

Week 0-1 · Discovery & fixed-scope SOW

A working session on your data, permissions and boundaries. We map sources, egress rules and success criteria, then sign a fixed-scope statement of work with the evaluation bar written in.

Week 1-3 · Data characterisation & on-tenancy build

We profile corpus shape, freshness and access control, then build the first assistant and connectors on-tenancy, on pgvector hybrid search, with permissions inherited from source systems.

Week 1-4 · Weekly demos, staging from week one

You watch working software every week on staging from week one, not slides. Demos are on your data, in your environment, against the agreed scenarios.

Week 3-8 · Eval harness wired to CI

Golden sets and groundedness scoring run in CI as a release gate. Every prompt, index or model change must clear the bar before it ships.

Week 8-10 · Handover & 30 days support

Runbooks, dashboards and cost controls are handed over with a walkthrough for your engineers, followed by 30 days of support while your team takes ownership.

Benchmark deliveries

Proven in this discipline.

Grounded assistant inside a government boundary

Deployed fully on-tenancy inside a government security boundary: P95 answer latency 2.1 seconds, groundedness score 0.94 on the golden set, and zero data egress findings at review.

Document intelligence at 2.4 million pages

A pipeline processing 2.4 million pages with human-in-the-loop gates at low confidence, provenance retained per field, and throughput tuned so review queues stayed within agreed service levels.

Enterprise connectors at 1.2 million objects

Go delta sync connectors indexing 1.2 million objects with ACL mapping faithful to source permissions, delta sync keeping the index current, and the compliance review passed first pass.

More evidence in the delivery history and the complete 734-engagement register.

FAQ

Asked before signing.

Where does our data actually stay?

Everything runs on-tenancy: models, indexes and logs inside your boundary. No corpus, query or answer leaves your environment, and we evidence this in a signed security review.

Which models do you use?

Open-weight models run in your environment where residency demands it, hosted models where they do not. We route per query by cost and risk, and you stay free to swap models.

How do answers inherit permissions?

Connectors map source ACLs into the index, and retrieval filters by the asking user at query time. Users see only what they could open in the source system, verified by test sets.

How do you control cost?

Tiered routing sends routine queries to cheaper models, caching absorbs repeated questions, and token budgets cap worst cases. Per-team dashboards show spend, and the eval harness proves quality did not drop.

Talk to the engineers.

No account managers. Your message lands with the people who would deliver it, and you get a straight answer within one business day.

Book a discovery call