Service
Generative AI consulting
Generative models are useful where the work is reading, drafting, summarising or answering, and where a correct answer can be defined well enough to test. We scope those workflows, build them so answers stay traceable to a source, and evaluate them on real cases rather than demonstrations.
The business problem
Generative AI demonstrates extremely well and fails quietly. A system that answers plausibly on prepared questions can be wrong on the cases that matter, and without a held-out set of real examples nobody finds out until a customer does.
The second problem is context. A model asked to make a judgement that depends on rules, exceptions and history nobody has written down will guess, competently. That failure looks like a model problem and is not one, which is the argument in your agent is not confused, it was never told.
So the work divides into two halves: supplying the context, and defining how you would know the answer is right.
Where it applies
The workflows where generative AI earns its cost, and where we have built.
Internal knowledge assistants
Answers drawn from your own policies, procedures and history, with citations back to the source document.
Document intelligence
Extraction, classification and comparison across contracts, invoices, reports and forms at volume.
Customer support
Drafting and triage with a defined escalation path, rather than an unsupervised front line.
Sales and operations copilots
Preparation, summarisation and next-step suggestions inside the tools people already use.
Retrieval-augmented generation
Grounding answers in your corpus so they can be checked, with permissions inherited from the person asking.
Content and research workflows
First drafts and literature reads where a human remains the editor of record.
Enterprise search
Search that answers a question rather than returning ten links, across repositories that do not talk to each other.
Report and briefing production
Recurring documents assembled from systems of record, with the numbers traceable.
What we will do
The pattern is consistent across use cases, and most of the effort is not in the model.
- Define the answer. A rubric for what correct means on your cases, built with the people who currently do the work.
- Assemble the context. The rules, exceptions and precedents the answer depends on, held in a format you own rather than only as embeddings in a vendor index.
- Build retrieval that respects permissions. The system sees what the person asking is entitled to see, and no more.
- Keep answers traceable. Citations to source, so a reviewer can check rather than trust.
- Evaluate on real cases. Including the ambiguous ones and the ones that should be refused, with a frozen slice nobody optimises against.
- Choose the model on evidence. Often a smaller one, per the reasoning in how to choose a model for production.
- Set the refusal and escalation path. What the system does when it lacks information or authority, which matters more than what it does when confident.
What you will receive
- A working system in your environment, with prompts and retrieval logic in your repository.
- An evaluation set of real cases with agreed answers, including cases that must be refused.
- Answer quality, refusal rate and cost per answer tracked alongside availability.
- Permissions and retention behaviour documented, including what leaves your estate and what does not.
- A measured comparison against the current process, on the baseline agreed before the build.
- The context corpus in a portable format, so the system can be rebuilt on a different provider.
Expected business outcomes
Stated against a baseline, and only where the workflow supports the claim.
Faster answers
Time-to-answer on questions that currently require finding a person who knows.
Consistent reading at volume
Documents processed to a defined standard rather than to whoever had time.
Reviewable output
Answers a human can check quickly because the source is cited.
Fewer escalations
Where the assistant handles the standard cases and reliably refuses the rest.
Knowledge that survives turnover
Context written down as a system asset rather than held individually.
Known unit economics
A cost per answer you can plan against at real volume.
How the engagement runs
The evaluation set is built before optimisation, not after the first incident.
Frequently asked questions
Is retrieval-augmented generation still the right approach?
For most enterprise question-answering, yes, because it keeps the answer traceable to a source. The interesting work is rarely the retrieval mechanism; it is the context and the evaluation around it.
Will it hallucinate?
Any generative system can produce a confident wrong answer. The question is whether you can detect it. That means a held-out evaluation set, citations to source, and a defined refusal path.
Can we use our own documents without them training a model?
Yes. Retention and training terms are contract questions we help you settle before anything is connected.
Do we need a vector database?
Sometimes. Search relevance often improves more from better metadata and permissions handling than from a different index.
How do you handle permissions?
The system inherits the permissions of the person asking. Retrieval that ignores access control is a data exfiltration route, and we treat it as a security surface.
What about cost at scale?
We model cost per answer at expected volume and at three times it. For high-volume workflows a smaller model on a narrower task often wins on both cost and quality.
Relevant insights
Start a conversation.
Name one question your people cannot answer quickly today, or one document class that eats a week a month. That is enough to start.