Service
AI governance and evaluation
Governance is the evidence that a system does what you say it does, and evaluation is how that evidence is produced. The two are one engagement. We build the held-out test sets, run the security and privacy review, measure the things that drift, and leave you with documentation an auditor or a regulator can actually use.
The business problem
Most live AI systems cannot answer the two questions that matter under scrutiny: how do you know it is right, and what did it do last Tuesday. Without a held-out evaluation set there is no answer to the first. Without an audit trail there is no answer to the second.
That is a commercial exposure before it is a regulatory one. Quality drift is invisible while the service stays green, so degradation is discovered from complaints, which is the most expensive detector available.
Both the NIST AI Risk Management Framework and the EU AI Act ask for essentially the same artefacts a well-run engineering team would produce anyway. Our own standard is published at our Responsible AI Standard.
When this is useful
Governance work is usually bought at one of these six moments.
A system already live
In production, unevaluated, and nobody can say whether last month was better or worse than this one.
A regulatory question
An obligation under the EU AI Act, a sector regulator, or a client's security review.
A vendor product to assess
Something being bought, and no independent read on its accuracy or its data handling.
An incident
A wrong answer that reached a customer, and no way to establish how often it happens.
Before scaling
Widening a system's scope or autonomy, where the evidence has to precede the decision.
A risk committee request
A requirement to show oversight and auditability rather than assurances.
What we assess
Ten areas. We scope to what is relevant, and we say which ones we did not cover.
- Accuracy. Measured on a held-out set of real cases with agreed correct answers, including the rare and difficult ones.
- Reliability. Behaviour under load, on malformed input, and on the unhappy path the demonstration never met.
- Hallucination testing. Rate of confident wrong answers, traceability to source, and whether the system will say it does not know.
- Security. Prompt injection, tool permissions and data exfiltration through retrieval, probed against the OWASP Top 10 for LLM applications.
- Privacy. What data leaves your estate, retention, training terms, and whether personal data appears where it should not.
- Bias. Outcome differences across the groups your process affects, with the limits of the measurement stated plainly.
- Human oversight. Whether review points are placed by damage profile rather than convenience, and whether reviewers can actually check.
- Auditability. An evidence trail of inputs, outputs, tool calls and decisions, retrievable before someone asks.
- Monitoring. Quality, refusal rate and cost per decision as a time series, alerting on change rather than absolute level.
- Model and vendor risk. Deprecation notice, price mechanics, retention terms, and whether you could switch provider in a fortnight.
What you will receive
- A held-out evaluation set of real cases with agreed answers, plus a frozen slice nobody optimises against. This is the asset that outlives every model you use.
- A measured baseline: accuracy, refusal rate, latency and cost per decision as they stand today.
- A findings report with severity, likelihood and the specific remediation for each item.
- A security and privacy review of the AI-specific surface, written for your security team.
- Monitoring and alerting configured in your own stack, not a vendor console.
- A governance file mapped to NIST and, where relevant, the EU AI Act, with the gaps named rather than glossed.
Expected business outcomes
Drift detected early
Degradation found by a dashboard rather than by a customer.
Changes made safely
A model or prompt swap tested in a week, because the evaluation set exists.
Answers for auditors
Documentation and evidence available before the request rather than assembled after it.
A defensible risk position
Known accuracy, known failure modes, stated limits.
Reduced incident cost
Because the expensive failures are the ones nobody was sampling.
Scope widened on evidence
Autonomy or coverage extended where measurement supports it.
How the engagement runs
Two to six weeks depending on how many systems are in scope.
Frequently asked questions
Is this a compliance exercise or an engineering one?
Both, and they are cheaper done together. Most of what a regulator or auditor wants is evidence a well-run engineering team would generate anyway.
Does the EU AI Act apply to us?
It depends on the use case and where your users are, not only where you are. Part of the work is establishing which of your systems fall in scope and which do not.
Can you test for hallucination?
You can measure it, on a held-out set of real cases with agreed answers, and track the rate over time. Any claim to eliminate it should be treated with suspicion.
Can you evaluate a system we did not build?
Yes, including a vendor product. That is a common engagement and we have no incentive in the outcome.
How often should evaluation run?
On a schedule, permanently. A system that passed in March can fail in September because a model was updated or the request population drifted.
Do you certify or accredit systems?
No. We are not a certification body. We produce the evidence and the documentation an auditor, a regulator or your own risk committee needs to make a judgement.
What about bias in our data?
Measured against outcomes for the groups your process actually affects, with the limits of the measurement stated. A single fairness number is usually a false comfort.
Where does model and vendor risk fit?
Deprecation notice, price change mechanics, retention and training terms, and whether you could switch provider in a fortnight. Governance and portability are the same question.
Relevant insights
Start a conversation.
If a system is live and you cannot say how often it is wrong, that is the engagement. Two weeks is usually enough to find out.