How it works
A supplier assurance workflow that has run in production for two years — and the three independent checks behind every finding.
Your experts stop reading documents and start reviewing findings
Someone senior reads several hundred pages of a supplier's response pack, checks it against dozens of controls, and writes up what's missing. It takes weeks, it happens under deadline, and the quality quietly depends on how tired the reviewer was by control 58. We automated the reading — not the judgement.
Documents go in as they arrive
Scans, photographs, diagrams. Nobody retypes anything, and nothing is rejected for being messy.
Every control gets walked, every time
The last one gets the same attention as the first. Coverage stops depending on stamina.
A rival model challenges the result
Built by a different company, it reads the evidence and forms its own view before it is allowed to see the assessment. It can flag, never edit.
A third company's model validates the document
A separate question, deliberately outside the subject matter: does this hold together, and would it survive an auditor?
Seventeen mechanical checks, then your expert signs off
No opinion involved in the checks. Nothing is filed until a person approves it, and the system cannot be configured otherwise.
Every finding cites its page
An auditor opens the source and reads the same paragraph. No finding rests on "the AI said so."
Weeks compress to days, and coverage stops drifting between reviewers. Representative of a workflow in production for two years; client, framework and control set are confidential.
The blind model benchmark, and why we didn't promote the open model+
Those figures come from a blind run against the frontier model on fifty real cases taken from live work — not a public leaderboard, and not a demo set we chose.
And the honest part: we did not promote it to lead assessor. An earlier run of only four cases made several candidates look flawless. Expanded to fifty, one of them wrongly marked seven cases compliant — the worst error a compliance tool can make, because it tells you that you passed when you did not. Four cases is not evidence. The open model earned the challenger seat; the frontier model kept the pen.
Every candidate is judged on the same five questions — can it hold the whole problem at once, does it reason or only follow rules, does it use tools reliably across a long job, does it invent things under pressure, and what does it cost per assessment. Then one hard rule: clear a full blind run and report the false-pass count, never a headline score on a small sample. That rule exists because a small sample already fooled us once.
We tested the thing everyone sells — and it lost+
Vector search is the industry's default answer for document retrieval. We measured it on our own production corpus against plain lexical search, for the lookups compliance work actually depends on — control references, standard numbers, document identifiers.
Embeddings lose that badly not because embeddings are bad, but because the two approaches point in opposite directions. Semantic search wins conceptual questions; lexical wins identifiers. Anything sold as purely one or the other is leaving results on the floor. We publish the numbers because most vendors in this market don't have any.
A chat interface your team already knows how to use
No training programme. The difference is what sits behind it.
▸ response pack · page 72
Interface mockup Representative of the working product. Content shown is fictional.
This is a demonstration of the platform. The practice that designs, builds and hands it over — along with what that costs — lives on cplt.tech.