What a KYC Document Workflow Looks Like When the AI Runs In-House

The process everybody recognises

A new client is onboarded. Someone collects identity documents, proof of address, company registration papers and a shareholder register. Someone reads the ownership structure and works out who ultimately controls the entity. Someone cross-checks names against sanctions and adverse media. Someone writes it up. Someone senior reviews and signs.

It takes most of a working day per file, sometimes considerably more when the structure is layered across jurisdictions. It happens hundreds of times a year. The work is repetitive, but it is not mechanical, because the judgement at the end genuinely matters.

It is also made almost entirely of documents that belong to other people. Passport scans. Bank references. Corporate records. This is the collision at the centre of the problem: the process that would benefit most from automation is the one most difficult to justify sending to an outside service.

Where AI actually helps

Not everywhere, and being clear about that is what makes the rest credible.

Extraction. Pulling names, dates, registration numbers, addresses and shareholding percentages out of documents of wildly varying quality, including scans and photographs. This is genuinely tedious for a person and well suited to a machine.

Structuring. Turning a stack of corporate documents into an ownership tree, following holdings through intermediate entities and flagging where the chain breaks or where a document is missing.

Consistency checking. Comparing the name on the passport against the name on the register against the name on the bank reference, and surfacing the mismatches instead of relying on someone noticing at the end of a long day.

Drafting. Producing the first version of the onboarding memo in your own house format, with the findings and the gaps laid out.

Completeness. Checking the file against your own checklist and listing what is still outstanding, which is the single most common cause of a file bouncing back and forth.

Where a human has to stay

The risk assessment itself. Deciding that a structure is acceptable, or that a jurisdiction combination warrants enhanced due diligence, is a professional judgement. It is not delegated.

Anything ambiguous. Documents that conflict, structures that do not resolve, entities that cannot be verified. These should route to a person by design, not because someone spotted them.

Adverse media and sanctions conclusions. A match is a signal, not a finding. Deciding whether a name is the same person is a judgement with consequences in both directions.

The sign-off. Someone qualified puts their name to it. That does not change.

The honest framing is that this kind of workflow assists the people doing the work. It does not replace the professional judgement at the end of it, and any vendor telling you otherwise is selling you a problem.

The workflow, step by step

A private deployment of this kind of process typically looks like this.

  1. Intake. Documents land in a folder or are uploaded. Nothing is sent to an outside AI service; for supported workflows, processing happens on a machine on your own network.
  2. Extraction. Fields are pulled from each document and mapped to your own schema, with the source page recorded against every value so a reviewer can check it in one click.
  3. Structure resolution. Ownership is traced through the corporate chain. Percentages are calculated. Anything that does not resolve is flagged rather than guessed.
  4. Cross-check. Values are compared across documents. Mismatches are listed with both sources shown side by side.
  5. Checklist. The file is measured against your own onboarding requirements. Missing items are listed explicitly.
  6. Draft memo. A first version is produced in your house format, with findings, gaps and open questions separated.
  7. Human review. A person works through the flags, the mismatches and the judgement calls. The draft is evidence for the reviewer, not a replacement for the review.
  8. Sign-off. Unchanged.

The gain is not that step eight disappears. It is that steps two to six, which are where most of the hours sit, stop consuming a qualified person’s day.

Size the value before anyone talks about software

This is the part most projects skip, and it is the part that decides whether the project is worth doing.

Take the number of files you process in a year. Multiply by the hours each one currently takes end to end. Multiply by a fully loaded hourly cost for the people doing it. That figure is what the project is measured against, and it is usually far more interesting than any software price.

Then ask a second question that matters just as much: what does a delayed onboarding cost you in practice? Sometimes the answer is a client who went elsewhere, and that number dwarfs the first one.

Express the result in man-days and capacity rather than in currency alone. “This frees roughly the equivalent of one and a half full-time people for higher-value work” lands differently to a percentage.

What you need before starting

A workflow like this is not bought off a shelf. Before anyone can scope it honestly, you need:

  • A named owner of the process who can answer questions and make decisions
  • Ten representative files, sanitised, including the awkward ones. The clean examples teach nobody anything
  • Your actual checklist, as it is used, not as the manual describes it
  • The exception rules. What currently makes a file stop and go to a person
  • Volume and time data. Files per year, hours per file, who touches it
  • A willingness to test. Someone from the team has to run real files through it and say where it is wrong

If those six exist, the process is a strong candidate. If several are missing, the honest answer is that the workflow needs to be understood before it can be automated, and that assessment is the first piece of work rather than a formality before a proposal.

Why this one suits a local deployment

Three reasons, in order of how much they usually matter.

The material belongs to other people. Your client’s identity documents are not yours to send anywhere. That is a different argument from data security, and it is the one that tends to end the discussion.

Volume is predictable and high. A process that runs hundreds of times a year with long documents is exactly the shape where consumption pricing becomes uncomfortable and a flat cost becomes attractive. There is no per-token charge for local inference.

You may be asked where it was processed. Being able to answer that with a machine and a room, rather than with a contract clause, is materially easier.

A note on what exists today

Estha on Mac ships with a library of ready-to-use ChatApps and a growing set of AI Agents covering document processing, data analysis, presentations and meeting intelligence. Workflow agents built around a firm’s own process, of which KYC is one, are scoped and built per organisation rather than bought as a standard feature.

We say that plainly because the alternative causes problems later. A workflow agent for your onboarding process is a project with an assessment at the front of it, not a switch to turn on. What it is not is speculative: the shape above is well understood, and the assessment exists precisely to establish whether your version of it holds up.

Regulated workflows may require additional scoping, validation, audit trails and human approval, and those are quoted separately.

Frequently asked questions

Can AI do KYC document review?

It can do the extraction, structuring, cross-checking, completeness checking and first-draft writing, which is where most of the time is spent. The risk assessment, the treatment of ambiguous cases and the sign-off remain with a qualified person. Treat any claim of full automation with suspicion.

Is it safe to use AI on client identity documents?

It depends entirely on where the processing happens. Sending them to an outside AI service raises a question about material that belongs to your client rather than to you. A deployment where inference runs on your own machine, for supported workflows, avoids that question rather than answering it.

How much time does this actually save?

That depends on your file complexity and your current process, which is why the assessment comes first. The useful calculation is files per year, multiplied by hours per file, multiplied by loaded hourly cost. Any figure quoted before someone has looked at your workflow is a guess.

Do we have to change our systems?

Not usually. This kind of workflow works with files and exports to standard formats. Live integration with a CRM, practice management or registry system is not a standard capability and would be assessed separately.

What happens with an exception?

It routes to a person. That is a design requirement, not a limitation, and any workflow that cannot explain its exception path has not been designed properly.


Start with the assessment

If client onboarding is expensive, repetitive and too confidential for cloud AI, the first step is mapping the workflow: inputs, outputs, exceptions and expected value. Everything else follows from that.

One email a month on private AI. For people who cannot put their work in the cloud. No spam, unsubscribe anytime.

Talk to us about Estha on Mac

more insights

Scroll to Top