We operate a live verification stack on an ongoing basis: monitoring the signals that predict failure, handling vendor incidents and failover, working the exception queues, keeping per-check cost visible, and making changes as vendors and requirements move. It is a retainer for a system that needs continuous attention, not a project with an end date.
Why a verification stack needs continuous ownership
An integration is finished. A verification stack never is, and the gap between those two facts is where most of the operational pain in onboarding lives.
Consider what changes underneath a stack that nobody has touched in a year. Vendors deprecate endpoints and alter response formats. Upstream registries change availability patterns. Your own volume shifts, seasonally and as products launch, and behaviour that was fine at one volume stops being fine at another. Commercial terms come up for renewal on tiers that were negotiated against a different volume mix. Requirements are amended. Meanwhile the engineer who built the integration has moved to another team, and their reasoning left with them.
None of these produce an outage. They produce slow degradation: a success rate that drifts down a little each month, a manual queue that grows, a cost per verified customer that rises without anyone deciding it should. By the time it is visible on a dashboard, several months of applicants have already had a worse experience than they needed to.
Most lending teams know this. What they lack is someone whose job it is to watch it. Verification is nobody's full-time responsibility until it breaks, at which point it is briefly everybody's.
What the retainer covers
- Monitoring the signals that matter — success rate by check and by vendor, response time distributions rather than averages, the composition of errors, how often fallbacks activate, and the depth and age of manual review queues.
- Vendor incident handling. When a provider degrades, we decide whether to fail over, do it, tell your operations team what is happening in language they can act on, and take up the post-incident conversation with the vendor.
- Exception queue operations — keeping the manual review path moving, and more importantly feeding back what the exceptions have in common so the queue gets structurally smaller rather than just being worked harder.
- Cost visibility per completed verification, which is the number that actually matters and is rarely the one on the invoice, since billing is usually per attempt and attempts include retries and failures.
- Change management — vendor migrations, new checks, deprecations and requirement changes, executed under your change process rather than around it.
- A standing point of contact who knows your stack, so an incident does not begin with somebody reading the code for the first time.
How we operate
We work inside your systems, not alongside them. The code stays in your repositories, runs on your infrastructure, and changes through your review and deployment process. We are additional engineers with specific expertise and continuous context, not an outsourcer with a black box.
The rhythm is a weekly review of the operational signals with a short written summary of what changed and what we did about it, immediate handling of anything that cannot wait for that, and a periodic deeper review of the things that only show up over months — cost drift, vendor performance trends, and whether the check mix still matches the products you are actually selling.
We prefer to be judged on the composition of what is happening rather than on a green dashboard. A stack with high overall success and a manual queue quietly doubling is not healthy, and a report that shows only the first number will not tell you that.
What we do not do
We do not handle end-customer identity data. The verification calls flow between your systems and your vendors; we build, watch and change the code that makes them, and we do not become a processor in the path. This is deliberate and it constrains what we can offer — some operational tasks a client might like to delegate are not ours to take.
We take no margin on your verification spend. What you pay a provider is settled between you and that provider, so when we tell you a check can be moved, batched or dropped to bring that bill down, nothing on our side moves with it. Advice about cost is only worth having on those terms.
We do not publish figures about vendor performance, including from what we see in operations. What we observe in a client's stack is that client's information, and a number we published about a third party would be stale before it was useful. Judgement, given directly to the client, is what we offer instead.
The signals we watch, and why those
Uptime is the metric most teams start with and it is close to useless here. A vendor that is technically responding while returning failures for a particular check is up by every naive definition and broken by every one that matters.
What predicts a bad week is composition. Success rate broken down by check and by vendor will show a single degraded check that an aggregate figure hides. Response time distributions matter far more than averages, because verification timeouts live in the tail and an average that looks stable can conceal a tail that has doubled. Error composition tells you whether failures are genuine negatives or infrastructure problems, which is the distinction that determines whether real applicants are being declined. Fallback activation frequency is an early warning that a primary vendor is drifting before the drift is visible anywhere else. And manual queue depth and age reveal whether the automated path is quietly handing more work to humans than it used to.
Watching those five together is the difference between noticing a problem in its first week and noticing it in a quarterly review.
Who this suits
This works for lenders running a live book on a verification stack that matters commercially but does not justify a dedicated internal team — where an incident is expensive, but not expensive enough to staff for permanently.
It suits teams that have just completed an integration and recognise that the knowledge is concentrated in one engagement rather than embedded in the organisation. It also suits teams that inherited a stack from a previous partner and need someone to take ownership of something they did not build.
It does not suit an organisation that wants a vendor to hold the risk. We operate the layer and we are accountable for how we operate it, but the regulatory obligation stays with the regulated entity, and any arrangement suggesting otherwise would be misrepresenting what an integration partner can carry.
Where the stack itself is the problem rather than its operation, start with KYC API integration and move to a retainer once it is live.