AI model services
Fraud detection, credit risk, document extraction, transaction intelligence. Built on your data, evaluated against your metrics, and handed back as models you own and can run inside your own network.
Purpose-built models for the problems fintech actually runs on — where class imbalance is severe, the latency budget is fixed, and a wrong answer has a cost you can put a number on.
Take the behaviour you like from a large model and move it into a small one you can afford to run everywhere — cheaper per call, faster at the tail, and inside your own network boundary.
The part that decides whether a model ships. We build the evidence a model validation function needs before it will sign anything off.
Because the constraints are unusual, and most general AI consulting doesn't take them seriously. A model that is accurate but takes four seconds is useless in an authorisation path. A model that is accurate but can't produce a reason code is useless in a lending decision. A model whose outputs shift when your serving topology changes is a problem you will discover during an audit rather than in testing.
We work on the version of the problem that includes all of that — cost per decision, p99 latency, calibration in the tail, data residency, and whether the whole thing can be reconstructed eighteen months later.
If you want to see how we think before you talk to us, that's what the blog is for: our writing on small language models in fintech, credit copilots that survive compliance review, where language models help in fraud, and what the current research actually says.
How we work
Most model projects fail on evaluation, not modelling. So we build the way to measure success before we build the thing being measured.
One task, one decision, one metric that maps to money. We would rather ship a narrow model that works than scope a platform that doesn't.
A few hundred hand-checked examples from your real traffic, including the ugly ones. This is the step everyone skips and the one that determines the outcome.
Establish what a strong conventional approach achieves before reaching for anything clever. Sometimes that is the answer, and we will tell you so.
Run alongside whatever you have today, compare on real volume, and cut over only on the segments where the new model actually wins.
Weights, training code, eval harness and documentation — delivered as notebooks and repositories your team can read and re-run. No black box, and no dependency on us to retrain it.
Model weights you own — self-hostable, no per-call dependency on us or a vendor API.
Training and inference code — readable, versioned, and runnable by your team.
The eval harness — so you can tell whether the next version is better.
Governance documentation — written for the people who have to approve it.
A short call is usually enough to tell whether this is a week of work, a quarter, or something you shouldn't build at all. We'll say which.