date

category

Announcement

read time

3 mins

share p0st

Financial institutions have thrown enormous teams at case review — financial crime, fraud, disputes, all the back-office work where a mistake means real money lost, a customer harmed, or a regulator asking hard questions. AI has been part of that work for a while now, but only ever as an assistant. A human still reviews every case before it closes. That's not caution for its own sake — it's because no general-purpose model has been accurate enough, or consistent enough, to safely take that human out of the loop.

Today we're launching Research Lab, the research division inside Arva AI dedicated to closing that gap. The Lab builds the proprietary models and infrastructure that let banks automate their highest-risk decisions and its first outputs are already running in production at global financial institutions, including a top-10 US bank.

Why "human-in-the-loop" hasn't been enough

General-purpose models are good generalists. That's exactly the problem in a domain like financial crime, where the cost of being wrong is so asymmetric: miss something and you've got a regulatory breach; flag too aggressively and you've buried your analysts in false positives. Off-the-shelf models weren't built for that trade-off, so banks have kept humans reviewing everything, using AI as a research assistant rather than a decision-maker.

As Rhim Shah, Arva's founder and CEO, puts it:

"Banks keep humans in the loop because no AI has been accurate enough to remove them safely — that's the problem the Lab solves. Our models and AgentCore let us automate these decisions with the accuracy and control banks require, and this is just the start."

What the Lab has built

Over 5,000+ hours of research, training, and evaluation went into the Lab's first two releases, both already live at banks today.

Proprietary models for the highest-risk parts of a decision

Rather than trying to build one model that handles an entire case end to end, the Lab is going component by component — starting with enrichment (the research step where an analyst pulls together everything relevant about an individual or business) and expanding from there into transaction analysis and evidence-based reasoning, the parts of a case that require the most judgment and carry the most risk if they're wrong.

AgentCore

AgentCore is the infrastructure layer underneath those models. Every correction an analyst makes, every insight they surface, every case outcome — AgentCore turns that into a system improvement. Crucially, none of it goes live untested: every change is backtested, evaluated, and versioned before it ever touches a real decision. For banks operating under strict model-risk and audit requirements, that governance is as important as the accuracy gains themselves.

The results so far

The first model out of the Lab is Arva Intel, used to research suspicious individuals and businesses online — the enrichment step analysts spend so much of their time on. In independent evaluation against frontier general-purpose models, Arva Intel came out 13% ahead on precision.

That figure comes from scoring the individual components inside a decision, not just the final case outcome — a stricter measure, since benchmarking only the outcome can hide where the accuracy actually comes from. It's the same standard the Lab is holding every future model to.

What's next

The Lab is starting where the stakes and the case volumes are both highest: financial crime and fraud. From there, the plan is to extend the same approach into payment exceptions, disputes, and other customer-related investigations — anywhere a bank is currently relying on a human to make a high-risk call because nothing else has been accurate enough to trust. The team also plans to publish its benchmark methodology and research in the coming months.

Read the official press release on Business Wire.

The latest

News & Research

Research notes, product updates, and company developments — straight from the team.