#91 - Have foundation models dethroned ML?
Are foundation models finally catching up to machine learning for fraud scoring?
Seems like every major payments vendor wants you to think so this year.
Stripe, Mastercard, and Revolut have all announced some version of a “fraud foundation model” in the last 12 months - and behind most of these claims sits remarkably little actual data.
But a couple of weeks ago I encountered a new published research from the team at Coinbase.
And let me tell you, it was an interesting read.
The claims so far
Before we deep dive into this new study, let’s recap what other prominent publications have we seen recently.
Stripe went first, although I’ll note that it was an unofficial linkedin post from their ML lead, and I already picked that one apart. They claimed amazing results, but gave no details other than a tiny table and some visual graphs.
Mastercard followed around Nvidia’s GTC conference with a “Large Tabular Model” - trained on billions of card transactions. Impressive in scale, but again - thin on actual detail about how the company evaluated it.
Revolut is the interesting exception: they published an actual paper in April 2026, unveiling PRAGMA: trained on roughly 40 billion banking events across 25 million users, claiming a 20% improvement in fraud detection.
In Revolut’s paper you can find a methodology section, but as you read through it, it’s very hard to ascertain what was exactly tested, how, and compared to what.
PRAGMA is also not a fraud model. It’s a general transformer model trained on banking data, that underwent some optimization for several use cases, fraud included.
Revolut claimed +16.7% precision and +64.7% recall versus their internal “external fraud” model. These are impressive numbers, but that’s about all we get in terms of detail.
I don’t doubt them, I just can’t verify them.
Then there’s the recent paper from Coinbase researchers, submitted in December 2025 and revised this spring. And it’s fundamentally different.
In two ways.
The first is that it names its champion algorithms, specifies dataset sizes and fraud ratios, reports metrics across four datasets, and includes latency numbers.
The second is that it’s not really a foundation model. It’s a “normal” LLM, guided by retrieved examples, but not pretrained on fraud data from scratch.
Not exactly apples to apples.But it’s the only paper that lets me judge for myself. That’s worth a careful read.
So I did just that.
What the research tested
The paper introduces a method called FinFRE-RAG for adapting LLMs to tabular fraud detection. It tested four datasets, but three are what the authors themselves call “toy” datasets in their “Limitations” section.
I’m ignoring those and focusing on the one closest to reality: IEEE-CIS: it has 590k transactions, a 3.5% fraud rate, and 393 raw features.
Against their model, the team ran three classical ML algorithms often used for fraud detection: Random Forest, XGBoost, and TabM - a modern deep tabular network.
The results
The team ran many(!) experiments and tracked several KPIs. You could see them all in the table below:
Let me draw your attention to the column in red (the dataset we’re interested in) and the two highlighted experiments: the best performing LLM-based model versus XGBoost.
Long story short - with a slightly better precision, XGBoost still managed to catch 20% more fraud (82% vs. 68%).
Now, I want to be clear on something - does that mean that any transformer-based model would do poorly against decision trees models in fraud detection?
Short answer - no.
Slightly longer answer - foundation models that are trained on domain data can beat ML models.
But it is more complex than transformers vs. decision trees. Because foundation models aren’t necessarily used for scoring.
They are used to enrich your features set. The same set that is then fed into an ML model to produce a better score than before.
In fact, that is exactly what Revolut’s PRAGMA was designed to do.
Side Note: I can also reveal that at Sardine we’re about to share some interesting (and much more detailed) news on how foundation models can be used for fraud detection. Stay tuned.
So back to the question - are foundation models going to dethrone ML?
No. They are going to make them even more powerful.
The bottom line
I give the researchers real credit.
They published a result that doesn’t flatter their own method on the one dataset that resembles production data.
That takes integrity and I highly respect them for it.
But I read the results more skeptically than they seem to. The gap between the best LLM setup and XGBoost isn’t as small as they like it to be.
Compared to the approach Revolut took, one might think that the performance gap is driven by using an LLM instead of training a foundation model from scratch.
But that’s just half of the story.
The other half is trying to replace the scoring mechanism instead of focusing on creating new features for models to consume.
And that shift is what turned me from a skeptic to a believer.
That’s what I’ve now seen firsthand in Sardine’s own research. Coinbase’s paper, if anything, confirmed the opposite approach doesn’t work.
Are you playing with foundation models for fraud detection? Hit reply - I’d be interested in speaking to you.
In the meantime, that’s all for this week.
See you next Saturday.
P.S. If you feel like you're running out of time and need some expert advice with getting your fraud strategy on track, here's how I can help you:
Free Discovery Call - Unsure where to start or have a specific need? Schedule a 15-min call with me to assess if and how I can be of value.
Schedule a Discovery Call Now »
Consultation Call - Need expert advice on fraud? Meet with me for a 1-hour consultation call to gain the clarity you need. Guaranteed.
Book a Consultation Call Now »
Fraud Strategy Action Plan - Is your Fintech struggling with balancing fraud prevention and growth? Are you thinking about adding new fraud vendors or even offering your own fraud product? Sign up for this 2-week program to get your tailored, high-ROI fraud strategy action plan so that you know exactly what to do next.
Sign-up Now »
Enjoyed this and want to read more? Sign up to my newsletter to get fresh, practical insights weekly!