Stripe’s Payments Foundation Model: How Data & Infra Create Compounding Advantage, w/ Emily Sands

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
25 September 2025 1h 21m
0:00 --:--
Episode Description
Today, Emily Sands, head of data and AI at Stripe, joins The Cognitive Revolution to discuss how the company built a payments foundation model that processes tens of billions of transactions into dense embeddings, exploring the technical architecture behind fraud detection improvements and the modular approach that enables rapid deployment of AI across Stripe's $1.4 trillion payment network. Check out our sponsors: Google Gemini Notebook LM, Linear, AGNTCY, Claude, Oracle Cloud Infrastructure.

Summary

Emily Sands, Head of Data and AI at Stripe, discusses the company's payments foundation model, which processes trillions of dollars in transactions into dense embeddings. This model significantly improves fraud detection, boosting rates from 59% to 97% for card testing, and enables modular AI deployment across Stripe's vast payment network. The conversation also explores Stripe's strategy for iterating faster than fraudsters, leveraging LLMs for data labeling, and the broader implications of AI for incumbent platforms and agentic commerce.

Chapters

Introduction to Stripe's Payments Foundation ModelEmily Sands, Head of Data and AI at Stripe, discusses how Stripe's new payments foundation model processes trillions of dollars in transactions to improve performance across its product suite.
Payments as a Distinct Modality and Data ScaleThe model treats payments as a unique modality, assembling extensive context from tens of billions of transactions and hundreds of signals, which is intractable for humans but ideal for neural networks.
Fraud Detection Breakthroughs and Modular AIStripe achieved a 97% card testing detection rate using the foundation model and employs a modular approach, exposing embeddings for rapid development across existing ML systems.
AI's Advantage for Incumbent PlatformsThe discussion highlights how AI strengthens the data advantage and product lead of incumbent platforms like Stripe, creating a compounding flywheel effect.
Staying Ahead of Fraudsters: Iteration and LLMsStripe shortens its iteration loop against fraudsters using dynamic risk thresholds, adaptive 3DS, and leveraging LLMs as judges to fill in missing or late-arriving ground truth data.
Multimodal AI and Agentic Commerce TrendsStripe's roadmap includes expanding to new modalities like financial time series and images, and the emergence of agentic commerce, where AI agents perform transactions on behalf of users and businesses.
Reliability in 'Talk to Your Data' ProductsStripe ensures accuracy and trust in its Sigma 'talk to your data' product by using well-structured data and providing natural language explanations of the LLM's reasoning and computations.
Stripe as the Economic Infrastructure for AIThe episode concludes with a recommendation for startups to use Stripe as their system of record and an overview of Stripe's strategic AI investments, focusing on being the best partner for AI companies.

Topics

Payments foundation modelFraud detectionCard testingModular AI deploymentProprietary data advantageLLMs as judgeDynamic risk thresholdsAdaptive 3DS authenticationFriendly fraudMultimodal AIAgentic commerceTalk to your dataStripe as system of recordEconomic infrastructure for AIAI investments

People

Steven Johnson (mentioned) Nathan Labenz (host) Emily Sands (guest) Patio11 (mentioned) Shakespeare (mentioned) Patrick (mentioned) Claude (mentioned) Meta (mentioned) Cisco (mentioned) Dell Technologies (mentioned) Google Cloud (mentioned) Oracle (mentioned) Red Hat (mentioned) Eleven (mentioned) Character AI (mentioned) Illumix (mentioned) Lovable (mentioned) Bolt (mentioned) Replit (mentioned) Cursor (mentioned) Retail AI (mentioned) Perplexity (mentioned) Hipcamp (mentioned) Vercel (mentioned) Collison brothers (mentioned) Mistral (mentioned)
Key Concepts (16)
Payments Foundation Model — A transformer model developed by Stripe that processes tens of billions of transactions into compact vectors (embeddings) to understand payment patterns and deliver improved performance across various products like fraud detection and authentication.
Payments as a Distinct Modality — The idea that payment data, while represented in text, behaves like a unique modality with its own syntax and semantics, requiring specialized models beyond traditional language models.
Contextual Understanding of Payments — To properly understand a single payment, the model assembles extensive context, including recent activity associated with entities like the buyer, card, device, and merchant, which is overwhelming for humans but ideal for neural networks.
Card Testing Detection — A fraud detection use case where the payments foundation model significantly improved detection rates from 59% to 97% by identifying patterns of fraudsters testing stolen cards amidst legitimate traffic.
Proprietary Modality Strategy — The concept of businesses training foundation models on their unique, domain-specific data modalities (e.g., health, cybersecurity, logistics) to gain a superhuman understanding and competitive advantage.
Modular AI Deployment — Stripe's approach of exposing payment foundation model representations (embeddings) as additional inputs to existing ML systems, allowing engineers to rapidly improve performance without rethinking entire systems.
AI Flywheel for Incumbents — The idea that AI strongly favors incumbent platforms with vast proprietary data, creating a compounding advantage where scale leads to better models, better products, more customers, and further data growth.
LLMs as Judge — Using Large Language Models to evaluate the quality and accuracy of labels or explanations generated by other models, especially in contexts where no definitive ground truth exists, such as for suspicious but not definitively fraudulent payments.
Dynamic Risk Thresholds — A fraud prevention mechanism where Stripe's Radar system dynamically adjusts blocking thresholds in real-time when an attack is detected, allowing revenue to flow freely during normal periods but tightening defenses during attacks.
Soft Blocks — Adaptive authentication methods, like Adaptive 3DS, where instead of a binary block/don't block decision, the system selectively requests additional authentication from users in cases of suspected risk, allowing legitimate users through while deterring fraudsters.
Blending Rules with Models — An approach to fraud prevention that combines the precision of rule-based systems with the comprehensive pattern recognition of machine learning models, allowing for nuanced decisions that balance risk and user experience.
Friendly Fraud — Suspicious payments or user behaviors (e.g., free trial abuse, reseller abuse, refund abuse) that do not result in traditional fraudulent disputes but are costly to businesses, particularly AI companies with high marginal compute costs.
Multimodal AI Roadmap — Stripe's vision to expand its foundation model beyond payments and text to incorporate other modalities like financial time series or images, enabling more comprehensive intelligence, such as assessing merchant fraud or detecting counterfeit products.
Practical Explainability — Stripe's focus on making AI outputs self-explaining through natural language descriptions of the model's reasoning, allowing non-data analysts to understand and trust the results, particularly in 'talk to your data' products.
Stripe as System of Record — The recommendation for businesses, especially startups, to rely entirely on Stripe's APIs and data as their primary financial system of record due to high uptime and comprehensive features, avoiding the complexity of building parallel systems.
Agentic Commerce — The emerging trend where AI agents perform commercial transactions on behalf of users or businesses, ranging from booking campsites to procuring SaaS services directly within developer environments.
References (36)
Google company
NotebookLM by Steven Johnson product
Linear company
AGNTCY project
Claude by Anthropic tool
Oracle Cloud Infrastructure company
Stripe company
Turpentine company
Complex Systems by Patio11 podcast
Shepard product
Radar product
3DS
Replit company
Shortwave company
Goodfire company
Anthropic company
Cursor tool
ChatGPT tool
Illumix company
Sigma product
Stripe Atlas product
Link product
Stripe Capital product
Stripe Tax product
Stripe Data Pipeline product
Lovable company
Bolt company
Retail AI company
Perplexity company
Hipcamp company
Vercel company
Mistral company
LeChat by Mistral product
Forbes AI 50
AI podcasting company
Notion company
Transcript (73 segments)
Speaker 1

This podcast is supported by Google. Hey folks, Steven Johnson here, co founder of NotebookLM. As an author, I've always been obsessed with how software could help organize ideas and make connections.

So we built NotebookLM as an AI first tool for anyone trying to make sense of complex information. Upload your documents, and NotebookLM instantly becomes your personal expert, uncovering insights and helping you brainstorm. Try it at notebooklm.

google.com.

Speaker 2

Hello, and welcome back to the cognitive revolution. Today, my guest is Emily Sands, head of data and AI at Stripe, the programmable financial infrastructure company that in 2024 processed $1,400,000,000,000 in payments or roughly 1.3% of global GDP for everyone from solo entrepreneurs to the Fortune 100 and which continues to grow at a blistering pace.

We begin by discussing the many fascinating details of Stripe's new foundation model for payments and how Stripe is using this model to deliver improved performance across their broad suite of products. While it might seem unassuming at first glance, I would argue that the payments foundation model has several important lessons to teach us. First, while payments are represented in text, the payments foundation model is not a language model in the familiar sense.

On the contrary, payments are treated as a distinct modality, and importantly, no payment is an island. To properly understand a single payment requires Stripe to assemble extensive context, including recent activity associated with multiple entities, the buyer, the card, the device used to make the purchase, and the merchant. So much context quickly becomes overwhelming to humans, but this is exactly where neural networks can shine.

And indeed, when Stripe first deployed this model to detect card testing, which is a process that fraudsters use to determine which stolen cards actually work, they saw a jump in their detection rate from 59% to ninety seven percent. Obviously, a massive win, not just for Stripe, but for the entire ecommerce ecosystem that collectively bears the cost of fraud. Now if you've listened to this show for a while, you know that one of my pet theories is that the surest path to superintelligence is to integrate today's reasoning models with models that are trained on other modalities that humans aren't well adapted to understand.

I'd say it's safe to say that the payments foundation model is superhuman when it comes to understanding payments, And this conversation left me wondering how many other businesses are training foundation models on their own modalities, as well as how many other interesting modalities might still currently be hiding in plain text. I can imagine that this proprietary modality strategy might work on any number of domains, including health, cybersecurity, logistics, energy, and insurance. But to be honest, I haven't found too many other examples of this strategy being used today.

So if you happen to know of any other foundation models being trained on any interesting proprietary modalities, please do ping me and let me know as I would love to do more episodes exploring this theme. The next lesson, perhaps as important to Stripe's success as the model itself, is the way they are using it. Rather than trying to design the foundation model to support all use cases directly, they are exposing payment foundation model representations and thus allowing engineers to use them as additional inputs to the many classification and other ML systems that they've already developed.

The richness of the foundation model signal makes everything else work better, but doesn't require a major rethinking of existing systems. Again, outside of social network companies, who I do believe make their user and content representations available in this way, I've not heard of other companies taking this approach, and it seems to me now that more of them should consider it. Finally, the most important lesson from a societal standpoint might be that AI strongly favors the incumbent platforms that have the data necessary to train such differentiated models.

The flywheel that Stripe has created here, which translates their incredible scale to commercial advantage, is allowing them to reduce the cost of fraud for their customers even as fraud is rising across the broader ecosystem. This makes Stripe the obvious choice going forward, which in turn further strengthens their data advantage and product lead. It is genuinely hard for me to imagine how anyone, aside from a few of the world's largest tech companies, could ever compete with Stripe.

Meaning that even as history begins to unfold at a dizzying pace in many respects, competition in many key markets may effectively come to an end. This isn't necessarily a problem. I've never supported punishing companies for their excellence, and I've never been convinced that we should break up American tech companies.

But it does seem like something that policymakers will need to think long and hard about as they envision the AI future and hopefully begin to imagine a new social contract. There's a lot more in this episode besides these key strategic insights, including how Stripe is designing processes to iterate quickly enough to stay ahead of fraudsters, including by using LLM as judge to fill in missing data, how they ensure reliability in their LLM powered Talk To Your Data product experiences, how developers can accelerate product development by treating Stripe as their payments database of record, what Emily and team are seeing in AgenTek Commerce today, and how they think about scoping their AI ambitions and investments. All in all, as you might expect from Stripe, it's a high alpha episode with practical lessons for rank and file AI engineers and big picture implications for executive level AI strategists.

So without further ado, I hope you enjoy this deep dive into how smart use of AI is transforming one of the world's most critical financial infrastructure companies with Emily Sands, head of data and AI at Stripe.

Speaker 3

Emily Sands, head of data and AI at Stripe. Welcome to the Cognitive Revolution.

Speaker 4

Thanks for having me.

Speaker 3

I'm excited for this conversation. Stripe obviously is a, global recognized leader in payments and doing some really interesting things in AI with high standards everywhere and, obviously, a lot of shared DNA with some of the big frontier AI developers. So a lot, to get into today.

For folks who wanna do a deeper dive into Stripe and the payments ecosystem, you did maybe six months ago now, a podcast with our sister pod complex systems with patio eleven.

Speaker 4

Patio 11.

Speaker 3

And I would definitely recommend that for folks who wanna do a, you know, deeper primer on the, you know, the the payments world, which is a, you know, a fascinating and Byzantine one with many rabbit holes to go down. We won't do nearly as much of that today. We'll kinda stay more focused on some of the cool new AI stuff that you guys are doing.

But maybe just for, like, a super quick primer, how would you describe the role that Stripe plays in the economy? And then we'll use that as jumping off point to get into the AI stuff.

Speaker 4

You said payments infrastructure. We started as payments infrastructure. Absolutely true.

We now build broader programmable financial infrastructure. So in plain terms, we give any business. Right?

It could be a teenager who's selling a Figma template, or it could be, you know, any one of now more than half of the Fortune 100 that run on Stripe, the rails and intelligence to move money online, and to grow faster. So last year, companies processed $1,400,000,000,000 through Stripe. And we'll talk about AI today.

You know, every one of those charges becomes training data for the AI systems that we'll talk about. But that flywheel also means that we're no longer just the payments API. We optimize the entire payments life cycle.

Yes. The the gory details are are covered with Patio 11, but it's the checkout UX, fraud prevention, bank routing, retries, even things like, you know, dispute paperwork so that businesses can really keep more of every hard $1, and scale up with very small teams. So we think of the tools we're building as structural growth tailwinds, and we're already seeing it in the data.

I mean, businesses on Stripe grew seven x faster, than the S and P 500 last year.

Speaker 3

Wow. Okay. A lot of good, nuggets there.

I have been a customer actually for what it's worth since nothing maybe the earliest early days, but pretty early days, like, at least ten years that I've been, a Stripe customer with my company Waymark. So we've seen a lot of the, evolution from the customer side. The biggest thing that has caught my attention in terms of what Stripe is doing with AI is the payments foundation model.

And I'd love to just spend, you know, a good chunk of time really going into details on that because one of the things that I have been fascinated with and kind of trying to see around the corner and better understand is to what degree are we gonna get a form of super intelligence via AIs that become sort of natively capable of understanding potentially a huge range of different modalities. And, you know, people are familiar now with, like, image generation. Of course, we had, like, text images with images image generation models.

Now those have kind of come together in this really tightly coupled, you know, deeply integrated way with the Nano Banana and other recent innovations in that space. And so I have this theory that, like, one thing that people really in general underappreciate is the degree to which training on these other modalities of data is just gonna create superhuman capability in these domains that, you know, are sort of familiar to us, but also in many ways, like, very alien. So maybe for starters, like, what can you tell us about sort of the the fundamentals of the payments foundation model?

Like, what does the data look like?

Speaker 4

but, you know, give us more detail on that. Like, what is transaction data when you really get into the weeds of it? Yeah.

And it's a good point. I mean, there's been a ton of coverage of sort of the large scale traditional LLMs and a lot less coverage of domain specific foundation models, of which the payments foundation model is one. For us, it's been really a step function change in the speed and quality with which we can deliver all of those optimization solutions I talked about in auth, in fraud, in disputes.

At its core, it's a transformer model that turns every payment, so the tens of billions of transactions that run through Stripe, into a compact vector. So it's like giving each transaction its own kind of latitude and longitude. And then once you have that map, you can use it for all sorts of downstream tasks, right, to figure out what's fraud, to figure out how to authenticate, to figure out what's a valid versus invalid dispute without having to train a new model from scratch every time.

And I think, you know, what what makes it work, you know, the reason you can build a domain specific foundation model in the payments context is Stripe scale. So we process about 50,000 new transactions every minute. And at that density, payments start to look in a lot of ways, not in all ways, but in a lot of ways, like language.

So there's kind of a syntax to a payment. Right? There's the card bins and the merchant codes and the amounts.

And then there's sort of an analog to semantics, so sort of how a device or card gets reused over time. And so in the same way that, like, language transformers are learning embeddings and words with similar meanings clustered together, so the sort of premise of the payments foundation model is just what if every charge or sequence of charges, and we can talk about that too, had its own vector in a similar space. So the inputs are, you're right, just the raw payment signals as they come in, the card details, the merchant categories, the IPs, but also those sequences.

So what a given card or device or merchant bin or customer has been doing in the last few minutes or the last k transactions. And it's actually that history that turns out to be a huge unlock. And then from those inputs, the model produces an output, which is just a reusable embedding.

Right? It's a it's a dense vector for each payment or short sequence, and then we can layer lightweight classifiers on top for real time detection. We also have a slower, higher latency variant that generates explanations through a text decoder.

And I think we'll get to a stage where that can be real time ish as well, but we're not there just yet.

Speaker 3

Cool. Okay. That's there's already a number of interesting things there.

In terms of scale, the blog post that introduced the payments foundation model said tens of billions of transactions, and then it also indicated hundreds of subtle signals. Could you go is there a, like, you know, couple of examples of sort of the long tail of these signals that kind of illustrate just, like, how how much information the model is able to ultimately take in that might be, like, hard for a person to rep, you know, because we can classically handle, like, seven, you know, items in working memory. Right?

So what are we missing, you know, with with our feeble human working memories that the model is able to take in? Yeah. And from there, I'm I'm kinda interested in the overall scale of data.

Sounds like it's getting into the trillions of tokens, which would be, like, not at the high end of text foundation models, but, like, not too far off, maybe, like, one order of magnitude less.

Speaker 4

estimates with you on that. Yeah. Your your math is legit.

I'll answer the second question first and that first question second. Yes. Your math is is legit and the data is very different from the free form text that you'd use to, you know, train a model to write like Shakespeare.

Right? So payments data is highly structured and dense, and so we actually build a custom tokenizer that compresses the numeric and categorical signals really efficiently. So, yes, the dataset is big, but it's also packed with kind of this purpose built information that's incredibly rich for the set of tasks that we care about in our context.

You asked about, like, what's hard for a human to eyeball. I think the thing that's hardest for the human to eyeball is the you know, looking across those dimensions, not within any one payment, but within any, you know, combinatorial sequence. So, you know, if you think about it, what you need to look at in order to figure out if, say, a fraud attack is happening or how to get a payment, authenticated has very little to do with that particular transaction and everything to do with where that transaction sits vis a vis the transactions that have come around it.

So you're not looking at, like, a single screen. You're looking at, like, you know, a clip of a movie. But there are a lot of different clips that include that screen that are relevant to look at.

Like, you wanna know what I was doing. You wanna know what the merchant was doing. You wanna know what my card was doing.

You wanna know what my IP was doing. And so that's really where the model sings, making it really efficient, not just to look at the individual payment. Like, it's hard to do with the scale of 50,000 a minute, but, you know, a human could, I suppose, if you had enough humans.

It's really about the sequences that make the problem intractable for humans, but also very hard for sort of traditional ML approaches where you have to hand engineer features to capture what's happening in each of a range of different sequences. And so in our context, the foundation model pays off dramatically because it expands really three things. One is how much data we can learn from.

Like, we can learn from literally all of Stripe's history, not just a task specific subset of history. It changes how richly we can learn. Right?

So these dense embeddings capture very subtle interactions that manual feature lists, like counterfeatures, wouldn't capture. And then third, which is more about sort of how we work internally, but it changes how efficiently we can build. Right?

Once you have a shared embedding, then spinning up a new model becomes a weekend project, not a quarter project, and that means we can sort of open the aperture for the types of ML powered solutions we can build.

Speaker 3

Yeah. Cool. I really like the idea of kind of multiple clips.

So and I take it that that just basically reflects the reality that there are obviously multi there are multiple parties to any transaction, and I'm kind of inferring that, like, the pattern of behavior of each of those different parties is really where the sick the strong signal is. It's not the if if you looked at this particular transaction in isolation, you might not get much. But when you combine recent history for all of the parties to a single transaction, the combination of those recent histories is really what tells you what you need to know.

Do I have that right? Wait.

Speaker 4

me using my card in Boston that tells you it's fraudulent. But if I just use my card on my device at my home IP, which is, by the way, at Palo Alto, not in Boston, and tend to be buying on, you know, things that are totally different from what you suddenly see someone doing in Boston, that's a red flag that that's actually, like, fraudulent use of my card. Or conversely, if you see someone using you know, rotating across a small number of cards to buy thousands of accounts from a given AI provider, maybe the card is truly there, but you're almost certainly gonna see some sort of reseller refund abuse happening where they're trying to steal your compute.

And so it's and and it becomes more complicated when you add more entities like the merchant, where there can actually be internal collusion happening. And so you're exactly right. Like, it's it's not about how any one entity acts in isolation.

It's about, like, an entity is a an individual or a card or a merchant, and then that's, the node. Right? And then the edges are the transactions.

And it's, how how much sense do these edges make in relation to each other and in relation to the combination of edges that we've seen in the past?

Speaker 3

Does that go out I can imagine that that web could extend easily farther or you could sort of imagine including sort of the rendered judgment on previous transactions. So for example, if I am trying to buy something from you and you're and the model is looking for the signal of fraud, you could also say, okay. Well, all these transactions that you have recently done as a seller, maybe you just sort of have the, you know, the determination if they were fraud or not fraud.

Or you could even look at, okay. Well, who who are all those buyers? I'm like, what's their so how how far out does this sort of path through the graph have to go to get you the you know, what is the what is the sort of shape of the curve in terms of scale versus diminishing returns?

Yeah. Totally.

Speaker 4

so I talked about the scale of the Stripe network. Right? It's like this $1,400,000,000,000.

But it's not just a big network. It's also a very dense network. So for example, 92% of cards that a merchant sees for the first time, Stripe has seen before on another merchant.

So okay. Well, in those cases, you don't have to do very many hops. Although, you do wanna validate that nothing's changed about the card or how the card's being used in the time since.

But fraud and conversion are kind of tail events in some sense too. Right? Like, if you can get 1% more conversion or one or 2% less fraud, that goes a long way.

So you get really far from the dense network, but you also wanna be able to traverse wide for more novel traffic that you see.

Speaker 3

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 2

AI's impact on product development feels very piecemeal right now. AI coding assistants and agents, including a number of our past guests, provide incredible productivity boosts. But that's just one aspect of building products.

What about all the coordination work like planning, customer feedback, and project management? There's nothing that really brings it all together. Well, our sponsor of this episode, Linear, is doing just that.

Linear started as an issue tracker for engineers, but has evolved into a platform that manages your entire product development life cycle. And now they're taking it to the next level with AI capabilities that provide massive leverage. Linear's AI handles the coordination busy work, routing bugs, generating updates, grooming backlogs.

You can even deploy agents within Linear to write code, debug, and draft PRs. Plus, with MCP, Linear connects to your favorite AI tools, Claude, Cursor, ChatGPT, and more. So what does it all mean?

Small teams can operate with the resources of much larger ones, and large teams can move as fast as startups. There's never been a more exciting time to build products, and Linear just has to be the platform to do it on. Nearly every AI company you've heard of is using Linear, so why aren't you?

To find out more and get six months of Linear business for free, head to linear.app/tcr. That's linear.

app/tcr for six months free of linear business. Build the future of multi agent software with Agency, a g n t c y. Now an open source Linux Foundation project, Agency is building the Internet of Agents, a collaboration layer where AI agents can discover, connect, and work across any framework.

All the pieces engineers need to deploy multi agent systems now belong to everyone who builds on Agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent to agent communication, and modular components for scalable workflows. Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and 75 more supporting companies to build next gen AI infrastructure together.

Agency is dropping code, specs, and services, no strings attached. Visit agency.org to contribute.

That's a gntcy.org.

Speaker 3

this sort of reminds me a little bit of some of the stuff that Meta has done with their, like, joint embedding video models. I'm not sure if that is the right intuition for me to have, but it does seem like there's a clear difference here where you're not trying to predict the next token. Right?

It's it seems like it would be more of a dedicated so it's not like it's not like a autoregressive type model. Right?

Speaker 4

like a masked type situation where you you could you could imagine doing a training setup where it's like mask out whatever kinda randomly and have the model learn to fill in details. Yeah. Exactly.

So our v one did use a sort of BERT style, like, masked modeling setup that you're talking about. And then we paired that with a second stage, which is, like, explicit similarity fine tuning. And so most of the heavy lifting there was like, okay.

Let's curate the right sequences to learn from. Let's build the right encodings. Let's do post training with that kind of similarity objective so that near neighbors in this payment space cluster together and the oddballs separate, and you can start to reason about those oddball clusters.

And, again, the big unlock here was modeling short histories. Right? So what a card or device or merchant bin is doing over some number of minutes or last k transactions rather than the the isolated payment.

Now in the we call it v 1.5, but we're actually moving towards encoder, decoder setups and compressed memory sequence. So, like, a few vectors together, that makes it actually easier to catch subtle abuse in real time because you're not averaging across noise.

You're distilling the full story into this, like, comp compact representations. And so, yeah, the mental models like v one, mass modeling, plus similarity training, v 1.5 is compression first with a tight sequence embedding.

And then we can put kinda lightweight task specific heads on top, which are for the charge path use cases. Like, you know, if you think about the charge path, you gotta get the job done in tens of milliseconds at most. And so those lightweight task specific heads are are important for latency and speed.

Speaker 3

Yeah. Can you say how big the model is? I mean, ten milliseconds doesn't allow it to be that big, I would assume.

Although, they don't have to do a lot of steps in forward pass. The cuff, but the task specific heads on top are small.

Speaker 4

And, you know, I think when you when you reason about it, like, all you have to be able to do actually is place the new charges and sequences as they come through in this dense embedding space as they come in, which is a much easier problem than obviously, like, the upfront training.

Speaker 3

Yeah. And it's also just one forward pass of the the model, right, as opposed to having to generate a whole sequence. So you can that that first that one pass, nature of it definitely helps with the latency as well.

Yeah. That's really interesting. It also reminds me of one of the first language vision language models that I studied deeply was the Blip family of models.

And I remember that they had really amazing success with a frozen language model and then also a frozen vision model and just trained a few million parameter connector between the two to sort of bridge from one latent space to the other. And, of course, we've gone way past that now in vision language, but this was, an early twenty twenty three thing. And you were able to get, like, really quite good captions out of that setup even though neither of the foundation models that were used had anticipated that use case.

Speaker 4

Yeah. So it sounds like you're Lipping a while, but that's, yeah, that's an interesting that's an interesting analog.

Speaker 3

So it but it sounds like you've kind of created a similar situation where people internally at Stripe can say, okay, I have a new use case idea for this. I can train something really small. You mentioned it can be like a weekend project instead of a, you know, multi month project.

Speaker 4

like, very quickly. Like, that that step becomes a rapid iteration step. Yeah.

Exactly. And actually, you know, it can be even simpler, which is so so the embeddings themselves where most people start actually is most most modelers start is just taking the embeddings themselves, which are stored in Shepard, which is our feature engineering platform, and literally just, like, adding them as a feature to existing models, like, those existing models are, and saying, like, is there added signal from these embeddings? And I would say you get some false negatives there where, obviously, just, like, shoving the raw embedding into some number of, you know, SOTA models that have been, you know, iterated on over the last six years isn't gonna produce uplift.

But in other cases where it's sort of a lower priority model that's only in its v one state, like, you actually do get something straight out of the gate, and you can start to reason about, you know, we're we've been talking much about the payments embeddings. Like, how much signal do I get from, like, understanding the payment better? How much signal do I get from understanding the customer better?

How much signal do I get from understanding the merchant better for each of these use cases? And then that's also motivated where folks have kind of for which applications folks have leaned in harder.

Speaker 3

Yeah. That's really interesting. So just to make sure I understand that correctly, you've got, obviously, stress around for a number of years.

There have been many types of problems that you've brought machine learning to over time. Typically with a more classical feature engineering type of approach. Yep.

Speaker 4

Hundreds of production models. Right? Like point point solutions.

Speaker 3

And so now the foundation model embeddings can become just tacked on as additional features, rerun that training, and immediately back test against your set, and then you're like, okay. Cool.

Speaker 4

we were able to get additional signal. That's, yeah, really interesting. I mean, our This is standout application wasn't that.

Right? It was card testing where we literally, like it was a whole new approach to card testing with the foundation model, which, you know, I think card testing is like, you know, fraudsters trying hundreds of tiny authorizations, of iterating across stolen cards or literally just doing raw enumeration, like trying a bunch of cards. And they bury those attempts inside floods of legitimate traffic.

Right? A big retailer can have hundreds of thousands of charges come through, and then there's, like, a couple 100 or maybe a thousand peppered fraudster charges of, like, 30¢ or 50¢. And, like, classic models couldn't really pick up those kind of needles in the haystack.

And so our first application of the foundation model was just treat those sequences, again, like frames in a movie. Right? And suddenly, these 200 nearly identical requests, like, you know, same low entropy user agent rotating across proxies coming in about every forty seconds, like, they light up as an island in the embedding space and they get blocked.

And the impact of that one was huge. Like, our detection rate of car testing at large merchants went from 59%, which is, you know, not bad, but not great to 97% from that change. But then as we started reasoning about where else could it be useful, yes, just expose the embeddings and let them be added as features to these sort of traditional single task models was was the was the next step.

And that was never intended to be the final state, but it's a way to get signal on where is there sort of incremental value or incremental signal from these embeddings that requires very little lift.

Speaker 3

Yeah. Fascinating. That's a very modular approach to AI deployment, and I can't recall hearing any organization that has had a similarly modular structure.

Maybe Meta comes to mind as another one that might have a sort of user model that, you know, could then be bridged over to any other space or problem that you might wanna apply them to. But this is, like, a fairly uncommon setup, I would say. Right?

Speaker 4

that was basically doing horizontal models. Again, don't know don't know exactly how it worked or the architecture. You know, we're we've been talking about it in the context of payments, but we actually did we've done the same thing over the last year in the merchant space where we have it's called the merchant intelligence team, but it basically has MI serve.

So it it can go out and find anything in the in the web about a merchant and generate embeddings and be used to answer questions. And those merchant embeddings are also features in downstream, for example, merchant risk models. But it's a service where the model owner can ask merchant intelligence, the agent, to come up with a more custom embedding or a more custom insights.

So maybe you wanna know what payment methods the merchant offers or whether they have anything that's counterfeit. And and that's actually been another horizontal layer that's provided a ton of leverage for Stripe. Right?

Because historically, you know, you got a lot of use cases. You wanna know things about the merchant to understand supportability, like, you know, whether they they meet the requirements of the card networks and the issuers and the banks. You wanna know whether or not they're fraudulent.

You wanna know whether they've had an account takeover. You wanna know if they're creditworthy. You wanna know whether or not we should give them Stripe Capital, like, a loan, and you wanna figure out whether you should be going to market with them.

Like, all sorts of things you you wanna know about a merchant. And, historically, teams at Stripe were, you know, when when LLMs hit the scene, like, out building their own sort of custom versions of this. But what we realized is there's actually just, like, one service now that that does that much more efficiently than everyone rolling their own.

Speaker 3

The better lesson strikes again. I got a lot of different directions I wanna go, but where does ground truth come from on some of these questions, and how long does that take? Because I I sort of imagine, especially in a fraud detection situation.

Right? I mean, fraudsters, I always assume are gonna be some of the most clever people in the world diabolically so. But nevertheless, like, you gotta respect the, the smarts of some of these folks.

Right? So I assume that they are very savvy to, like, real world events. You mentioned, I think, the conversation with Patrick that, you know, somebody might have a flash sale and that sort of spike.

Like, obviously, you don't wanna turn them off when they're having a flash sale because that's a, you know, horrible experience and loss of business for the company running the flash sale. But at the same time, like, that's sort of potentially a really good target for a card tester to come in and try to do whatever it is they wanna do. And you're sort of, I imagine, in a kind of eternal arms race between fraud and fraud detection.

Speaker 4

set in stone. Like, that's a long process. Right?

So Well, it's a long process if you even get to a definitive answer. But something like card testing, right, like, actually, I said the first thing we did with the foundation model was deploy it for card testing. Actually, the first thing we did for the foundation model is deploy it internally for card testing, pass those labels to internal expert humans, have them go and validate the labels, then feed the validated labels into our traditional ML model for card testing.

And suddenly, our ML model traditional ML don't deploy the foundation model. Our traditional ML model for card testing started doing way better because finally, it had, like, a more comprehensive source of truth for the labels. So it was actually actually the first version, although hadn't hadn't hadn't revealed that fun fact before.

Yes. Attackers iterate and so do their models. And so our job is just to iterate faster.

And we are, and I'll talk about some of the ways we we get around the the late arriving labels or the missing labels altogether. But just to give you a sense of, like, how we're comparing to the fraudsters, like, industry wide, ecommerce fraud is up. I think it's up, like, 15% year on year.

But the dispute rate for the businesses that are running on Stripe are down 17% year on year. And that's because we, in a bunch of different ways, and I can give a couple of my favorite recent examples, are just consistently shortening the loop between, like, new tactic shows up and, like, defenses go and adapt. And that sort of loop shortening is happening in production and in some cases in real time.

So an example, that that our users are getting a ton of value from that we recently released is dynamic risk thresholds. So it's basically like, you know, radar's out. They've got their threshold score.

Block stuff above the threshold. But then when an attack starts, radar learns an attack starts, and it tightens the defenses. Right?

So it it kind of throttles. And that allows, like, okay. You know, revenue is flowing freely when you're not under attack, but then we're much more aggressively blocking when an attack arises because, again, like, an attack is never almost never a single event.

It's almost always like a a true a true cluster. And in that case, like, you know, the the model is learning the policy of how to act. Now it's not learning that policy online just yet, but it's learning the policy of of how to act.

Another powerful tool, and I think, you know, in payments, it's easy to think, like, I put in my credit card and then just, like, an objective decision is made to block me or not, But that's not actually true. And so we've been leaning in harder on what we call soft blocks. So adaptive three d s is an example here.

It, like, applies that three d s authentication. So, like, if you're in The US, most of the time, you don't get you don't get three d s. But we can Can you tell me what that is?

Because I don't feel like it. You might have defined it in the in the complex systems, but if if so, I could use a real question. Yeah.

You're just you're just like you have, like, a second like, a it's sort of like a two factor off that you would have. It's like a two factor off esque experience where you're verifying that to the credit card the bank network or the credit card issuer that it is in fact you. And this is very, very common in Europe, it's very, very common very, very uncommon outside of Europe.

And by the way, when it does happen, it often creates unnecessary friction. And so part of what we do at Stripe is figure out when we need to authenticate and when we don't. But also, with Adaptive three d s, we are pushing for authentication selectively in cases where we have a sense that the charge may not be good.

And so instead of just having this binary decision of block, don't block, you have this other arm you can go down, which is, you know, hit them with three d s. And what ends up happening is the good guys get through the three d s because they're excited to buy the thing and they're legit users, and the bad guys do not. And so a lot of the AI companies are using this.

You know, AI companies being hit with fraud is extra painful because their marginal costs are high, right, unlike for SaaS companies who care a lot care a lot less. And so, like, early adopters of adaptive three d s were, like, 11 and Character AI, and they're able to just dramatically cut down fraudulent disputes without any effect on conversion because, you know, three d s isn't isn't super heavyweight for the end user. In fact, US checkout users so it's it's a little different in Europe because a lot of those folks are already three d s, but US checkout users saw a 30% average drop in fraud, and they just, like, turn this on with a single click in the dashboard.

And then it lets us basically, like, learn the policy of who is worth 3 d s ing to balance conversion and fraud to maximize their profits.

Speaker 3

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 2

Today's episode is brought to you by Anthropic, makers of Claude. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you, not for you.

Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Regular listeners know that Claude plays a critical role in the production of this podcast, saving me hours per week by writing the first draft of my intro essays. For every episode, I give Claude 50 previous intro essays plus the transcript of the current episode and ask it to draft a new intro essay following the pattern in my examples.

Claude does a uniquely good job at writing in my style. No other model from any other company has come close. And while I do usually edit its output, I did recently read one essay exactly as Claude had drafted it.

And as I suspected, nobody really seemed to mind. When it comes to coding and agentic use cases, Claude frequently tops leaderboards and has consistently been the default model choice in both coding and email assistant products, including our past guests, Replit and Shortwave. And meanwhile, of course, Clawd Code continues to take the world by storm.

Anthropic has delivered this elite level of performance while also pioneering safety techniques like constitutional alignment and investing heavily in mechanistic interpretability techniques like sparse auto encoders, both internally and as an investor in our past guest, Goodfire. By any measure, they are one of the few live players shaping the international AI landscape today. Ready to tackle bigger problems?

Sign up for Claude today and get 50% off Claude Pro, which includes access to Claude code when you use my link, claude.ai/tcr. That's claude.

ai/tcr right now for 50% off your first three months of Claude Pro. That includes access to all of the features mentioned in today's episode. Once more, that's claw.

ai/tcr. In business, they say you can have better, cheaper, or faster, but you only get to pick two. But what if you could have all three at the same time?

That's exactly what Coher, Thomson Reuters, and Specialized Bikes have since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster?

OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50% less for compute, 70% less for storage, and 80% less for networking.

And better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is the cloud built for AI and all of your biggest workloads.

Right now, with zero commitment, try OCI for free. Head to oracle.com/cognitive.

That's oracle.com/cognitive.

Speaker 3

So one way I like to frame, excuse me, some of these conversations is just in terms of, like, practical lessons that people can apply in their own, AI pursuit. So one takeaway there is add middle ground outcomes to your classifiers so that they're not binary, but try to find that sort of middle space if one exists that can where something other than the model itself can can step in to help resolve the most challenging cases. It's almost like Claude, you know, now these days can sometimes end the conversation.

Right? If it if it has to Fundamental information could help you and you can get it from your users in a low cost way.

Speaker 4

Don't constrain yourself to being a modeler. Like, be a product thinker and go figure out how to get that information. Right?

And the model's really good at deciding, and then that that's not brute force. You don't require that additional information from everybody, but then let the model decide where it needs more information and where it doesn't.

Speaker 3

Yeah. On the adaptive threshold concept, this suggests a state, basically, a sort of state world state that is maybe being fed into I assume it's not like that the model itself is calculating that on the fly. This would be a more, like, global variable sort of thing that the model would receive or No.

So so it's actually like, hey. This merchant is starting to see, like, clusters of scores creep up. Like, isn't that isn't that interesting?

Speaker 4

And when we look at the subset of transactions that have those higher scores, maybe they're still below the block threshold, but they're looking elevated. Is there anything about those that looks like it's something collusive or, you know, coming from a small number of attackers or some rotating across IPs or coming from a geography that they haven't seen before. And then once we get signal that, like, looks like there's a slice that's an attack, you can actually, like, start to lower the threshold from that subset for what it takes to block.

Speaker 3

it just sounds like there's a longer history at some point coming in to inform that kind of decision, but maybe not I guess the short history could be enough.

Speaker 4

but less at the individual transaction level. Like, it's it's basically detecting anomalies in slices of traffic. Right?

So, like, this geo, this these bins, this cart size, like, something anomalous is happening here. That anomalous thing has kind of elevated risk scores. Hey.

Like, it reads a bit like an attack. And, actually, what's interesting I mean, rules are really good in a lot of ways. Right?

And so, you know, maybe another general lesson is, like, rules are good, but they're also blunt. So figure out where you can blend rules with models. I mean, you'd asked earlier when disputes actually come in.

Like, disputes are super lagged. They can take days. They can take months.

Right? I'm the cardholder. I have to, like, see my bill, like, notice I didn't buy the thing, tell my bank.

My bank has to go and, like, file with the network. And so those labels for sure arrive late, but we don't wait. We use proxy signals, and those sort of weak labels show up way earlier all the way to real time issuer feedback.

And it can be very so real time issuer feedback would be like, CVC mismatch. Right? Like, the the CVC code, you know, the the little three, four digit credit card code doesn't match or, like, the the ZIP code doesn't match.

It'd be easy to write, like, a blunt rule that said a CBC doesn't match or if the ZIP code doesn't match, block it. But you'd be blocking a bunch of good revenue because, like, who doesn't sometimes fat finger their CBC or their ZIP code in a hurry or on their phone or whatever. And so we have these risk based radar rules, which are like, okay.

Take the model score, combine it with the issuer's real time responses, and make a decision based on that intersection. So, like, if it's looking marginally risky and the CVC is wrong, for sure block. But if it's like a pretty known good user and they fat fingered a thing, like, let them through.

And I think that blend of rules and models is I think it's easy for modelers to put their nose up at rules, and it's easy for rule makers to put their nose up at models that aren't fully explainable. But in plenty of context, blending the two actually does does far better.

Speaker 3

So you actually do let transactions go with a wrong three digits? What was that called? City soup.

Yeah.

Speaker 4

Nathan, like, I, like, I know you're good. You've bought from this person before. Maybe you even use the same credit card.

You're coming from a legit IP. And, like, I I feel good about you in a lot of ways. And, yeah, the issuer comes back and says, hey.

There's a mismatch. And we say, hey. Let it through.

And then by the way, once we let it through, we also have to get the issuer to let it through. And there, we actually have data sharing with the issuers where we pass them our risk scores so that they can also understand why we passed it through, and that motivates them to also pass it through when they when they see our signals. So it's kind of a two step.

Speaker 3

Very interesting. Let's go back to how you are tightening the iteration loop. Again, I think this is something that basically everybody who's developing AI products could stand to get better at.

So what have you guys found to be effective needle movers in shortening your cycle time?

Speaker 4

mean, this this isn't one for us where it's like there's some magical, you know, reinforcement learning that we need to be implementing online for every single use case. I think it's actually for us been quite context dependent. The things that matter are having enough labels and having good labels and having those labels fast enough.

And, actually, like, you can get pretty creative about what the label is. We talked about some examples. We also talked about human generated labels.

But another thing that we've been leaning into is LLMs as a judge. So especially for context where there actually is is no source of truth. So a simple example, we've been talking a bunch about fraudulent disputes, but there's a lot of suspicious payments that never result in a fraudulent dispute.

Right? Maybe the person starts a free trial and then they cancel, or they ask for a refund, or they, you know, just spin up a bot account but never even get to the checkout page. That type of friendly fraud is actually really costly to businesses, and it's like almost half a business.

Think 47% of businesses say friendly fraud, which is a total misnomer because it's not friendly, hurts them more, hurts their business more than sort of stolen card credentials or what most people think of as fraud. And that's actually you know, there's a lot of AI companies running on Stripe. That cost of friendly fraud is particularly true for these AI companies.

Right? Very different than SaaS. Again, they have they have inference costs.

They have compute costs. Therefore, they have, like, very high marginal costs. And so when someone is engaging in free trial abuse or reseller abuse or refund abuse, it's super expensive to their their unit economics.

Anyway, so built on the foundation model, we now have these suspicious payments that we identify. So these are fraudulent e things, but not in the traditional going to result in a fraudulent dispute sense. And when we pass those over, for example, to the AI companies, we wanna be able to describe to them why they're flagged as suspicious.

So, oh, it has an enumerated email or it is, like, cycling through small number of IP addresses or whatever. So those labels are so sort of those explainers are generated by the foundation model. But then the question, of course, is like, well, how do you know if they're right?

And so we have this, like, LLM as a judge that sits on top that looks at every transaction label combination and asks, given everything you know about this transaction and everything you know about the cluster to which it belongs, how do you feel about the quality of the label? And what ends up happening is that there's, you know, a a large share of labels that are good enough, trustworthy enough that we pass them over to name AI company du jour to decision on. And there's some small number that are, like, too noisy, and we're like, okay.

Like, we gotta go work a little bit to make that label stronger. But I call out that example because these are like there's no source of truth. Like, I like, could I mean, you and I could manually go through, I guess, but at the transaction level, we're not going to.

And so it's been really helpful to have LLMs kind of as a judge where there's where there's no where there's no clear north star.

Speaker 3

Yeah.

Speaker 2

so

Speaker 3

that's fascinating, but I'm I'm still kind of confused about one thing, which is well, I'm probably confused about a lot of things. But the thing I'm focused on being confused about right now is when I try to advise people on AI broadly or when I kinda try to give people the lay of the land, one of the things I tell people is AIs are not very adversarially robust. They are, you know, really good these days at the happy path.

If, you know, dial in the performance and, you know, you control the inputs, you can in many, many cases, you can get to superhuman performance on routine tasks. However, if you don't control the inputs and you're exposing your AI system to the world, you do have to be mindful about the fact that the these systems are not adversarially robust. People can usually find some, you know, some weakness.

Right? And that's even been true. We did an episode once on superhuman Go playing AIs that were beaten by really simple attacks that, you know, no human would ever fall for, but which the AI, even though it was superhuman when playing Go in the normal way against, like, other, you know, high quality Go players, it was just totally blind to this certain class of attack that was found through this sort of adversarial optimization.

And so it seems like you would be in this environment where, you know, you've got, like, you've got it on, like, hard mode kinda everywhere. Right? Because anybody can come test the system from kind of any position.

You can't really deny people the ability to, like, try a payment. So they can kind of gray box you. Right?

They can test from a bunch of different angles and try to see, like, what's gonna get through, what's not gonna get through. And presumably, there's always some, you know, vulnerability that you're not aware of that they can systematically attack or or try to find through these these sort of attacks. And my guess would then be, like, the only way to really deal with that is to just constantly be identifying and iterating.

But that that sounds still hard. Despite everything you've told me, it still sounds hard to be as responsive as you would need to be given especially that the actual ground truth is so lagging. So, like, how do we not maybe we do.

How do we not, like, just bleed a ton of money in one incident after another as attackers figure out that there's, like, some gap and then just, you know, jam as much as they can to exploit it for a while, until it's closed? Like, how does that not end up being a huge problem? Yeah.

Speaker 4

A couple a couple thoughts. So one is we expose capabilities through products and APIs to our users, not through raw weights. So for sure, anybody can try a payment and test and see if they can do a workaround.

But just to be clear, like, we're not actually exposing the model for them to for them to, you know, have an attack surface against. Right? So, I think that the products and the APIs actually better meet the user needs, and they also kind of narrow the potential attack surface.

So just clarification one. I think for you know, it's it's interesting to think, like, what's the relevant alternative? Like, what is fraudster's job?

Fraudster's job is, like, find loopholes and exploit the system. Like, that's that's what they make their money on. And so the relevant alternative isn't, you know, perfectly airtight.

The relevant alternative is, like, baseline approaches. And, actually, when you start to think about foundation models or LMs, they you know, the payments foundation model, for example, like, it's actually a lot more nuanced, the the type of information that it's using to decision versus, like, you could think of, an early an early transaction fraud model that's using, like, last seven day counters. And, like, the fraudster figures out that, like, as long as I'm eight days out, I'm safe.

I'm just gonna do everything, like, on day eight and then hit them hard and then go seven days back. So so to some extent, I think sort of traditional ML is is easier to get around, whereas the the founder's model is more comprehensive. But the other thing that we certainly have long done and continue to do is a layered approach.

So it is not there's not just a single set of defenses. There's a set of model defenses. There's a set of rule based defenses.

There's the soft blocks I mentioned. There's the user's own defense set, which can also vary, like, all the way to to how they treat you at sign up or how they block bots at sign up. And so, fortunately, for us, unfortunately, for the fraudsters, they're not fighting against, like, one model.

They're fighting against a whole system that is, I guess, until I said it, opaque to them.

Speaker 3

Yeah. Why don't you tell us just about kind of any other big use cases? You mentioned, you know, like, off fraud disputes.

There's, some interesting talk to your data product experiences in Stripe. You know, what what stands out to you as the most interesting applications? Not even necessarily from a, like, what moved the most money, but, like, what would be most interesting to the AI engineers audience in terms of just interesting implementation details or surprises, you know, quirky stuff that you've learned along the way?

Speaker 4

Yeah. I mean, we've we've talked a lot about sort of transaction level understanding and, you know, the the path there, if you think about, like, modality, like, it's it's mostly payments plus text. You have these, like, structured payment signals with and then, like, language using contrastive learning, and you align the two, and you got the text decoder and whatever else.

But payments plus text is, like, only the start. Right? So the system's actually designed so that new modalities are just considered, like, tools that the router on top can invoke.

And so, like, if you wanted to add another encoder maybe for financial time series, which I'm very interested in, but I don't have anything yet that I could share. Or for images. Right?

It doesn't require, like, a whole kind of rewriting of the system. It's just like a modular expansion. And I think the multimodality like, I'm starting to see it really shine at the merchant level.

So, actually, yesterday, I was testing two lightweight agents. Neither of these are in production. So, well, you know, just just full disclosure.

But, like, the team, has them in shadow, and one crawls merchant sites to assess fraud, and it's relentless, like, incredibly relentless. The other spots counterfeit products, and it just does so, like, literal orders of magnitude better than trained human reviewers that we have at Stripe doing the same thing. So it'll find, like, there's a print shop and there's thousands of items in the print shop and the agent will, like, patiently zero in on, like, the one, like, spider Gwen sticker was, like, an example I was staring at yesterday with, like, no sign of of official licensing, like, rut rut.

Then it also knows, like, oh, this this other site, the Canada Goose, that's, like, marked as with tags is secondhand, and so it's actually it's actually fair game. So I think that kind of, like, that kind of multimodal road map is is interesting. Not for a multimodal in and of itself, not for the technology in and of itself, but for where it's, like, gonna unlock gonna unlock real value.

Speaker 3

Cool. On the talk to your data thing in particular, you know, that that's something that I think a lot of people have tried to do for themselves or they've tried to use a product to do it. It strikes me that the where most people have kind of gotten stuck there.

Right? It's like, well, I was able to get GPT whatever or Claude whatever to be, like, pretty good, but it still made some mistakes. And I didn't really feel like I could confidently give somebody who wasn't a proper data analyst this this tool and, you know, be confident that they would get good insights out of it.

So you guys have that problem, you know, at maybe the biggest scale in the world. How did you think about, like, what is the right threshold of accuracy for I talk to your data model? I assume you you didn't achieve a 100% accuracy on this sort of thing.

But what was the threshold that you felt you had to get to, and what was needed to, you know, to keep dialing in until you actually got over that threshold to where you could deploy?

Speaker 4

one of the reasons that, you know, this sort of talk to your data was interesting to us in Stripe's context is, like, a lot of what a business wants to know is captured in Stripe data. Like, you know, who's selling what, for how much, to whom, like, who's retaining and churning their subscriptions, etcetera. And so that's thing one.

Then thing two is the data is actually, like, very well structured because it it has to be. Right? Like, it's it's generated, you know, from the transactions that are flowing through Stripe that are incredibly, like, robust and well documented, and the schemas downstream of that make sense and are well documented as well.

And so, you know, a lot of this talk to your data stuff, it's like the garbage in garbage out problem where, like, my my tables aren't well labeled, my fields aren't well labeled, maybe, like, the underlying data actually isn't deduped. And so you can't really tell if the issue was, like, that sort of text to SQL or if it was actually, like, the underlying data was bad and or the data structure was not understandable. So we kind of were able to, like, leapfrog that, which is great.

But still, LLMs so you're I think you're referring to our Sigma system. LMs do make mistakes. And so our approach there is actually, like, if we have reasonable confidence, we'll provide it, but we I don't know if you've ever used it.

We overlay on top a natural language explanation of what we're doing. So, like, we thought you wanted to know whatever. You asked, like, you know, how did Black Friday this year compare to Black Friday the last two years?

And then they'll be like, okay. These are the dates we use for Black Friday. These are the time stamps we use.

Because by the way, like, most things on Stripe happen in UTC, and many people aren't reasoning about their business only in only in UTC. We looked at Black Friday over the last three years. Here's how we computed the percentage growth.

Like, you're like, that's boring. Doesn't everyone compute the percentage growth the same way? But that actually allows someone who's not a data analyst to build comfort in in in the output versus saying, like, either YOLO ing it, just like taking it and running with it, or throwing their hands up and saying, can't I can't trust anything because I don't know what's happening under the hood.

Like, you just wrote a SQL query for me, but I have no idea how to interpret it. So, you know, I think when I think about Talk to Your Data, it's like, well, is your data interesting to talk to? If yes, like, make sure it is well structured, well documented.

And if it's not, invest in that before you invest in, like, you know, the natural language interface on top, and then just make sure the natural that the LLM is, like, explaining what it's doing with which they're, of course, very good at doing now. And that allows you to open the aperture a bit in terms of, like, less certain questions you're willing to answer because you know anyone can read the natural language and make a call on on whether or not that was the the right approach.

Speaker 3

For folks who wanna do a a double click on the process of getting the data into shape, the episode with the CEO of Illumix was really good on that. And just for what it's worth for you, they have built basically canonical structures of enterprises across, like, a bunch of different categories, like, you know, drug company, for example. They've kind of built out a vast representation of data that in their, you know, studied opinion represents, like, the canonical drug company.

And then when an actual drug company comes to them, they do this painstaking process of mapping all of their actual data with all of its idiosyncrasies onto the canonical version that they've kind of made work well, and that mapping becomes the kind of cleanup process that gets them, you know, the reliability that customers obviously ultimately want. Pretty interesting. That's capability Yeah.

Yeah. Please.

Speaker 4

there's also kind of an interesting feedback loop here with users. Right? So if you're a usage based billing company, the types of metrics you wanna know to reason about your business or to share with your investors are generally very similar to the types of questions that all the other usage based billing AI startups also wanna know.

And so that both means that we can really make great the subset of questions that matter in a given domain. But, also, you know, forget the natural language to SQL interface or talk to your data. We can just push those commonly asked questions over onto the dashboard and even benchmark you.

We have smart benchmarking now. Benchmark you on those metrics versus a peer group. And by the way, that smart benchmarking is one of the applications of the, merchant intelligence service, which is, like, figure out which websites are like this website in terms of would be good comps because have similar have similar user bases and are at a similar stage of their development.

Speaker 3

Yeah. Cool. There's a good pattern there as well for sure.

I've been thinking about that in the context of agents lately, and there's kind of the choose your own adventure agent, you know, where you, like, give it a bunch of tools, here's some MCPs, whatever, have at it. And then there's the, like, sometimes better described as a workflow, maybe with a, you know, a couple forking, you know, decision points that people also call agents in many cases. And I'm starting to see the emerging pattern be like, have that choose your own adventure agent sort of at the top level of user interaction.

But then in terms of the things that it's choosing, like, make those actually, like, pretty detailed workflows in a lot of cases where you know that as long as it makes the right choice at a high level, that the process that's gonna be kicked off is one that you've really deeply understood, you know, dialed in for accuracy, confirmed for yourself is that is gonna work, you know, reliably. So I think that's another you're kind of talking about a push model instead of a pull, but nevertheless, that there's a sort of isomorphism, I think, between those structures. Explainability is obviously huge.

One thing I was interested in asking is, are you doing any mechanistic interpretability? Are there, like, sparse autoencoder type things now happening on the foundation model so that you can, like, learn in a semantic way, like, what new features the thing is learning? Yeah.

Not literally.

Speaker 4

like, individual neurons in the way some research groups are. I think our focus is really on making the outputs self explaining in the way that we and our users, where it's user facing, can actually trust. And so, you know, mechanistic interpretability is really important when, like, you're releasing the full open ended model into the wild.

In our case, we control both application and the environment, and so our priority is, like, that really practical explainability that's tailored to the payments application. So, like, when the foundation model flags a transaction, it doesn't just say high risk. It says, you know, gibberish email and enumerated name pattern and device concentration.

And that's actually the layer of explanation that lets the fraud analyst or even another system, like a follow on agent, act confidently. And then as we're talking about a little bit ago, like, in many of those cases, there's actually no ground truth for that explainer. And that comes up more and more as we expand to new domains.

Like, oh, we're we're actually detecting fraud further up your customer funnel, all the way when someone's, like, creating an account on you and well before they're entering credit card details. And so that's where things like the LLM as a judge framework are really helpful. So it'll, like, look at that and the cluster of similar events and the tag definitions and then output, like, how confident is the LLM basically judging how confident is it that the cluster really matches the label.

So, no, we are not, like, peering in neuron by neuron, but we are focused on interpretability at the output level, and that's really valuable for us. And then, you know, for example, like, you're in your dashboard and you're seeing a bunch of suspicious users, you wanna know exactly why we flagged them as suspicious so you can decide how to action. And so in our setting, that's that's really what matters most.

Speaker 3

Gotcha. Okay. You mentioned usage usage based billing.

And this led me to a very practical question around how you would recommend people build on top of Stripe today. Ten years ago or so when I first became a Stripe customer, it was already a respected company, but not, you know, such a foundational part of the economy as it's become. So we were like, well, we don't wanna switch off of this one day or who knows, whatever.

We'll like have kind of our own database of all the transactions and all that kind of stuff. And, you know, Stripe will of course have their view of it, but we'll like maintain our view. That was a lot of work then With usage based billing, it sounds like an even more, you know, challenging project now, especially for your proverbial, you know, couple people that are doing a hackathon and wanna kinda get something started.

Yep. Do you recommend that people so the alternative I have in mind, which I wonder, ultimately if you recommend is like, could I just leave all of that to Stripe and just basically make nothing but API calls, trust Stripe to be like real time ground truth across the board and like not even have a sort of financial side to my database, but just purely do that, like, real time API calls?

Speaker 4

Totally. Totally. Don't even have it.

Like, Stripe's Stripe's APIs run at six nines of uptime. Right? So they are safe to use as your system of record for, like, very critical flows.

Our usage based billing APIs process 100,000 events per second, and they've got all the built in monitoring and alerting and invoicing. The alternative is also, like, pretty painful. Right?

Like, building if you're gonna, like, build your own mirror of Stripe's data, it's pretty complex. You gotta sync across all the events. You gotta build your own monitoring systems.

You gotta keep everything reconciled. And especially if we're talking about, like, a start up, that's just, like, a lot of work that doesn't create any kind of differentiated value. Right?

A lot of these companies are taking off have, like, five, ten, 20 employees. Like, they shouldn't be spending an ounce of that of that limited capacity on this stuff. And then on the flip side, you know, if you if you treat Stripe as your source of truth, you get real time signals that you can actually act on in the Stripe ecosystem.

Right? Like, the, like, the billing threshold has been exceeded or whatever without having to have this whole parallel system. And we talked about Sigma system earlier, but with products like Sigma and Stripe data pipeline, you can still run all your analytics and all your reporting sort of without doing the job of of building your own your own warehouse.

You know, you you might be wondering about the downsides. Like, historically, the biggest downside of not mirroring, of like, just leaning into Stripe was, okay, but what about when I wanna join, basically, Stripe data to, like, my own business objects? And now you actually can.

So you can actually extend Stripe's objects with metadata. So, like, a lot of users will whatever. They'll, like, attach their own order ID or shipment ID to an invoice.

And for, you know, most companies, that closes most of the gap. Now it's different if you're a large enterprise and you've got a whole bunch of other follow on systems that are off Stripe and data sources that are off Stripe. But for most startups, it's both simpler and just a lot safer to let Stripe be the system of record.

Speaker 3

Yeah. Cool. I imagine some companies have gotten pretty big, revenue terms over the last however many months, while still doing just that.

I don't know if you would wanna highlight any by name or if that's, too secret. But when I see the curves, you know, from folks like Lovable, Bolt, Replit recently, obviously, things like Cursor, one starts to wonder, you know, in the head counts that they have.

Speaker 4

trusting, which six nines gives you pretty good reason to trust. Yeah. And, you know, that's just for the system of record.

Right? So Lovable, it's a great example. They hit a 100,000,000 in ARR in their first eight months, and their stack is basically a a case study in all in on Stripe.

Right? So they they incorporated the business. Before they monetized, they incorporated the business with Stripe Atlas.

They, from the very beginning, used our optimized checkout suite. So, like, the the front end customer facing services is our optimized checkout suite, which allowed them to localize payments in over a 100 countries and get something like a 150 payment methods out of the box. They leaned on billing for subscriptions, so they didn't build their own billing system or, you know, have to contract with another third party.

They leaned on Link, which is like our our one click consumer checkout for fast checkout. By the way, nerds love to buy from nerds. So, like, our concentration, on Link of AI buyers is is very, very high.

They leaned on Radar for fraud prevention. They leaned on Sigma for analytics. And so, really, like, Stripe took care of the financial plumbing and so that Lovable is just really focused with that small team on on product and growth, which, of course, they nailed.

But there's, like, small there's smaller ones. Right? Retail AI, they build, have you used them?

They're like CallAgents. Yep. Yeah.

For customer support. And so they launched last year. I think they have, like, you know, over 10,000,000 in ARR in their first year.

We mentioned LINK concentration. LINK actually powers 38% of their payments. So 38% of their payments run through our consumer network where the individual has an identity and their payment methods are saved on file, and it's literally a one click checkout for them.

They use us for smart retries. So, like, you know, when the transaction, usually, recurring bill fails, like, we retry it at the optimal time, which allows them to recover about 60% of their failed charges. They use us for Stripe Tax, which keeps them compliant in a 100 countries.

And so it's just a great example of how these AI companies, very lean teams, growing fast, going global, and really just, like, scaling up to look like a much bigger company was right behind them.

Speaker 3

I've heard you talk a couple times, and then we don't have too much more time. So just to hit on a couple last topics. I've heard you talk a couple times about the time you spend getting new clothes for your kids.

That's mostly something in all honesty my wife does in our home. Lucky you. Yes.

I I I would flatter myself that I, you know, do my share in other ways, but she's definitely better suited to pick out, you know, what will make the kids look cute. What are we what's interesting you know, folks who listen to this podcast will sort of know the basics. Right?

That, like, perplexity has a shopping thing and whatever. And, you know, we know what MCPs are. Are there, like, any recent developments or, you know, is this really happening?

Or is it still, from what you've seen, like, kind of the wouldn't it be cool if one day this were real sort of phase of agentic commerce?

Speaker 4

It's kind of both. Right? Like, it's definitely still early.

There's still a ton we're sorting out about how's this actually gonna work and how quickly is it gonna take off. But and we're seeing meaningful traction. You mentioned Perplexity rates.

You can, like, discover and book hotels, you know, directly inside the app. But it's it's not just, like, big guys like Perplexity. Like, Hipcamp is a little site that uses agents with virtual cards to book campsites off platform.

I'm from Montana. It's, like, impossible to get into Yellowstone National Park. I hate using their website, although I love the park and value that they're not spending a ton investing in in tech.

But Hipcamp is, like, actually solving that. And then, you know, I think it's easy when people think about a gen to commerce to think about commerce buying kids clothes consumer. But on the developer side, we're seeing the same trend.

Right? So developers now in Cursor can buy Vercel services right inside their editor. That's a brand new channel.

Right? Like, this really embedded commerce directly in the workflow, and and Stripe powers those transactions too. So, you know, we're not totally new to this.

Like, it was it was last November when we launched our agent toolkit, but we still have thousands of downloads each week. And I think just looking at the pace of adoption and looking at who's testing, I think agent commerce will be a major channel far sooner than than most people think.

Speaker 3

With that embedded stuff, like Vercell in Cursor Yep. It seems like that's much more about, like, the connective tissue of the user has an intent, and it's a question of how it's gonna get executed as opposed to any sort of, like, autonomous, you know, decision making by the agent or any sort of, like, meaningful delegated discretion to the agent.

Speaker 4

I'm gonna actually trust you to go figure out what to buy and and execute on it at this stage, or is is that still not really materialized? Yeah. So so I can't name names, but this idea of, like, a business in a box.

Like, I wanna build this business, and I actually don't know what third party tools and services I need. I just want the business in a box and go spin up the business. And that's not just the payment provider or the front end service or the, you know, bot protection or the, you know, HR system.

But, like, give me my whole business in a box, I think, could be an interesting direction. Now getting that right for the whole world of businesses that might be created is hard. Getting that right for, you know, a pretty focused AI wave that's coming online isn't isn't a crazy thing isn't a crazy thing to about.

So I agree with you that sort of the option sets in consumer is broader, and so there's more job to be done for the agent to select from this very broad option set. But, like, SaaS procurement is also very inefficient. Maybe we underestimate how inefficient it is, and that's not just in the selection of vendors.

That's also in, like, the pricing and negotiation with vendors. And so I don't think it'll be tomorrow, but I think there will be a there there.

Speaker 3

Cool. I'll keep watching out for that. Last question.

Just about kind of the the future of platforms, the future of scale, the future of kind of market power. I think back often to the weak Danthropic deck from, like, two years ago where the claim was made by Anthropic. We believe that the companies that train the best models in, like, 2025, 2026 may have such an advantage that nobody will be able to catch them from there.

Why? Because presumably, like, the models will help train its successor with all these data filtering and synthetic and constitutional AI and whatever. And, like, once you've got Claude four contributing to the training of Claude five, like, anybody who doesn't have Claude four and is still, like, you know, sourcing everything through Scaleai or whatever is just at a massive disadvantage.

It seems like that basically applies to Stripe as well. Like, is there any is there any hope for anybody to ever compete with Stripe given, you know, the 1.3% of global GDP flowing through the system and the massive data advantage that already exists?

Or are we now sort of in a future where we just need to rely on the Callison brothers to continue to be, you know, good actors? Like, it seems like this position is, like, almost unassailable.

Speaker 4

I think financial services is a big broad space, and there are a lot of services that one can provide in that space. In the context of data, the 1,300,000,000,000.0 a year is a lot.

Right? And volumes grow growing, like, 38% year over year. That's like a massive growing dataset.

The real advantage, I think, isn't the raw size, though. To your comment earlier, it's more the compounding loop. And in our context, that loop is the more data we process, the better better our models get.

The better our models get, the more value we deliver to businesses. Incentives are super aligned. The more value we deliver to businesses, the more the businesses grow, which means the more transactions they run through Stripe, and kind of that loop compounds year after year.

And, you know, obviously, that's why we talked earlier about why it's hard to make horizontal bets, but that's why we can make horizontal bets. Not just because we have scale, it's because we're in a position to harness that scale to create even better products, which then is is the feedback loop. And we're pushing this further.

Right? You may have heard at Sessions last year, we announced a big push for modularity. And so now products like Radar, our fraud prevention product, or our billing product, or that optimized checkout suite for your users are available multiprocessors.

So they don't just work on Stripe transactions. They work on transactions or on billing plans or on checkouts that are happening outside of Stripe too. And that actually gives us window into an even bigger data network and kind of further reinforces that loop.

So, you know, I think there's a lot to be done in the financial infrastructure space, and I think there will be plenty of players playing important roles there. But I think we are quite differentiated in the intelligence that we can serve to users, and it's just really fun to see how that intelligence in turn helps them grow more profitably.

Speaker 3

On the other end of that, do you ever think about trying to compete at the, like, foundation model level? This is something that obviously not many companies are really able to do, but given the depth of ML experience and the unique dataset that does exist and just the reputation of the company, you know, I sort of expect that if there was a special fundraising round to raise $10,000,000,000 to go, you know, train a, you know, Stripe one to try to compete with Cloud five and GPT, whatever, like, the money would be there. I guess, do you ever think about going that hard or how do you think about calibrating, you know, just how ambitious to be with the AI investments?

Speaker 4

Yeah. I mean, Stripe has always leaned into new technology ways. Like, back when we were founded, it was the platforms and marketplaces wave that got us a lot of the way here.

Today, it's the AI wave, and our mission is to build the economic infrastructure for AI. That shows up in today for big bets being the best partner for AI companies. So just helping them monetize effectively and scale globally and manage billing and manage tax and manage fraud.

And, you know, two thirds of the Forbes AI 50 already run on Stripe, and we're very focused on co building, whether it's usage based building or whatever the next wave is, co building with them, and and being the best partner. The second is enabling agentic commerce. So we only talked about it briefly, but, like, agents are going to be buying on your behalf, and we want that to work really well for the whole ecosystem.

Yes. For the consumer. Yes.

For the seller. And yes. For the platform or or commerce facilitator.

The third place we're really focused in in the world of economic infrastructure for AI is making Stripe native inside the AI enabled tools that developers already use, whether that's Vercel or Replit or Cursor or Mistral's LeChat. Right? Like, payments should show up right where the work is happening.

So that's thing three. And then fourth is what we talked about today, which is, you know, deploying our foundation model across the network to improve fraud detection, yes, to boost authorization rates, yes, but also expanding the intelligence layer that we provide to every user. So those are the four big investments.

I don't know. I'm not gonna say that there could never be a fifth, but today, we're really hyper focused on on the economic infrastructure for AI, not being an AI model shop directly.

Speaker 3

Gotcha. Cool. This has been excellent.

I really appreciate the time and the depth. Anything we didn't touch on that you would wanna leave people with or just any concluding thoughts? Nope.

Super fun. Thanks, thanks so much for having me. Emily Sands, head of data and AI at Stripe.

Thank you for being part of the Cognitive Revolution.

Speaker 4

Thanks so much.

Speaker 2

If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts where experts talk technology, business, economics, geopolitics, culture, and more, which is now a part of a 16 z. We're produced by AI podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.

ing. And finally, I encourage you to take a moment to check out our new and improved show notes, which were created automatically by Notion's AI meeting notes. AI meeting notes captures every detail and breaks down complex concepts so no idea gets lost.

And because AI meeting notes lives right in Notion, everything you capture, whether that's meetings, podcasts, interviews, or conversations, lives exactly where you plan, build, and get things done. No switching, no slowdown. Check out Notion's AI meeting notes if you want perfect notes that write themselves.

And head to the link in our show notes to try Notion's AI meeting notes free for thirty days.

Shared via Hopper