Brex’s AI Hail Mary — With CTO James Reggio

Latent Space: The AI Engineer Podcast
17 January 2026 1h 13m
0:00 --:--
Episode Description
From building internal AI labs to becoming CTO of Brex, James Reggio has helped lead one of the most disciplined AI transformations inside a real financial institution where compliance, auditability, and customer trust actually matter.We sat down with Reggio to unpack Brex’s three-pillar AI strategy (corporate, operational, and product AI) [https://www.brex.com/journal/brex-ai-native-operations], how SOP-driven agents beat overengineered RL in ops, why Brex lets employees “build their own AI sta

Summary

James Reggio, CTO of Brex, outlines the company's three-pillar AI strategy (corporate, operational, and product AI) aimed at enhancing workflows, reducing operational costs, and delivering AI-powered customer features. He details Brex's approach to building a multi-agent orchestration framework for its Brex Assistant and emphasizes the effectiveness of SOP-driven agents over complex RL models for financial operations. The episode also covers fostering an AI-native culture, managing code quality with AI, and the evolving role of engineers in the age of generative AI.

Chapters

Brex's Three-Pillar AI StrategyJames Reggio introduces Brex's AI strategy, encompassing corporate AI for internal tooling, operational AI for cost reduction, and product AI for customer-facing features.
James Reggio's Career JourneyJames Reggio discusses his transition from mobile engineering leader to CTO, highlighting how founder experience and general business leadership were key to his career progression.
Brex Engineering & AI Team StructureBrex's engineering organization is structured around product domains, with a dedicated 10-person AI team focused on LLM applications and fostering widespread AI adoption across the company.
Brex's AI Platform ArchitectureThe discussion covers Brex's internal LLM gateway for prompt management and observability, its adoption of Maestro for agent development, and the use of various vector databases like pgVector and Pinecone.
Multi-Agent Network for Brex AssistantBrex developed a multi-agent network where sub-agents communicate in multi-turn conversations to power the Brex Assistant, aiming to automate employee tasks like expense management and travel.
Operational AI: Impact & LessonsOperational AI has delivered immediate business impact by automating tasks like fraud detection and underwriting, with the key lesson being that simple SOP-driven agents often outperform complex RL models.
Managing AI Code Quality & EvolutionThe conversation addresses the second-order effects of AI-generated code, such as maintaining code quality, ensuring long-term maintainability, and the evolving craft of engineering.
AI Fluency & Knowledge ManagementBrex emphasizes building a unified knowledge base to ground LLM applications and implements an 'AI Fluency Levels' framework to upskill all employees, including operations and leadership, in AI usage.
CTO Concerns: Headcount & AI ImpactJames Reggio discusses common concerns among CTOs regarding the impact of AI on headcount, the junior-senior level mix, and the nuanced effect of AI on overall engineering capacity and productivity.
Call for Multi-Agent Network InnovationReggio concludes with a call to action for the industry to focus more on multi-agent networks and agent-to-agent interactions, believing it unlocks more sophisticated and fluid AI capabilities.

Topics

AI strategyCorporate AIOperational AIProduct AILLM applicationsAgentic codingMulti-agent systemsFinancial AIFraud detectionUnderwriting automationKYC processesExpense managementTravel bookingRegulatory complianceCode qualityEngineering cultureAI adoptionPrompt engineeringEvaluation frameworksKnowledge managementAI fluencyHeadcount planningCTO challenges

People

James Reggio (guest) Alessio (host) Swiggs (host) Pedro (mentioned) Camilla (mentioned) David (mentioned) Brad Taylor (mentioned)
Key Concepts (15)
Three-Pillar AI Strategy — Brex's framework for AI adoption, covering corporate AI (internal tooling), operational AI (cost reduction), and product AI (customer-facing features).
Founder Mentality — The idea that ex-founders bring initiative and agency to companies, but also the challenge of retaining them as long-term employees.
Quitters Welcome — Brex's employee value proposition that celebrates employees who leave to become founders or department heads, fostering a culture of growth and support.
AI Center of Excellence — A centralized team (10 people at Brex) focused on LLM applications, designed to explore new AI possibilities somewhat independently from core product teams.
Agentic Finance — The concept of building production-grade AI agents to automate financial workflows and tasks, as described by Brex's CEO.
LLM Gateway — Internal infrastructure built by Brex to manage, deploy, version, and evaluate prompts, handle data egress, model routing, observability, and cost monitoring for LLM applications.
Multi-Agent Orchestration Framework — Brex's internally developed framework for agents to 'DM' and have multi-turn conversations with other agents to coordinate and complete complex tasks, especially for the Brex Assistant.
Brex Assistant — An executive assistant-like AI for every employee, designed to completely disappear the Brex UI/UX by handling tasks like travel booking and expense documentation via agents.
Build Your Own AI Stack — Brex's approach of allowing employees to choose their preferred foundational models (ChatGPT, Claude, Gemini) and agentic coding tools (Cursor, Windsurf, Cloud Code) to foster experimentation and competition.
SOP-driven Agents — The realization that simple, auditable, Standard Operating Procedure-driven agents are more effective for operational AI tasks (like underwriting) than complex reinforcement learning models, especially in a regulated financial institution.
Knowledge Base for LLMs — The critical need to build and curate a corpus of internal product and process documentation to ground LLM applications and prevent hallucinations about Brex's current offerings and policies.
AI Fluency Levels — A framework developed by Brex's operations team to create learning pathways and upskill employees in AI, transforming roles from SOP execution to prompt refinement and eval building.
AI-Native Interview Loop — Brex's revamped interview process for engineers that requires using agentic coding to complete a project, also used internally to upskill existing engineers.
Audit Agent — A finance team agent launched by Brex that ingests SOPs and looks for patterns of waste, fraud, or abuse across multiple expenses, raising potential violations.
Review Agent — A companion agent to the Audit Agent that applies 'wisdom' to potential violations, deciding if they are significant enough to be turned into a case based on factors like dollar amount and user compliance history.
References (29)
Stripe company
Banter company
Convoy company
Cursor tool
Clocko tool
ChatGPT tool
Maestro framework
TypeScript
Kotlin
Elixir
pgVector tool
Pinecone tool
LangChain framework
LangGraph framework
ConductorOne tool
Claude tool
Gemini tool
Windsurf tool
Cloud Code tool
Codex model
Creptile tool
Danger Systems project
Retool tool
ChatGPT Five model
Sierra by Brad Taylor company
Veracy company
Microsoft company
Salesforce company
DoorDash company
Transcript (54 segments)
Speaker 1

We have, like, three pillars for AI strategy. We have our corporate AI strategy, which is how are we going to adopt and, like, buy AI tooling across the business and basically every single function to be able to 10 x our workflows. Then we have our operational AI strategy, which is how are we going to buy and build, solutions that enable us to lower our cost of operations as a financial institution.

And then the final pillar is the product AI pillar, which is like, how are we going to introduce new features that enable Brex to be a part of the corporate AI pillar of our customers? It's like we wanna build features and be a solution that somebody else is saying to their board, hey.

Speaker 2

Hey, everyone. Welcome to the Laid in Space podcast. This is Alessio.

I'm their Kernel Labs, and I'm joined by Swiggs, editor of Laid in Space. Hey. Hey.

Hey. And we're here with Jasper Giosi to have Brex. Welcome.

Hey. Thank you for having me.

Speaker 3

from up in Seattle where I I've been a little bit. It's cold up there. Yeah.

And we have an atmospheric river hitting the the the city right now, so a lot of it. Yeah. Well, yeah, it's we're getting we're getting the full on winter effect right now.

Well, you're you're here to we talk about the sort of AI transformation within Braxton's a lot of interesting tidbits that we were gonna draw from your article, but also your background. You've got a wide array of experience from Stripe to Banter to Convoy. Mhmm.

And I I think also mostly, I'm interested in your journey as as one of the rare people that have transitioned from, like, a mobile engineering leader to a CTO, which I think is also a bit more rare. I I used to have this comment in the past where there's a career ceiling for people who work on client only things Mhmm. Where usually they don't hit CTO.

Whereas they typically promote the the back end people or the back end clouding for people to CTO.

Speaker 1

Yeah. You know, it's it's something that I I hear fairly fairly frequently because there aren't that many folks with a front end background to reach this level of leadership, and it's exciting for me to be able to represent that group. But I I'll say that even though my resume kinda reflects that I've been more on the the front end of things, it's probably more my experience as a founder a couple times over that actually helped me get to this this level of my career working for somebody else.

Becoming CTO is very much like a leadership and and, like, general business role as much as it is a technical role. And so I think it was more the skills that I built from starting companies and and trying to build those up, made me a decent fit and enabled me to get the nod from from Pedro to take this on as my predecessor left about two years ago. Yeah.

One thing I'm curious of you guys' commentary. This is a little bit broad Mhmm. Unscheduled, but a lot of startups are bragging about how many ex founders they have.

Speaker 3

which is what what you did, to be your employees and to to take initiative in the company.

Speaker 2

becoming anti signal sometimes. I don't know if you've thought about this. I think it's more about the turn for me, especially when people are hiring ex founders.

It's like, if you're truly of the founder gene, it's kinda hard to just stay somewhere. It's like an IC for too long. And then it's like, alright.

I joined this thing, and then in one year, I'm back to being a founder. I'm curious for you. Yeah.

What was your I'm sure you thought about leaving and, like, doing another company and say In fact, that was that was the the alternative. Was considering even at the time that I got the phone call where they made me the offer to become CTO. I was thinking about leaving to go start a company.

Speaker 1

And, you know, I think what's interesting about it, we we actually launched, sort of like a new recruiting and employee value proposition for Brex a couple months ago called Quitters Welcome, where we actually intentionally are leaning into this idea that we have a disproportionate number of folks who go on to become founders or, like, heads of a department when they leave our company, and and we celebrate that. It's actually something that I'm very proud of. And that means that, like, we we welcome in people who want to get a different experience.

I think that there's certainly, like, a lot of founders who don't make it don't scale their own businesses to this to the scale that we've achieved at Brex, so there's something to be learned when they come in. And then we're very happy to, like, support people on their way out. And so I I actually really like hiring former founders or future founders.

The one value proposition I find that's most relevant, because a lot of the folks we're hiring as AI engineers, are kind of folks that are either, like, winding down their companies or or considering maybe running AI startup. The the thing that resonates the most with them is that we oftentimes can give them problems to solve that are interesting, problems that may maybe they even want to want to, like, build their own startup around, but with instant distribution. Right?

Like, that that is the that is the allure. Just like, can come into this business and build, like, financial AI applications and instantly have that deployed to roughly 40,000 customers across, you know, the Fortune 100 down to, know, you tens of thousands of startups. So that that's what is, I think, appealing to founders.

But but the the challenge then is making sure that we set them up for success in in an environment that still feels a little bit like the startup that they might build themselves versus, like, something that's too corporate. Yeah. Instead of doing your own company and then coming to you and be like, can I integrate into Brex?

You gotta get all the data. Yes. Exactly.

Speaker 2

the engineering team structure?

Speaker 1

Yeah. So we have about 300 people in engineering, like 350 total across EPD. And for the most part, we structure around our product domains.

And so this means that Brex is a corporate card. It's also a corporate bank account. It's expense management, travel, and accounting.

And so we we actually have sort of full stack product domains that are roughly, like, thirty, forty people for each of those that have everything from, like, the low level infrastructure up to the the web and mobile experiences. Mhmm. That's generally, like, the structure of of our of our engineering organization.

And then we have, naturally, like, a an organization that focuses on infrastructure, security, IT. And then there are two additional centers of excellence that we've kinda built that kind of violate that org design where we've felt the need to to put more focus or, like, operate slightly differently. And AI is one of those areas where we have another team of just roughly about 10 people who are focused primarily on LLM applications.

And we wanted to create a bit of a separation there because the way that we were thinking about this, and this is actually something we did this summer, is we we paused and asked ourselves on on our AI journey towards, like, infusing our product with AI and generating customer value. We asked ourselves, like, what would a company that was founded today to disrupt Brex look like? And then we tried to basically use the answer to that question to form this team internally.

So it's a little bit off to the side.

Speaker 2

you know, LLM features, but but we have this sort of off on the side right now in a centralized manner. What's the difference in AI adoption for those teams? So, like, are the people on the LLM team, like, much bigger Cursor users, Clocko users, or, like, do you see similar diffusion?

Speaker 1

It's actually fairly fairly uniform across the entire engineering department. It's actually kinda funny. Like, one of our our largest cursor users is actually an engineering manager.

So, like and I I think that this also just, like, speaks to our core value of operate at all levels where we want all of our EMs and everybody in leadership to still basically do the job that they're managing, manage the work. So it actually is I I I think the journey of getting everybody into using agentic coding was not sort of exclusive to, like, the AI group.

Speaker 3

Yeah. I in fact, I think this podcast was actually set up because I called Outreach to Pedro Mhmm. Because he's he tweeted this.

I I I assume this is the synerexist. Yep. He says, I started a new company inside Brex to build the future of Vagintic Finance.

No BS, just builders building nine eight six and pushing production grade agents to 30,000 finance teams, now 40,000. And then he actually has, like, a little job description, which I think is really interesting. I'll skip that and go straight to Brexit's automated growth, five x, and cut burn 99% in the past eighteen months.

I assume that's a mix of internal AI automation and other stuff. Mhmm.

Speaker 1

put some headline numbers up front to impress people Yeah. Before we dig into the details. Yeah.

Absolutely. And you're you're correct. That's the that's the team that we have, this, like, AI team.

You're actually, what was that? Very young team. Yeah.

It's very young. I mean, it's and it's been really interesting. The the composition of the team is, like, very young, like, AI native, like, 20 year olds who basically grew up with the tech, kind of paired off with more, like, staff level software engineers that have been at Brexford a while who can kinda navigate, like, the existing code bases and, like, understand the product and the customer deeply.

Like, we've formed this really couple of tight tight knit pods in the AI org where there's, like, three people. Somebody who has, like, more of a product, a customer focused background that, like, staff engineer who knows where the skeletons are and then, like, a a much younger, like, AI native engineer who can just do things with with agents that, like, the rest of us dinosaurs maybe don't don't can't either dream of or, like, or where our I I think I think part of it is, like, sometimes the too much experience or too much knowledge of how to solve a problem and actually being impediment to thinking differently about it and thinking about it from, like, an AI first lens. But, yes, we we've been we've been slowly growing that team just in the same way that, like, a pre seed startup, you wanna be very, very careful about talent density and, like, very deliberate.

Like, only hire when you absolutely need it. And so, yeah, at this point, it's just about 10 people, and I think it was probably four or five people. I think everybody was actually in the photo that was attached to that tweet when Pedro put that out a couple months ago.

Yeah. We'll put it up. You that's a photo at 01:20AM in a on a Friday.

Yes. Oh, yeah. Yeah.

Because we we always do we always do, like, Friday Friday demos and and, like, that's a time for everybody to get, like, kind of exec review time. And so Everyone's in Seattle? Those folks were all in Seattle, but they're actually geo geographically distributed.

We have a couple folks here, a couple in Sao Paulo, a couple in Seattle.

Speaker 2

How at Decibel, we have this, AI center of excellence, which are basically the people running these teams across companies. Yep. How do you make VR engineers not feel like you're not special?

I think that's something that I hear a lot. It's like, hey. You know, why aren't these people working on all the GoLM things?

And, like, I'm stuck working on, you know, the KYC integration with whatever. Yeah. You know what I mean?

It's like, how do you build that culture? You know what's interesting?

Speaker 1

really optimized our engineering culture around business impact actually causes it to cut in the other direction where where folks some folks don't want to work on the AI products because it doesn't have as much clear direct, like, business impact right now. Doesn't doesn't impact revenues directly. And so I I think folks, for the most part, we've we've enabled folks who have as strong as our work on on AI products to to join that team.

Like, somebody somebody transferred out of our expense management organization to come over there because they're really passionate about taking, like, their knowledge of, like, policy evaluation and and bringing it into the the AI team. But the most part, I think everybody understands, like, how their work ladders up. And maybe there's some, like, friendly rivalry because, like, the folks who say, work on your card product, they they drive 60% of our direct revenue, and so they now they're pretty happy with that, and and they don't feel like they're being left out.

And I will also say, as you probably saw in this this piece that we we put out with first round, there is a lot of smaller applications of LLMs peppered throughout all of our product and operations teams. It's just some of the more novel, like, agentic layer that sits on top of Brex that has been put together, like, in this in this sort of isolated team. So it's not like folks aren't getting to to build with LLMs or use LLMs on a daily basis.

Yeah. Maybe run people through the Brex agent platform. We'll put the diagram in the video where you had the LLM gateway.

You had, like, the whole MCP layer. We just had David, the creator of MCP right before you. So this is very timely.

Yeah. Yeah. How did you start building that?

What's the architecture? Yeah. The architecture, you know, I I think simple is is elegant, and we we've had basically an LLM gateway and and a basic can rolled platform from the very early days.

In fact, right before being tapped to become CTO, I was leading, like, a AI labs team internally in the wake of, like, the announcement of ChatGPT. You know, everybody saw this through technology and said, hey. What are we gonna do with it?

And so one of the first things that we did, I think, January 2023, that would have been, was try to put together some internal infrastructure that made it possible for us to deploy deploy manage version and eval prompts, and then be able to manage, like, data egress and model routing and have some very basic, like, observability and cost monitoring in an LLM gateway. So that's that's infrastructure that we stood up, and it still continues to power a lot of those smaller, more, let's say, like, precise applications of LLM. So, like, for instance, we've, we set up a completely automated pipeline for, evaluating, customer applications to get them onboarded instantly to Brex, which is something that used to require human intervention either for underwriting or KYC.

But now we basically have a series of of agents and and particularly, like, research agents that will go and do the work that humans would normally do. And so that's running on top of this this hand rolled framework. And then for the agents on Brex that we announced in our fall release, which is like this agentic layer that we're building that sort of sits on top of Brex and can embody workflows that a finance team would normally hire humans for.

We've actually started using Maestro for that as, like, the kind of primary primary framework for for accelerating us. We actually built everything in TypeScript, which is another, like, technology choice that's answers the question of, like, what would we do if we started Brex today, but isn't the case for all of our existing back end code, which is either Kotlin or Elixir. And then we have we have a mix of pgVector, Pinecone, and, like, I think what we've seen is we're always we're always reevaluating the tech and framework choices as we go, because the half life of code has declined so significantly with agentic coding.

It's actually quite, easy for us and for anyone else to to kind of try on for size a variety of different pieces of tech to to figure out what is going to be most ergonomic for solving the problem. Double click on Maestro, that's a new choice, an interesting one. Yeah.

I mean, I think that the main the main reason that we adopted Maestro is that it provided the ergonomics that we were actually that the ergonomics of Maestro are quite similar to the internal LLM framework that we built two and a half years ago. Whereas, like, LangChain was available at the time two and a half, three years ago. It didn't quite feel right to us when we were trying to it it kind of addressed the things that weren't the the pieces that we we needed to address, which was, like, being able to have really simple observability and and logging, tracing.

LangChain didn't do it? It I mean, at that time, it didn't. I think it was really I think it was Oh, they fixed that.

Yeah. No. They certainly did.

They certainly did. But but but so, like, we we did I'm trying to remember because this is now ancient history. We evaluated Linkchain, turned off of it, built our own thing.

And then as we were looking, we we kind of want to deprecate this internal framework that we built because at the end of the day, it's not leveraged for us to maintain that. And Master ended up fitting the bill for for the the feature set that we were looking for. And I think what what's been interesting is about half of the the applications that we're we're building right now on the the agent layer are running on Maestro, and then the other half are actually still running on, like, yet another internally developed framework, which is a framework that's focused more on networks of agents.

So sort of multi agent orchestration versus more, like, strict, like, you know, single turn or, like, workflows, which are easier to use, like, either Lend Graph or Maestro. Tell us about your multi agent framework. I mean, that's what are the design considerations?

Why why why is this the first we're hearing about it? Yeah. Yeah.

So it's funny. I I a big big reason why we haven't written more about this is that it continues to evolve quite a bit. And I I feel like we we actually had a blog post that we were going to put out in conjunction with the fall release talking about how we built this.

And by the time that we finished, you know, the blog post and had all the package ready, it was already, like, halfway outdated. And so the way that this has started to emerge is this multi agent network approach to implementation was when we were trying to scale up our sort of consumer grade Brex Assistant. If you think about, like, Brex and our customers, there's really, like, two very broad personas that we serve.

We serve members of a finance team who are generally, like, going to be do it, like, in roles like accountant or controller or head of T and E. For those folks, they are going to be interacting with agents that are much more specific to their roles. But then the other broad cohort of of users we have are, like, employees of companies that have deployed Brex.

So, you know, you go join a new company. That company uses Brex. You get your Brex card.

And our goal for employees is for Brex to completely disappear. Like, the best UI UX for Brex is just the card. Like, every single thing that you have to do in the software beyond just swiping the card is like an opportunity for AI to to eliminate some work for you.

And so what we thought was the right approach to solving that for that was to was to embody, like, an executive assistant for every employee. Because I, as an executive at Brex, I have an EA, and she knows enough about me. She has access to my calendar, my email, has all the context on when I'm traveling and for what business purposes.

And so she's basically able to do everything that I would be obligated to to do in Brex, be it, like, booking travel or, like, doing expense documentation. And so what we wanted to do is we wanted to build, like, that EA connected to the same data sources and see if we couldn't simulate that behavior so that, you know, you basically you're interfaced to Brex's SMS in the card. And when we started building that out, you know, the most naive, like, architecture for that would be to have an agent with a variety of tools and maybe maybe do some some rag to ensure that it has, like, appropriate context for the conversation.

But what we were finding is that the wide range of different product lines that exist on Brex made it difficult for one, like, agent to perform well, being responsible from everything from, like, expense management to finding and booking travel to answering policy and procurement questions. And so that's when we started breaking down the problem, and into into a variety of sub agents that sit behind an orchestrator. And, obviously, this is something that can be implemented using LangGraph, or Master even has the notion of these as, like, network switches and beta.

But what we found is that it was easier for us, when it came to being able to build evals for the system. We we kind of just hit the eject button and built our own framework, which is one in which, we have agents that are able to, to basically DM with other agents and have multi turn conversations amongst themselves to coordinate to to complete a task to or, like, to complete an objective. And what's what's been nice about that is it means that, like, you can have your Brex assistant.

There's, one single one single, like, point of contact between you as an employee and the Brex product. And then behind your assistant, if the company has, like, expense management turned on, you have that. If they have reimbursements, there's another agent for that.

If they're they have travel attached to the Redo agent for that. It actually also then facilitates like, our conception here is that, you know, it's like generally, like, software encapsulation patterns taking, like, sort of projected into the agent space. It also makes it easier for us to have, like, the team that owns and understands travel, like, be the ones to go and iterate on that without needing to worry about, like, redressing the total system, or needing, like, one team to own every single possible act action you could take as an employee.

And I'll say that, like, I'm still of the mindset that somebody will build a a great framework, and we've they ultimately migrate to it. But or it might be us that we I'm ultimately over the source of this. Right?

Like but, but for us, like, this is, this has worked out quite well in, like, lieu of, like, a couple other approaches that we we tried along the way that just didn't perform well, which was to, you know, overload the the agent with a variety of tools or contextual, like, context switching where we try to say, oh, this conversation looks like it's more about reimbursement. So let's, like, update the the prompt with more reimbursement context. Like, that was that was another approach that we took that didn't perform as well as actually having a reimbursement agent that it would collaborate with.

Speaker 2

What about MCPs

Speaker 1

as, like, sub agents? Oh, yeah. It's another pattern.

The key thing there is that we there's actually a lot of value in having, like, multi turn conversations from, like, the orchestrator or the assistant to, like, the sub agent. Whereas, like, you know, a tool call is basically just like one RPC. And so oftentimes what will happen is, you know, let's say let's say the the the user reaches out to their Rex assistant and says, hey.

Like, am I allowed like, how much am I allowed to expense per person for dinner tonight? I'm taking my team out. And the the you know, your assistant's gonna then reach out to the policy agent.

Maybe the policy agent needs to know in order to answer that question, maybe it needs to know, like, whether this was, was, like, a customer event, a team event, or whether you're traveling. And so it may actually send instead of and it can't just answer the question, so it's gonna reply back to the assistant and say, hey. You I need you to ask this clarifying question.

And so then the assistant will return to the user as clarifying question, and then they'll basically have this sort of multi multi turn conversation across multiple agents versus it just being encapsulated in, a single call and response tool call. And so there are still, like all the all the sub agents have a ton of tools. But I I think of, like, the MCP and and tool usage as being, like, the interface to all of our conventional imperative systems, not at the the AI space.

Yeah. That's the conversation we were having earlier, whether or not it should be an agent to agent Mhmm. On Gol as well.

Yep. Or, like, yeah, there should be, like, a chat back. Exactly.

Exactly. And that's the thing. It's like, okay.

And one of the ways that we actually grafted this into Astra before we we built our own framework was to was to make every sub agent a tool, and then the input was just natural language. The output was natural language. And the if you needed to have multi multi turn, you would basically just put the full, like, our conversation in as you kept calling calling the sub agent as a tool.

And it's just like at that point, you're like, okay. The ergonomics are kinda the framework framework is fighting me on this. It's actually helpful for us to basically conceive of it as an org chart and, like, it's the agent org chart with with, you know, my EA is DMing other specialists and having brief conversations to support me as their client.

Yep.

Speaker 3

deep dive. Thanks for indulging. I feel like you guys are not afraid to make your own tech, which I think is a competitive advantage.

I really like that culture. Maybe I'll we should go a bit breadth first as well. Of course.

Because I think we also deep dive a little bit too much in in one area. There's and we'll we'll put up the chart, but I'm also very interested in, like, the the sort of internal agent stuff Mhmm.

Speaker 1

scope. So please feel free to just, like, go into your spiel on it. Yeah.

Of course. So one of the things that I was trying to do at the beginning of the year as CTO, you know, I think it really fell to me to articulate what our AI strategy was as a business. You know, every every board of director was, you know or every every member of our board was like, hey.

What's your AI strategy? And while we were doing a lot of good things, we'd really go, he's got it. Well, yeah.

And And and if I didn't, I I'd be in trouble. I think he also was counting on me given that I was doing the AI organization before CTO to to have That's true. But but a big part of it was, we we were doing a lot with with LLMs.

It was more like these little one off features and, you know, hey. Like, maybe mix in some suggestions here or maybe do a little bit of ops automation over here. But it wasn't it wasn't easy to to kind of create, like, a verbal framework of all of these investments.

And without that framework, then we weren't able to, like, set a set a a vision or a road map for for investments. So what we did at the beginning of the year is we took everything that was going on as well as all of our ambitions, all of the good ideas, as well as, like, the problems we were trying to tackle as a business this year, throw it all on the table and see if there were some ways to cluster it into a framework that made sense to the business, to our board, to ourselves. And we came up with I I think this is not particularly novel, but has helped us quite a bit.

We have, like, three pillars for AI strategy. We have our corporate AI strategy, which is how are we going to adopt and, like, buy AI tooling across the business and basically every single function to be able to 10 x our workflows. And we have our operational AI strategy, which is how are we going to buy and build solutions that enable us to lower our cost of operations as a financial institution?

Because I think it it's fairly intuitive. Like, financial institutions like ours face a lot of regulatory expectations, and there's just, like, a high ops burden for running our business. And so it's sort of like a lot of kind of internal use cases, like being able to do, like, fraud detection, underwriting, KYC, be able to handle dispute automation on card transactions.

Those those types of operational investments are our Ops AI pillar. And then the final pillar is the product AI pillar, which is like, are we going to introduce new features that enable PRACs to be a part of the corporate AI pillar of our customers? It's like we wanna build features and be a a solution that somebody else is saying to their board, hey.

We we adopted BreX, and this is part of our corporate AI strategy. Yeah. Yeah.

And so it's it's kind of has this nice little feedback loop, and we we basically, within the company, split you know, did a little bit of divide and conquer where folks in IT and on our people team were more or less spending more of the effort driving on corporate AI, really, like, looking for making procurement decisions, like creating a culture of experimentation where we spotlight and incentivize people for trying to sort of improve their personal workflows using AI. And then the the pieces that I've been more involved in have been operational and product. We were just talking about products here, which is like the agents on Brexit stuff.

But I think that the operational AI investments have been some of the the most sort of immediately impactful, to the business because we have hundreds of people who work in our operations organization. And it's actually something that differentiates us because our CSAT and the quality of our our support and service is very, very high. It's something we're very proud of.

And so trying to figure out how can we automate significant portion of this and use LLMs in a way that doesn't degrade the customer experience. And then also kind of addresses, like, what is the future of the roles of the people who we already have working full time for us? So this is where Camilla, our COO, who kinda co wrote the the piece with first round with me, she's been leaning really aggressively to help every member of the operations organization start rethinking their role as being not people who kind of execute against an SOP, but are people who are going to, like, build prompts, build evals, and, like, be become more AI native in, like, the way that they do work.

And so a lot of the engineering we've done has been to enable folks, say, in in fraud and risk to be able to to refine prompts and and add additional automation to their workflows.

Speaker 3

Yeah. And this secret fourth pillar, the the platform.

Speaker 1

Yeah. Yeah. Exactly.

Yeah. That is the that is the thing that ties it all together exactly is is the is the platform. And I think what's been really nice is that even though the platform is kind of a loose loose loose term because it consists of a wide variety of technologies, as I said, like, haven't been too religious or dogmatic about everybody needed to be on one particular thing.

What we've seen is that by making a variety of sort of ergonomic options for building with LLMs available, it, like, really has made it easier for for us to make a quick leap forward on operational AI. Like, we as soon as we put our mind to it, we said, like, look. No.

We wanna hit 80% automated acceptance rate for all all startup and commercial businesses that apply for practice. Like, we want a decision within sixty seconds that's fully touchless. No humans involved.

We were able to break that down and then actually build the build the agents, build the tools on top of that platform really quickly, and it and a lot of those tools are the same tools that our Product AI agents use as well. I was pretty sold on the Conductor. I don't know if this is under exactly that bucket, the Conductor one.

Mhmm. Oh, yeah. Provisioning command.

I was like, yep. I want that. Yeah.

That was actually I'd love to talk about that. So that's that's actually on the corporate side. And I think that this goes back to maybe another intuitive, but but I'd say, like, bold decision that we made, which is that we're not going to we're not gonna try to pick winners in the horse race between the foundational model providers or the the agentic coding tools or, like, basically anywhere where there's there's an active horse race.

What we do in instead of, like, trying to pick a single solution is we will procure, like, a a small number of seats, like, multiple solutions, and then we'll give employees the ability to pick whatever one they want to use. And so for instance, like, we allow employees to basically go to in in Slack and use Conductor one to get a ChatGPT, a Claude, or a Gemini license. And basically, can just, like, build your own stack where you pick your you pick your, like, chat chat provider.

As a as a dev, you can pick, you know, between, like, cursor, Windsurf, Cloud Code, Credits, like and and you can basically craft your your stack to your preference and easily switch between them. And what that does for us is when we're going to like, obviously, we have sort of enterprise agreements in place for all of them for the sake of, like, the, you know, the the privacy and non training guarantees. But it's fun because when we go to renew these contracts, it it we can basically resist the need to, like, do a wall to wall deployment.

We can say, hey. Look. Like, usage trends, they our our employees are voting with their fee.

They're voting with their dollars, and, you know, maybe maybe your tool isn't as as hot as it was a year ago. Does it give you a dashboard of what people are choosing? Yeah.

Actually, we look at that we were looking at that as we're going into budgeting over next year. Very interesting. I would love to see that those what what's, you know, anything that's like really up, anything that's really down?

It's fascinating how how different the landscape is every every three three months. And I think one of the one of the interesting challenges we had early on was getting folks to just, like, try these tools, try to incorporate, like, a genetic coding. You know?

And I like, early on, I say, twelve to eighteen months ago now, like, get folks to to just take the time to try a new workflow. And now at this point, I think what we're seeing is, like, even if, you know, a new model hits the same, like, when when Codex came out and everybody was like, oh, Codex is is better at at CodeGen, but it's a bit slower. Like, I find fewer folks are, like, kicking the tires on new things because, like, the they're just so comfortable with ergonomics of their current workflow that that, you know, some folks are just like, I wanna just stick with Cloud Code because I know it now.

I've been working with it for, like, nine months, so I don't need to to keep keep switching. I don't need I don't feel the incessant need to keep trying new things because I've I've gotten I'm an iPhone person, and I'm just, like, gonna stay with an iPhone even, you know, even though there's some really sexy Android hardware out there. Do you have one of the big numbers, like 80% of all of our code is written by AI?

Or but how how do you measure it internally? Yeah. No.

Not really. We we I mean, I what we do is we'll we'll measure, like, the attributions on the the number of commits that that have the and, like, co coauthored with. And we pull some of those stats, but I don't index heavily.

Like, in fact, I don't index on those at all. I don't and, honestly, like, I I don't know how I in honest like, honestly calculate that number. Yeah.

I agree. Yeah. And so so I the thing that the thing that we're we're really just you know, we're at the point now with, like, our AI agentic coding journey where now we're trying to solve the second order effects of, like, a little bit too much slop, maybe a little not enough yeah.

Exactly. Not enough, like, rigor and code reviews. We're trying to the adoption is there, and now we have to figure out, like, how to mature in our usage of these tools so that we you know, quality or, like, long term maintainability doesn't suffer.

As well as, like, maybe one of the other fact facets of being able to generate a lot more code more quickly is, like, the the drift between team members as far as, like, understanding of the the the code that's in their services increases is, like, everybody's moving faster and more independently. It that is another sort of risk that we're starting to see. Like, know, an incident response where folks don't know they don't know a service as well as they they used to because it's changed so much in the past couple months because everybody's moving more quickly.

Speaker 3

Yeah. This has been a major topic for me this year on code based understanding and SLAB, because obviously, it's so much easier to generate code, but then now we have to review it. Mhmm.

And to some extent, you can't really fight AI with more AI. You can't just be like, oh, just throw throw an AI reviewer onto the AI code and you solved it. And and so so you do need to just scale human attention, and I I think that's something I've been pushing a little bit in terms of like, well, you're you're just gonna like, every engineer is just gonna own more code.

Yep. Period. And and be parachuted in, and be expected to ramp up, and be be productive, and also fix bugs, and if you're on, you know, pager duty or whatever to just because that I mean, everyone's gonna try to be more efficient, and you're supposed to see ROI productivity.

Because if you don't, then what's the whole point of Exactly.

Speaker 1

Exactly. And I and I think it's funny you're going back to the point of, you know, you could you could add AI on top to solve the problems that AI introduces, and there'll be you just keep you that's like an endless chain. And so Well, no.

I mean, the the the the cold rabbits of the world, the the graphites of the world would say, yes, actually, you can. And so that's the little bit of the the tension there. Yeah.

You know, I I I've been thinking a lot about how the craft of of engineering is evolving, and and I will say that I feel further away from being able to predict what what it looks like than I I did this past summer when I spent a bunch of time. I actually basically went on leave for a month and joined the, joined the the the team that, the the AI team that we were building just to go and build alongside them. And I felt like it was really important for me to deeply understand the problems in the tech.

But and so that was me. I was I was, you know, writing, pushing code effectively nine nine six. And and I I went through so many different moments of realization of, like, oh my god.

This is going to change everything to, oh my god. This is just amplifying all the good and the bad in the industry to, oh my god. Engineers are not gonna have a job anymore to you know, it's like and so I I don't have any predict like, I felt like I had all the predictions back then.

And at this point now, I'm just very interested to watch the the phenomenon continue to unfold in front of us. And I will say, I was chatting with a bunch of really bright, you know, college juniors and seniors at a dinner we hosted last night. And while these folks are about to enter the industry, basically having kinda come up in the the era of agentic development and LLMs, and I asked them, like, what is your workflow when you're, like, building, like, building a project?

How do you how do you use agents versus, like, when you decide you're gonna actually just write code by hand? And I was surprised to hear the consensus was that most people there were using agents to collaborate on, like, building a design document and look at, like, collaborating on the architecture of the solution that they want to build and then maybe asking it to, like, emit, you know, a doc or an implementation plan, but then they'll go and write a lot of the code themselves still. So it's a little bit more of the the, the rubber duck co architect, use case that was most prevalent in that group.

I I was very surprised by that. I'm impressed. The kid the kids are alright.

Yeah. I know. No.

They still wanna they still wanna actually write the code themselves. It's interesting.

Speaker 3

Yeah. What we hear from, like, the Gen Zs at OpenAI, they they just YOLO everything into Codest. Yeah.

Speaker 2

I would say most of the code I generate is like yeah. But but I spend a lot of time on the doc. It's curious, like, when you're, like, younger in your career, it's like, you you don't really have all the mental models of the different patterns to instruct.

I feel like there's, like, overreliance, especially if you're doing the design doc. You know? I I feel like most of the senior engineers will spend more time on that.

It's like, even things like, you know, what columns should you index Mhmm. Depending on, you know, what queries we usually run on this table and things like that. It's hard for any AI to know that.

Right. You know? And it's like, I feel like the the role of, like, the more senior engineer should actually be more of this.

It's like spending time teaching the AI, and then the AI can teach the junior people in a way. Yeah. Yeah.

And it it everything everything looks like mentorship and management. At the end of the day. Right?

It's like you're breaking down tasks. You're you're supervising work. You're giving feedback.

Like, it's, you know, it's basically management. Except that there's agents are really bad at memory still. Like, basically have zero memory.

Right? And then it's it's Seattle 2025.

Speaker 3

What's going on? Yeah.

Speaker 2

Yeah. What's your internal stack for like preferences? There's like kinda like, you know, explicit preference you can use with, you know, agents that MD and all that stuff.

There's implicit preference with Linter rules and things like that in a way where it's like, it just happens. You don't have to tell it. How do you structure that?

Oh, no. You're talking about for agentic coding or with memory thin or like data platform? Yeah.

Yeah. For like the coding specifically.

Speaker 1

It's like and then we can kinda talk about, you know, the whole Brex platform. Yeah. Just just nothing nothing special.

Just a lot of explicit rules. That MD files. Yeah.

And then we have and we in linting, we still have like traditional linters in place for the couple of different language pool chains, and then we're we're we're big fans of creptile, and we use them for basically all of sort of the smarter than linting, like, agentic code review. That's been the one solution that we have aligned around that has served us extremely well. Yeah.

Good. Good. Reptile.

Yeah. No. We're we're huge fans.

They're they've built something really impressive. And I think the thing that constantly blows my mind about it is, the way that they're able to just have a really impressive signal to to noise ratio. Like, the the comments that it leaves are very, very high signal.

Like, never I never regret going through all, like, 65 comments it leaves on my on my diffs because it it catches so many things. Yeah. I found the Codex review to be really good.

I don't use Codex for code generation, but, like, the review product Yeah. Is, like, very good for some reason.

Speaker 2

I used to have when I was working in Rails, there was, like, this project called Danger Systems. Oh, yeah. It was kinda like a semantic linter.

Exactly. I feel like there should be more of that now. It's kinda like the rules are one thing in generation, but I want something in my CI that is like enforce these rules and call out where they're broken.

And then I can just copy paste that in an agent. But Yeah.

Speaker 1

code base, like, because as I was saying, like, we were answering the question, what would you do if you built a, you know, a Brex disruptor today? And it's like, it wouldn't be to pick Kotlin and Elixir as the back end. And and so we actually went with the full, like, TypeScript stack, and and we we were building on all, like, public interfaces and really trying to make sure that this agent layer was, like, arm's length from from the the good and the bad of of the core of our product.

And and one thing, I think what we did early on, and I don't actually know if this is true because, again, the team keeps Right. Sort of iterating, but we we're having good good luck using Cloud Code, like, in a GitHub action to basically go and do do more of that dangerous style, like, code review. So have a prompt for it that went through all of the different facets that were more conceptual versus, like, rigidly enforceable by a linter and have it leave a big comment at the end with, you your conformance to the idiomatic coding patterns of the of the new repo.

Speaker 3

I wanted to spend some time. You said you wanted to d five on operational agents. Mhmm.

The customer support, onboarding, KYC, fraud, delinquent account disputes. This is, I imagine, the bulk of it Yes. Of of the work.

Anywhere where there's a good story about maybe when you started audio, it gonna be this way, and then you discovered through building or through customer contacts that it had to go a different direction. And so that difference in beliefs is something that people can learn from.

Speaker 1

we believed at the beginning that using RL for credit decisions would actually be a like, would be the way that we would end up or, like, credit and underwriting, like, how much of a of a limit should we give to this business? That reinforcement learning would be the way that we would go about building a model that effectively would decision in the way that a human underwriter would. And it turns out that it was we made this big investment.

We were working with some outside, like, the like, a company that specializes in this, and the performance we ended up getting was inferior to just building a, like, a web research agent. Yeah. And so so I think what what we took away what what has been most evident in operational AI is that in operations, you need to be able to break down problems really granularly and be able to form SOPs that humans can repeatedly follow and and thus can be audited, because so much of the the responsibilities and operations is to, is to have audible repeatable processes that help to ensure that we're operating in a compliant manner.

And that actually translates just so cleanly to LLMs that we haven't needed to use too many sophisticated techniques in in operational AI. It's been it's been relatively simple, like, new tool, like agents, or maybe even a lot of problems can be solved, which is like a single turn chat completion. And so the fact that we didn't well, we did one one sort of attempt to overengineer and use more sophisticated techniques, And we we we discovered that, in fact, the solutions are a bit more more plain and and less technically sophisticated.

The the challenge is really articulating and refining prompts to reflect reflect execution of the SOP and, like, reflect all of the sort of institutional knowledge that isn't written down so that agents can properly replace, like, the the humans or the contractors that we would have making these decisions.

Speaker 2

How do you decide what is worth, like, spending a lot of time building versus what you think some of these models are just because some of these tasks are so generic. Mhmm. They're not really about Brex.

Yep. Like, you can assume the models will be good at it versus some of them are, like, very specific to you.

Speaker 1

like, the the tasks that are most common for the broadest number of customers. And the some of them are are are fairly fairly intuitive, like being able to research a customer to look to assess, like, legitimacy of the business and whether that business would fit our ideal customer profile for for onboarding because there's certain types of businesses that we either legally cannot serve or we are not comfortable being able to serve. So that's the type of really kind of basic research, and, like, a a relatively straightforward problem, that isn't hyper Brexit specific.

The things that are a little bit more specific to to us or or companies in our sector would be preparing documentation for a network card dispute. Like, if if you go and dispute a transaction on on your your personal card, you will provide evidence to your card issuer. The card issuer then has to put together, like, a three or four page word document that goes to, the card network and then eventually goes to the acquiring bank.

And and all of that is, like, much more specific to our business. It's a huge operational overhead for us, and that's something that we we decided to automate later because it's not as, it's not on the critical path of, like, serving the vast number of our customers. Like, disputes are expensive, but not very common operational process.

And so they're lower on the stack. And, I think we're we're getting there right now. But this year has basically been us just kinda, like, looking at every single process, just kinda stack ranking.

And I will say, like, the thing that got us started down this path was we wanted to expand our ideal customer profile to support more biz like, a wider variety of commercial businesses, which tend to be businesses that aren't growing as quickly. So they're not like tech startups, which have a lot of growth, and they're not usually, like they're not enterprises, which also tend to have a lot of growth. It's more like a lawyer's, a law firm, or a dentist office.

These types of, like, solid businesses that we should be able to serve and underwrite, but the cost to to onboard them and the cost to serve if you have all all the humans in the loop, make them ROI negative. And so that was the first sort of use case of of AI within our ops ops organization that then led to us really understanding we could automate much more than that. Is this Burks going back into question.

Yeah. Yeah. So never never let let let that die.

You know? No. We I think the way we've thought about this is we want to always, like, offer our product to customers where we believe we have a, like, a an offering that is well suited to the to the needs of those businesses.

And I I would say that still for very small, businesses, our offering isn't it's not built for that. It's built for it's built for companies that have some degree of scale, typically have at least sort of one person, if not a couple people in the their finance team. So we consider these to be more like the the commercial segment.

And so it rhymes with with SB, but our approach back then was was a little bit more naive. And I would say we also we were just going for a volumes like a volume game there. Our our internal controls were not as strong.

We didn't have as much experience, like, underwriting those businesses. And so it it was really it ended up being a huge burden for the business, almost existential, for us to have those tens of thousands of customers that all were, ROI negative. So we're we're trying to basically scale to serve more businesses outside tech and outside of, like, the out market segment, but but do it thoughtfully.

So I think right now, our, you know, our minimum threshold is is, like, $1,000,000 a year in recur in in annual revenue or or, like, $10,000 or more per month in in card transactions is kind of being, like, the low end of our ICP, which is obviously not what you would think when you think of a small business. Like, small businesses tend to still be smaller than that. Oh, wow.

That's really small. Okay. Yeah.

Yeah. Mid market. Yeah.

Exactly. And it's funny. It's just like the the the names of these segments, you know, it's like, you what what we could sort I don't know.

Yeah. No. I think I think, like, that's it's like yeah.

It's like lower mid market. And it's funny though, because when what we call enterprise may be another, you know, what sells what we call enterprises business that Salesforce might call a mid market. Right?

Like, because it's just that it depends on the scale of yourself as a business when you use these terms. And all of these things are built in the Brax agent platform, like all of these automations that people build? Yes.

Exactly. Yeah. And in fact, the most of the operational AI is running on that original platform that we have.

And we we built it one element of it that I didn't mention is that it it also most of the UI UX for this platform is built in Retul. And so, like, you you can basically go into Retool, and there's, like, a a prompt manager, a tool manager, an eval manager, and that's sort of where much of this was built. And the goal with that was, again, to make it more accessible, more ergonomic to to get started, but what an a secondary effect of having a more, like, visual set of tools for this is it's enabled members of the ops organization to go and do prompt refinement themselves.

So you don't need engineers to go and and refine the prompts or or even, like, test new foundational models when they come out. I think that that's another fun thing when, like, a new when a new model drops, folks will go into the the platform and basically run the evals on the new on the new model and kinda see, like, can we get better performance here, or does this have different different latency or different, like, cost characteristics?

Speaker 3

Yeah. You want the domain experts or the people directly using the tool, not the engineers who are Yep. Sort of somewhat removed from the tool.

Yep. Yeah. I I I I do wanna highlight to listeners that a lot of the Braxton agent platform are just things that every company should have, basically.

Promotional system, which we talked about, where the Topian experts are doing it, multi model testing, evaluation and benchmarking frameworks, API integrations for automated workflows, NCP based artist architecture with Brexit's external AI products. This one is obviously very Brexit specific. One thing I did wanna highlight that I was semi impressed by, because nobody people very few rarely talk about this, is knowledge base for understanding Brexit's business.

Yeah.

Speaker 1

So do you wanna expand on that? Yeah. And this is an area where we've only scratched the surface here.

But Yeah. But a big a big challenge that that we face is that the world knowledge or the knowledge that's built into the model about about, you know, what GPT five thinks Brex does and how it thinks our business operates is actually quite different from what our business offers today or how our product works. And so we've had to to work on building a corpus of sort of product documentation, process documentation, and, like, curate this set of information to basically ground a variety of our LLM applications, including, like, that Brex Assistant, which is like the, you know, the assistant that employees will will talk to.

It's like, don't want it to to hallucinate features that we don't have or, like, give give wrong information there. And similarly, like, some of the operational agents need to be grounded on, like, what our ICP is. Because if you ask, you know, ChatGPT five right now, like, what types of businesses does Brex onboard or, like, what types of businesses does Brex serve?

It might not give an accurate explanation to that to that question. It might it might say, hey. We're a corporate car for startups, which is what we did, you know, seven years ago.

And it might say, we're only we only serve enterprises. And so that has been an interesting challenge. And I think we're what we've been trying to do there is I'm actually going to be spending time with with folks talking about this next week internally about, like, can we refresh our strategy and kinda unify it?

Because we have a lot of product documentation that's internal for, like, our operations and go to market teams. We have a bunch of product documentation that's external for our customers. We have a lot of go to market sort of enablement material that's more sales pitchy.

And we have documentation that is put into Sierra, which is the, you know, the the chat assistant that we use for for frontline support. Like, all of this ideally could draw from the same source, but right now, it's right now, it's a little bit fragmented. It's just something that we're trying to invest in, though, because I think at the end of the day, the duplication of of efforts is just, like, is is is wasteful, and it's absolutely necessary to to get this right.

Speaker 3

Just to deduplicate Sierra, meaning the Brad Taylor startup. Yes. Exactly.

Yeah. Yep. So I would expect that you have you built so many other agents.

That's that's one you can build yourself.

Speaker 1

Okay. For us. I think what what's interesting about the the SIRA that has been really helpful is that, again, it's really easy for like, the UI and UX of basically administering a SIRA agent is something that's really accessible for the ops and CX strategy team, which are, like it's much more low code and more sort of workflow and DAG oriented, and the we have engineers kinda going and giving it tools to to take actions.

But the most part, like, it's nice to not have to build build the UX for somebody to manage something like that. And I think the fact that Sierra speaks the language of of customers yeah. Exactly.

Speaks the language of CX. They can do all the reporting and the telemetry and stuff that that our, you know, VP of CX would like to see. Just, you know, it's just one fewer thing that we have to build.

What about evals?

Speaker 2

How do you build evals? Who manages them?

Speaker 1

Well, it depends on it depends on the application. So on the on the operational AI side, those evals are are basically baked into the in the platform around every every prompt or every agent. And for the most part, I think most of these use cases kinda come online, like the v one of, like, our our, commercial underwriting agent or the v one of our our start up KYC agent are co developed between, like, a subject matter expert in ops and, like, an engineer, and they're gonna kinda co develop, an initial eval set.

But then from there, generally, in ops, you're always doing QA, be it, like, on humans or on, on on the LLM, decisions. And so whenever like, as part of our QA feedback loop, whenever there is a a mistake, that's usually almost always gonna result in, like, another eval being written as, a regression test. So all of that within Ops dot ai is pretty pretty straightforwardly managed.

On the product AI side, that's where it starts getting a little bit more challenging because the multi multi agent network, is quite challenging to evaluate. And so what we do there is we try to adopt some of the state of the art for multi turn evals where we will, we'll basically have an agent embody the user and, like, you know, have basically the the end user agent is given an objective, and then we basically have it run a multi turn conversation and then use l l a misjudge at the end to do all of the different asset assessment. The one other thing that we do technique wise that is interesting is sometimes you want to you don't want to do, like you know, I think these multiturn evals are kinda like integration tests.

They they sometimes test more than than what you want to to assess. And so sometimes what we'll do is we'll also pre can, like, an initial preamble to a conversation, and maybe a couple turns will be handwritten. And we'll basically set set the the, eval to start, and we'll see if, we're able to, like, isolate certain certain behaviors.

So, it's it's still, like, a work in progress. And I say, like, at the end of the day, a lot of the just periodic human review and and, like, looking at at cases where, we've detected as we go into, like, summarize. Like, what we'll do is we'll reflect on a conversation after a certain amount of time has has passed where we'll summarize it, like, extract facets like, did it seem like the user accomplished their their objective?

And we'll just manual when a lot of the cases when that's that's failed and decide to write another eval for it. Mhmm. Are all the evals supposed to pass, or do you have a set of evals that are, like, someday the model will be good enough?

And, like, how is it gonna change over time? Yeah. It's interesting.

I don't know if we have any that that are like, oh, someday, I hope it'll be good enough to do this. But it's like there there are the evals that are are blocking because they would indicate, like, a a regression, an unacceptable regression. So these tend to be just accuracy related evals, but then there are others that are more about, like, tone and coherency and these types of things where they're they're more subjective, and we were just looking at those over time as a as a metric.

But the the team is actually interesting. Think we're gonna get an a big update on, like, how the team is thinking about evals tomorrow and, like, our Friday our Friday review. So it's this is an area where I'd say the largest challenge like, the the largest change we needed to make in how we were executing sort of as, a lab or an incubator back earlier this year to, like, where we are now where we've we've shipped and, like, we're trying to to increase the rigor has been around, like, avoiding regressions and having more and more increasingly robust evals.

Speaker 2

Yeah. Work with a company called VeracyIs that does user simulations. And Mhmm.

I I I think, like, that's what's been interesting. Some of these things, they just don't expect. Like, the customer does not expect the model to do.

Mhmm. But they wanna track the saturation of the model in a way, if that makes sense. And I feel like most companies know what they don't want to happen.

Yep. But it's almost like they don't they cannot quite articulate, oh, I want in the future the model to be able to do this. They can do it today, but I'll I'll keep running this eval.

Speaker 1

really interesting to me. And I I'm gonna take that away and and and start thinking about this because there are there are going to be certain I mean, we already seen this where where, users will ask the assistant for help with things that we don't support yet or we haven't implemented yet. It's like those are opportunities actually for us to build a like, effectively write a test that's going to be fail like, failing Right.

For for weeks or months and and eventually will go green, but is a way for us to actually kinda show, like, the progression of sophistication of the assistant. I I really I really like that as an idea. Yeah.

I I wonder how you also catch hallucinations and of things that it doesn't have. That's usually the that's usually the problem, is it? Yeah.

It'll it'll it'll pretend like it can assist with something, and it'll like, one thing that is really annoying that has been tough to to prevent is that the the assistant, because it is used to speaking to other agents that can support it in, like, accomplishing various tasks, if you ask it to to help with a task that it thinks it probably should have an agent to to to work with, it'll just hallucinate that it they always like, oh, yes. I'll I will, like you know, I'll reach out to the finance team on your behalf to to pass this question along, but it's not doing anything. There's, like, no finance team.

There's no way for it to do that. This is something that comes up a lot. It's like, would you like me to ask the finance team?

And there's no there's no actual tool You put guardrails for that? Yeah. Yeah.

That was that was something that we had to Like, rejects. Like Oh, no.

Speaker 3

like, things that could get us into trouble. Yeah. Really extreme ones.

Yeah. I I just yeah. It's surprising when I I guess two years ago, was first kicking around the idea of all these things, I would have said that probably guardrails would be more prevalent, especially in finance use cases, But surprisingly, they're not.

Yeah.

Speaker 1

Like, coded. Yeah. Exactly.

Here's some red jacks. There's this tilt. Yeah.

Exactly.

Speaker 3

you just get, like, the inline 500 error. It doesn't even tell you that it can't help. It just, like, craps out.

Like, I we kinda built a couple of those circuit breaker or, like, the ability to put those circuit breakers in, and and I don't I'll believe we're using them for anything. One last thing I wanna get your thoughts on was AI fluency levels Mhmm. Which you guys have a framework of user advocate builder native, and everyone goes through it, including Camilla.

Mhmm. And I just think it's interesting. I think it's a model that other people are thinking about adopting, but they're worried about rolling it out.

That everybody's gonna be banned. And then I And, also, like, how do you have, like, this in house training course that you keep up to date? Just tell us more about it.

Yeah.

Speaker 1

more ahead of even engineering on this front as far as, like, trying to create create, like, learning pathways for this. And I think that part of the reason why they're ahead of us is that in operations, they're much more they have to be able to operate training at at scale. Like, training is a big very big part of of how people build aptitude around their their job function within ops.

Whereas, like, in EPD, a lot of it is sort of getting hands on, building experience, like, going a lot and getting mentored, getting code review. But it's been really neat because I think we've really like, we created an environment. We managed to by by speaking openly about the the transformation that we saw would happen in this industry towards AI sort of displacing a lot of, a lot of the operations and CX roles.

And we were just honest about it. And I think what what in the same breath that we said, hey. A lot of these job responsibilities will go away.

We also said, we don't anticipate that meaning that your job has to go away. It's just that your job has to change. And so the the fluency framework and then the the, like, the training and support and, like, the positive sort of culture where we celebrate people making progress has been really helpful for, like, avoiding a culture of fear or, like, oh, you have to do this or you have you're gonna get this is gonna go in your performance evaluation.

I think the It does. It well, it's not like as rote as like, oh, like, what is, you know, what is the like, how much are you using AI, and is it is it enough? It's it's more I think we've built a pretty, like, positively framed culture where we'll we'll do, like, spot bonuses for for people who have, like, particularly novel uses of AI on in their day to day.

In our company, all hands, every two weeks, we'll do an AI spotlight. And it's very rarely somebody in EPD for the most part. It's folks in ATMs, ops, finance, the people organization showing off, like, how they're building agents, you know, in ChatGPT or on Glean or how they're they, like, just found some new use case that they thought was helpful.

So we're trying to create create, like I think at the end of the day, like, we've hired a bunch of really smart people who like, I have full confidence that that that this type of work is in within the reach of anybody who's motivated to, like, sorta challenge themselves. And so we've we've done that. Then in engineering, there's one other thing that I wanna call out because I think that this is kind of fun is that we adapted our interview loop to be more AI sort of agentic coding native.

So instead of we had, like, a coding and a system design question that we basically have revamped into a project where we'll give you, like, a brief before you come on-site and then, an additional sorta spec when you do when you start. You know, we expect you to use agentic coding to complete the the task. In fact, it's, like, kind of impossible to get all the way through it if you don't.

And so we're evaluating, you know, your knowledge. Like, we're kinda watching how you work. We're evaluating whether you understand the code that's coming out.

We we you know, we're kinda probing at you as you go. But what we did in order to kinda bootstrap the process of all of our existing engineers, like, getting familiar with agenda coding is that we as soon as we had the interview ready to ship, we started we said everybody in engineering, including all the managers, are gonna have to go through this interview. And so we reinterviewed everybody internally.

And it's like it's one of those things where it's like, it it's not a we didn't, like, keep a score or, like, or, like, you know, I don't have any data on, like, who passed or failed or what they what they scored. But what we found is, like, as people would take it, it would actually cause them to have moments of realization where it was like, oh, I I can uplevel my skills around. So, like, I have like, I want to be better at this.

And so we're trying to find, like, a way like, a variety of techniques that kinda push the culture along. And I think as I reflect on, like, the year, because this is the year where we really put all the effort into it, I'm really satisfied to see the ascent to which everybody's leaning in on a on a daily basis. Going back to, like, even I was shocked when we were looking at our cursor logs that, like, the number one user is is an engineering manager on a for a InfraOrg.

It's like, that that is super cool to me.

Speaker 3

doing their job differently. I guess my I I had a closing question or I guess a parting question, and this is broadening out from Brex. Yeah.

And this is just you interface with other engineering leaders all the time. Mhmm. Did we not cover anything that other CTOs are having as top of mind today?

Like, their number one problem is underscore.

Speaker 1

The thing I find myself discussing with with folks that and I I don't wanna shy away from, like, scary topics. In fact, we're just just kind of on one that was adjacent, which is like, how do you evaluate somebody's, like, progression towards being more AI native? The the the the cousin to that question is it's like, will we need as many people Yeah.

To operate our businesses? Like, are there layoffs coming? Are how are we how are we thinking about, like, headcount growth?

Junior versus senior. Junior versus senior. Yes.

Exactly. Like, level mix. And I still have more questions than I have answers there.

I think I think what has been really interesting is that I view agentic development as being something that amplifies all the all the good just as much as it amplifies all the bad. And it amplifies sloppiness, poor architectural thinking, misunderstanding of of the requirements. Like, there are for all of the the acceleration of good outcomes, it also accelerates bad outcomes.

And think what has been interesting is that there has been when you sum that altogether, there's less of a obvious, like, capacity increase. It's it's more it's more nuanced than that. So I'm not looking at headcount planning as as we think about it next year as as being something like, oh, well, because AI is giving us so much more leverage, we we don't need as many people.

We've actually the thing I'm really proud of in in my tenure as CTO is that we we haven't grown engineering at all. What we've done is we've we've grown the business significantly, but we've been able to build, like, greater efficiencies in in how we execute, like, how we how we we think about building, how we road map, what we choose to do and what not to do, that we're able to to serve significantly more customers with more lines of business without needing to grow engineering headcount. I think that that's kind of the way that we're gonna just continue on this road is like, I like having 300 engineers.

Like, I would love love to just, know, a year from now have 300 engineers, but we're still, you know, thirty, fifty, 100% more efficient. That that that is the thing that comes up with with other engineering leaders. And the other part of that conversation is, like, how much is AI getting blamed for this sort of ordinary performance oriented risk?

You know? Like, if if Microsoft is letting go of, like, 4,000 people as a business, what, they have a 150,000 employees, I believe, Is that really, like, AI causing that, or is it them just using it as a way to to avoid some harder, like, perf management decisions? I'm not entirely sure, but I'm I'm listening more than I'm speaking on the on this topic because every time I feel like I have a pretty firm point of view, some new anecdote or experience comes in that kinda challenges or invalidates it.

Yeah. Well, you know, it's I I take these signals as it's my job to go find people who think they have answers and surface them. And you may or may not disagree, but at least you have something to use as a straw man in in your work.

Exactly. Exactly. And I and I think as as an industry, was just early innings on on on this transformation.

So I'm looking forward to seeing, you know, listening to this this podcast episode a year from now and and and seeing, you know, what we got right, what we got wrong, and what's different because so much changes quarter over quarter. Yeah.

Speaker 3

is a very well established pattern. I think internal platform is very well established pattern.

Speaker 1

fluency thing is something that people are figuring out that I think you guys are hit on. I'm happy to hear that. It'll be my feedback.

Yeah. Any final call to action for things that you wanna buy? Like, what should people build for you?

Like, problems you're trying to solve that you would love people to reach out for to to help with? The call that I'd make is for folks who are interested in in multi agent networks to to get in touch with us because I I do feel like this is something where where we're we're innovating in in service of of our customers and where I I feel like the the frameworks, the tooling, and the the research is is is there. There's actually quite a lot of, like, interesting papers and things that we lean on, but I would love to would love to see more of that, like, encoded in the in the what's available at large in the industry because I feel like my intuition has been that trying to craft LLMs into deterministic workflows and DAGs is is kind of underselling, like, the power that they have to actually find and execute more in a more sophisticated, like, fluid way.

And and I and I just want to see, like, the the industry lean in more on on these agent to agent interactions.

Speaker 3

Okay. So I'll I'll dive in a little bit here because I I have a minor opinion.

Speaker 1

You keep using the word networks. Yep. Is that a reference to a specific paper, or it's your term for it?

It's just our it it's our term. And I think that that is that's actually the term that master uses as well for it. It we yeah.

Initially, we used to call them agent runtimes internally, and that we just, yeah, switched to networks.

Speaker 3

And then I think the other thing I wanted to get a clarification on is, is it mostly a full agent talking with a full agent, or is there a kind of like a orchestrated boss agent talking to a sub agent? And I think that does matter for a subset of people who are building all these things. Because when you say multi agents, sometimes people don't agree what that means.

Speaker 1

Yeah. So it's it's a tree more than it is a graph. So it is like yeah.

We have When you say network, it feels more of a graph. Yeah. But it seems more directional as a tree.

Like, there there is a hierarchy. There's a hierarchy. Yeah.

But there but there are some violations of that. Like, one of the one of the interesting use cases, and this is where, like, the power of of having an an assistant for every employee plus having agents that run and and embody members of the finance team is really powerful because there's this interesting use case that that we brought to market, which is that, one of the finance team agents that we we, launched is an audit agent, where, like, an audit agent kind of embodies the work that a lot of larger finance teams will do to look for patterns of waste, fraud, or abuse or, like, systematic avoidance of policy that isn't as obvious with a single expense. Like, you can evaluate a single expense and the metadata around it to see if it if it's within policy or not.

But what if you start seeing an employee often make a large number of, like, $74 transactions when receipts are required at 75? Or what if you what if you see certain things like, oh, okay. There's actually a fair number of, like, DoorDash expenses during business hours from this individual, like, on on days that an office launch has provided, or maybe you see, like, rideshare patterns that are are where you have to look at a broader context.

So we built this audit agent that can, like, ingest your SOP and and look also ingest your This is a Sykes' customer's SOP. Exactly. Yep.

And and what it does then is it's it's basically always looking for potential violations. And what it does is it it is extremely zealous. Like, it it wants to have a minimum number of false negatives, so it will raise a large number of potential violations.

And then a separate agent, a review agent, will then apply wisdom the wisdom of, like, is this important enough to follow-up on? Is the dollar amount in question high enough? Does this user seem to have, like, a high compliance behavior more generally?

It makes a judgment call about whether it's worthy enough to take that violation and make it into a case. Then once it's made into a case, generally, what happens is that you need to get more information from the individual. So if humans were doing this, there'd there'd be some outsourced team that's, like, looking for all the potential violations.

Then you have some full time employee on the finance team who's who's looking at all the violations. Oh, these are the ones that are important. We need to follow-up on it.

Now what they do is they hand it off to somebody who will go and Slack that employee and be like, hey. What's going on here? And so what we have is, like, the audit agent looks for violations.

The review agent decides whether it's worthy enough to turn into a case. And then from there, when the case is filed, the that that will trigger an event to the brex assistant for that employee. And, like, any additional information about, like, the business justification can be collected.

Or maybe the assistant already knows because it in its conversation history with the employee knew something about why this this expense looked out of out of policy. And so you start having the the network becomes interesting when you have the finance team agents communicating with the assistant or various employees, and then behind there, you have other other sub agents. And so then you start seeing, like, more of a graph emerge.

But when you look at just what serves the employee, it looks more like a tree. Amazing. Well, I didn't know you were gonna go into that level of detail.

Yeah. Sorry about that. No.

No. No. No.

I'm I'm actually really glad I asked. Like, that is very impressive, and I hope you get more content about that. Yeah.

Absolutely. We're we're really excited about it. I think it's it's been it's been good to finally figure out a use for for agents and have the technology be as, like, as robust as it is to start realizing this vision because it's something that we we kinda dreamt of a couple years ago.

And the tech like, to your earlier point, the tech just wasn't there when we were trying to make the make the a similar concept work for the GPT 3.5. I was like, nope.

Right. We were hallucinating tool calls in back in that day. Awesome, man.

Thanks so much for joining us. This was fun. I really enjoyed it.

Happy holidays, guys. Thank you for having me. Thank you.

Shared via Hopper