At OpenAI DevDay 2025, Sherwin Wu and Christina Huang from the OpenAI Platform Team discuss the launch of AgentKit, a suite of tools for building and optimizing AI agents. Christina demonstrates building a customer support agent live in eight minutes, while Sherwin shares insights on the Apps SDK, MCP protocol adoption, prompt optimization, and the evolving AI developer ecosystem.
Hey, everyone. Welcome to the Laid in Space podcast. This is Alessio from the Kernel Labs, and I'm joined by Swix, editor of Laid in Space.
Hello. Hello. And we are here in the OpenAI Dev Days video with Sherwin and Christina from the OpenAI platform team.
Welcome. Thank you for having us. Yeah.
It's always to be here. Yeah.
It's so it's such a nice thing.
and this is, like, the first time that's been, like, so well organized that we have our own little studio podcast studio in in the Dev Day venue. And it it's really nice to actually get a chance to sit down with you guys, so thanks for taking the time. Yeah.
I feel like we Dev Day's always a process, and, like, we only had three of them, and we try to improve it every time. And Yeah. I actually I I know for a fact that I think we have this podcast studio this time because the the podcast interviews and the interviews with folks like yourselves last time went really well.
And so you wanna lean into a little bit more and glad that we're able to have this studio for y'all. We we were kneeling on the ground interviewing, like, Michelle last year. I don't know what you're saying.
I I I just saw it postproduction. I thought it was We we had to have people, like, cordon off the area so they wouldn't walk in front of the cameras. Yeah.
People would just come up. Hey. Good to I'm like, we're, like, recording.
We're it's nice. I guess if you guys have been to three, like, what what stood out from today, or what, what's your favorite part? I feel like the vibes are just a lot more confident.
Like, you are obviously doing very well. You have the numbers to show it that, you know, I just every year, in Dev Day, you report the number of developers. This year, it's 4,000,000.
I think last year was, like, three. And and I have more more questions about that and that kind of stuff. But also, like, just, like, very interesting, very high confidence launches.
And and then also, like, I think that this is the community is clearly much more developed.
in my mind. I don't know about you. Yeah.
And we were at the OG Dev Day, which was the DALI hack night at OpenAI in 2022, and think Sam spoke to, like, 30 people. So I I think it's just crazy to see the Yeah. Honestly, I I think it's, like, it's kinda similar to this podcast studio, which is I think we've had a number of Dev Days now.
We're we honestly were, like, slowly figuring things out as a company over time as well and both from a product perspective and also from how we want to present ourselves at Dev Day. At this point, we've had a lot of feedback from people. I actually think a lot of attendees will get an email with a transfer feedback as well, and we actually do read those and we we act on those.
And, like, one of the things that we did this year that I really liked were all of those, like there was, like, some art installations and, like, the the little arcade games that we did, which was, you know, came up with an via, like, engaging with the feedback from the Yeah. The arcade games are so fun. I love, like, the theme of all the ASCII art throughout.
This is my first SF dev day, but I've been to the Singapore one. That was actually my first week. That's when I spoke Yeah.
I saw I saw you there. That was my first week of OpenAI. They're really in the deep end.
They're on a plane to Singapore. Yeah.
Yeah. That's awesome. Well, so, you know, that's congrats congrats on everything, and, like, kudos kudos to the organizing team.
We should talk about some developer API stuff. Yeah. So the we're we're gonna cover a few of the things.
You're not exactly working on Apps SDK, but I guess what should people just generically take away what should developers take away from the Apps SDK launch? Like, how do you internally view it?
So the way that I think about it is I actually view OpenAI since the very beginning as a company that has really valued opening up our technology and bringing it out to the rest of the world. One thing we talk about a lot internally is our mission at OpenAI is to, one, build AGI, which we're trying to do, but two, you know, potentially, you know, just as important is to bring the benefits of that to the entire world. And one thing that we realized very early on is that we, as a company, it's very difficult for us to just bring it to every truly every corner of the world, and we really need to rely on developers, other third parties to be able to do this, which is, you know, Greg talked about the the the start of the API and, like, kind of how, you know, that was formulated.
But that was part of, you know, that mentality, which is we needed to we need to rely on developers, and we we need to open up our technology to to the rest of the world so that they can partake for us to really fulfill our mission. So the API, obviously, is a very natural, you know, way of doing that where we just literally expose API endpoints or expose tools for people to to build things. But now that we have, you know, ChatGPT with its, I don't know, like, 800,000,000 weekly active users, I forgot the stat that we Yeah.
Shared. I think it's, like, now the fifth or, like, sixth largest website in the world. And the number one and number two most downloaded on the Apple App Store.
Oh, yeah. With Sora. Yeah.
But that one, like, it moves around all the time, so it's kind of hard to celebrate and, you know, like You just screenshot it when it's good. Yeah. Yeah.
We definitely screenshot it and shared it earlier when it was when it was good. But kind of going back to my main point is, like, we've always kind of engaged with developers as a way for us to bring the benefits of AGI to the rest of the world, and so I view this as actually a natural extension of this. Candidly, we've actually been trying to do this, you know, a couple of times with last Dev Day with GPT, two Dev Days ago with, I'm sorry, two Devs ago with with GPTs and plug ins.
Plug ins, which was I think not tied to a Dev Day. So I view this as like, again, we we love to deploy things so iteratively, and I view it as like just a continuation of that process and also engaging deeply with developers and helping them benefit from some of the stuff that we have, which which in this case is ChatGPT distribution.
And when so Apps SDK is built on the MCP protocol. When did OpenAI become MCP build? I'm sure internally you must have had, you know, design discussions before about doing your own protocol.
When did you buy into it, and how long ago was that?
I think it was in in March, I wanna say. It's hard for me to remember kind of, like, the exact March was the takeoff of MCP. Yeah.
Yeah. So we we built the agents SDK, and we launched that alongside the responses API in early March. And I think as MCP was was growing, that felt like a really and, you know, we're building kind of a new agentic API that can call tools and just be much more powerful.
MCP was kind of like the natural protocol that developers were already using to bring all the tools into their system.
I think, like, in March is when we added in MCP to Agents SDK first and then soon after with kind of our other Yeah. I think there was, like, a tweet or something we did where it was, like, OpenAI, you know, is Yeah. There was definitely a moment.
I think there was a specific moment in a specific tweet. But what I will say, though, is, like and this is honestly, like, credit to the team at Anthropic that that kind of created MCPs. I really do think they treat it as an open protocol.
Like Yeah. We work very closely with, I think, like, David and and the folks on the, like, you know, consortium, and they are not, you know, really viewing it as this, like, thing that is specific to Anthropic. They really view it as this open protocol.
There is like, it is an open protocol. The way in which you make changes feels very open. We actually have a member of our team, Nick Cooper, who is sitting on kind of like that that steering committee for MCP as well.
And so I think they are really treating it as something that is easy for us and other companies and everyone else to to embrace, which I think they should because they they do want it to be something that is very embraced by all. And so because of that, I think it makes it a little bit easier for us to to embrace it. And, honestly, it's a great it's a great protocol.
Like, it's very general. It's already solved. Like, why would you make Yeah.
Yeah. It's very general. There's obviously still more to do with it, but it was very easy for us to, you know, integrate because of how how how how streamlined and how simple it was.
Yeah.
is, you know, like, I always see, like, in abstractly when you sort of wireframe a website or an AI app, it used to be that the initial AI integration on the website would be you have the normal website and then you have a little chatbot app. And now it's kind of inverted, where there's ChatGPT at the top layer, and then there's a ton of the website embedded inside of it. And it's kind of like that inversion that I honestly have been looking for for a little bit, and I think it's really well done.
Like, actually, all, like, the integrations and, like, the custom UI components that come up, you had, like, Canva on the Keynote Yeah. There, and it looks like Canva Mhmm. But, like, you can chat with it in in all your the context of your ChatGPT.
That is an experience I've never seen. Yeah. And I think that that's a kind of back to the iterative, like, learning that we've had.
That I think was because we've learned a lot from plug ins. So, like, when we launched plug ins, I remember one of the feedback that we got. I don't know if if, you know, people here really remember plug ins.
It was, like, March Oh, yeah. '23. Yep.
I'm like, one of the points of feedback was, oh, you can integrate we tell we told, like, you know, these companies that you can integrate these plug ins into ChatGPT, but they really didn't have that much control over how exactly it was used. It was really just like a tool that the model could call, and you were just, like, really bound by ChatGPT. And so I think, like, you can kind of see the evolution of our product with this.
And, like, this time, we realized how important it was for companies, for third party developers to really own and, like, steer the experience to make it feel like themselves, help them, you know, like, really preserve their own brand. And so and, you know, I actually don't think we would have gotten that learning had we not, you know, had all these other steps beforehand.
Awesome. Christina, you were the star today on stage with the agent kit demo. You had eight minutes to build an agent.
You had a minute to spare, and then you had so many shoes in the Yeah. With the download.
I don't know how much time I killed on that. Was extremely frustrated with the download screen. Was like, the UI bug is what, like, takes the demo down and be slow.
Yeah. Think it was a full screen, yeah, like, focus thing. But I heard the window wasn't in focus or something.
Yeah.
Maybe you wanna introduce Agent Kit to the audience. Yeah. So we launched AgentKit today.
Full set of solutions to build, deploy, and optimize agents. I think a lot of this comes from working with API customers and realizing how hard it actually is to take to build agents and then actually take them into production, hard to get kind of that confidence and the iterative loop and writing prompts, optimizing them, writing evals, all takes a lot of expertise. And so kind of taking those learnings and packaging them into a set of tools that makes it a lot easier and kind of intuitive to know what you need to do.
that today for people to try out and see what they build. Yeah. So I I had I find it hard to hold all the building blocks in my head.
But, actually, chronologically, it's really interesting that you guys started out with the agent SDK first. Mhmm. And then, then you have agent builder.
You have a connector registry. You have Trackit, and then you have the eval stuff. Am I missing any major components?
That that those are the main moving parts. Right? Yeah.
Yeah. Umbrella. Got it.
Got it. Got it. Yeah.
So, like, it's it's weird how it develops, and it's it's it's now become the full agent platform. Right? And I think one thing that I wasn't clear about when I was looking at the demo was it's very funny because, like, what you did on stage was build, like, a live chat app for Dev Days website.
Yeah. Did you get a chance to try it out? Yeah.
I tried to try it out. Was awesome. And, actually, I kinda wanted to ask, like, how did the Where's merch?
Yeah. Exactly. Was like, where'd you where'd you click the merch?
Anyway, I I and and this is very close to home because I've done it for my conferences, and, like, it's it's a it's a very similar process. But, like, I I think while it's not obvious is, like, how much is gonna be done inside of agent builder. I see there are some actually very interesting nodes that you didn't get to talk about on stage, like user approval.
Mhmm. That's, like, a whole thing.
you know, like, transform and set state, like, there's there's, like, a kind of like a Turing complete machine in here. Yeah. Yeah.
So, I mean, I think, again, like, this is the first time that we're showing agent builder, and so it's definitely the beginning of what we're building. And human human approval is, like, one of those use cases that we wanna go pretty deep on, I think. The node today that I showed is pretty simple, like, binary approval.
Approval object. It's similar to kind of what you'd see for MCP tools of, approving that an action can take place. But I think what we've seen with much more complex workflows from our users is that it's actually quite advanced, like human in the loop interactions.
Sometimes these could be over the course of weeks. It's not just kind of simple approval of a tool, there's actual like, decision making involved in it. And I think as we, like, work with those customers, we definitely wanna continue to go deeper onto those use cases too.
Yeah. What's the entry point?
export, like, just the segment, like, the use cases. Yeah. I mean, I so I think, like, the two reasons that you would come to agent builder are one, kind of more as a as a playground, right, to kind of model and iterate on your systems and write your prompts and optimize them and test them out.
And then you can export it and run it in your own systems using Agents SDK, using kind of other models as well. The second would be kind of to get all of the benefits of us deploying that for you too. So you can kind of use maybe like natural language to describe what type of agent you want to build, model it out, bring in subject matter experts so that you really have this canvas for iterating on it and getting feedback, you know, building data sets and kind of getting feedback from those subject matter experts as well, and then being able to deploy it all without needing to to handle that on your on your own.
And that's a lot of the philosophy around how we're building it with ChatKit as well. Right? You can kind of take pieces of it.
You can have a more advanced integration where it's much more customized, but you also get a really natural path of going live without like, with really kind of easy defaults as well. Yeah.
Do you see it as a two way thing? So I build here, I go to code, then maybe I make changes in code, and then I bring those changes back to the agent builder. Like, what's the Eventually, like, that's definitely what we wanna do.
So maybe you could start off in code. You could bring it in.
We'll also probably have, like, ability to, you know, run code in in the agent builder as well.
I think just a lot of flexibility around. The the one thing I'd say too is a lot of the demos that we showed today, I think were, like, you know, aired on the side of simplicity just so that the audience could kind of see it. But, like, if you talk to lot of these customers, like, they're building, like, pretty complex.
Like, you got to, like, zoom out on that canvas quite a bit to kind of, like, see the the full flow. And that and then for us, we, you know, we were kind of, like, working with a lot of customers who are doing this. And then, you know, if you turn that into, like, an actual agent's SDK, file, it's, like, pretty it's pretty long.
And so we saw a lot of, like, benefit from having the visual setup here, especially as the as the setup grows grows longer and longer. It would have been a little difficult to kind of showcase this, but even on, like, the minutes. Right.
Yeah. You can do it in eight minutes. But, like, even with some of the presets that we have on the Yeah.
we launched today as well alongside just like the Canvas is a set of templates that we've actually gathered from our engineers who are working in the field with customers directly of the kind of common patterns that they have in our own, basically like playbooks when we're working with customers on customer support, document discovery,
so kind of publishing those as well. Data enrichment, planning helper, customer service, structured data Q and A, document comparison, that's nice, internal knowledge assistant. Yeah.
Yeah. And I think, like, we just plan to add more to those as we can kind of build those out. I always wonder if there should be so you're not the only agent builders, but obviously, by default, being an OpenAI, you are a very significant one.
Any interest in, like, a protocol or, like, interop between different open source implementations of this kind of pattern of agent builder?
I think we we've thought about it, especially around, I'd say, agents SDK. I would actually say maybe even, like, zooming out a bit more from from just this is, like, yeah, we we, like we're also sitting here and kinda, like, observing, like, things being made over and over again. Even, like, besides, like, agent workflows, we're kind of launching what the industry is trying to do with responses, like, we've done with responses API, like, stateful APIs.
And so, you know, obviously, we were the first one to launch responses API, but, like, couple of other other people have kind of adopted. I think I think Grock has it in their in their API. I think I saw LMSYS just did something recently in Wells, but not, you know, not everyone.
And so, unfortunately, I don't have a great answer today of, like, yes or no, but we are kind of, like, assessing everything and trying to see, like, hey. You know, there's a there has been a lot of value with MCP, with hopefully, our with our commerce protocol as well. ACP, yeah, that's the I definitely did not forget the name.
stateful API integrations if they wanna use different models. Yeah. And I think that's one of the so it's not exactly a protocol, but one of the things that we launched today with Evals too is ability to use, like, third party models as well and kind of bring that into one place.
And so I think definitely kind of see where the ecosystem is at, which is, you know, using multi models and kind of having Third party models as in non non OpenAir models? Yeah. Yeah.
We're work with, Evals, Sergeant today. Yeah. Okay.
Got it.
Mhmm. Where we're working with them, and then you can bring your OpenRouter setup. And then with that, you can actually you know, you write your Evals, using our datasets tool or use our dataset tool to create a bunch of Evals, and you'd actually be able to hit a bunch of different model providers, you know, take your pick from wherever, even, like, open source ones on together, and see the see the results in our product.
Yeah. That's awesome.
Speaking more about evals. Right? Like, I think I saw somewhere in the the release docs that you basically had to expand the evals products a little bit to to allow for agent evals.
Maybe you can talk about, like, what you had to do there.
Yeah. I can't answer. Yeah.
I I was gonna say, so the I actually think agent evals is still a work in progress. So I think we've, like, made maybe 10% of the progress that we need here. For example, I think we could still do a lot more around multimodal Evals, but the main progress that we made this time was kind of allowing you to take traces, so the Agents SDK has really nice traces feature where if you run, if you define things, can have, like, a really long trace, allowing me to use that in the eval's product and be able to grade it in some way, shape, form over the the entire the entirety of what it's supposed to be doing.
I think this is step one. Like, I think it's good to to be able to do this, but I I think our road map from here on out is to, you know, really allow you to break down the different parts of the trace and and allow you to eval and, like, kind of, like, measure each of those and optimize each of those as well. A lot of times, this will involve human in the loop as well, which is why we have the the human in the loop component here too.
But if you kind of look at our Evolves product over the last year, it's been very simple. It's been much more geared towards this, like, simple prompt completion setup. But, obviously, as we see people doing these longer genetic traces, like, you know, how do you even evaluate a twenty minute task correctly?
And it was like, this is a really hard problem. We're trying to set up our Evalse product to move in that way to to help you not only evaluate the overall trajectory, but also individual parts of it. Yeah.
I mean, the magic keyword is rubrics. Right? Everyone Yep.
judge rubrics. Yep. Yep.
Yeah. Obviously, we're just this is above. Yeah.
Okay. Great. Yeah.
The other thing I think online, I see the developer community are very excited about is sort of automated prompts optimization, which is kind of evals in the loop with with with prompts. What's the thinking there? Where where is things going?
Yeah. So so we have automated prompt optimization, but, again, like, kinda I think this is an area that we definitely want to invest more in. We, I think, did a pretty big launch of this when we launched gbt5, actually, because we saw that it was pretty difficult as new models come out to kind of learn all the quirks about a new model.
Yeah. The prompt Right. There's, we have a big prompting guide, right, for every model that we launch, and I think building out a system to make that a lot easier.
We definitely wanna tie that in, like, completely with evals. We should be able to improve your prompts over time, improve your agents over time as well if they're made in the agent builder based on the evals that you've set up. I think we see this as a pretty core part of the platform of suggested improvements to the things that you're building.
I actually think it's a really cool time right now in prompt optimization. I'm sure you guys are seeing this too. It's like, not only are a lot of products kind of, like, gearing around this, so, like, kind of what we're we're thinking about, but I also think, like, there's a lot of interesting research around this, like GEPA with, like, the Databricks folks are actually doing really cool stuff around.
This. We're obviously not doing any of the the cool GEPA optimization right now in in our product, but would love to would love to do that soon. And, also, it's just an active research area.
So, like, you know, whatever Matej and the Databricks folks, like, might think about next, what we might, you know, think about internally as well. Whatever new prompt optimization techniques come out, I think we'd we'd love to be able to have that in our in our product as well. And, yeah, and and it's interesting because it's coming at a time when people are realizing that prompt.
You know, like like, I feel like two years ago, people were like, oh, at some point, like, prompting's gonna be dead. No. Like Yeah.
You know? And it's like, you know It's gone up. Yeah.
Yeah. Yeah. And then if anything, it is, like, become more and more entrenched.
And I think that, you know, there's this interesting trend where, like, it's becoming more and more important, and then there's also interesting cool working done to, like, further entrench, like, prompt optimization. And so that that's why I just think it's, like, a very fascinating, you know, area to to follow right now and also is an area where I think a lot of us were wrong, two years ago because, if anything, it's only gotten more important. Yeah.
Shen Yue used to work at OpenAI now now as an MSL. We call this kind of, like, zero gradient fine tuning or zero gradient Yeah. Updating because you're just tweaking the prompts.
Like Yeah. It it is so much prompt that is actually like, you end up with a different model at the end of it. There's a lot of, like, things that make it more practical too, just, like, even from our perspective.
and, like, it is extremely difficult for us to run, you know, and serve, like, all of these different snapshots. Like, you know, Laura's great. MSL just, you know or sorry.
ThinkingLab just just published John Truman just had a cool blog post about this. But, like, man, it is, like, pretty difficult for us to, like, manage all of these different snapshots. And so if there is a way to, like, hill climb and, yeah, do this, like, zero gradient like optimization via prompts, like, yeah, I'm all for it.
And and I think developers should be all for it because you get all these gains without having to do any of the Yeah. You know, fancy fancy fine tuning work. Since you are part of the API you know, you lead the API team and since we you mentioned Thinky, I gotta throw a cheeky one in there.
API?
So yeah. It's a good one. So it's actually funny when it when it launched, I actually DM'd John Schulman.
I was like Really? Wow. We finally launched it.
So the because you used to work with him. Yeah. Yeah.
So we it's it's actually funny. So at yeah. So right right when I joined OpenAI, like, this has actually been, I I I think, a a passion project of John's.
Like, he's been talking about doing something in this, like, in in this shape for a while, which is, like, a truly, like, low level research, like, fine tuning library. And so we actually talked about it quite a bit when he was at OpenAI as well. It's actually funny.
I talked to one of my friends who said that when he was at Anthropic, he also, you know, worked on this idea for a bit, and I think now He's a man on a mission. Yeah. I mean I mean, John's, like, so great in this regard.
He's he's, like, so purely just, like, interested in the impact of this because it's one, like, a really cool problem, and then, two, it also empowers builders and researchers. Like, you saw all the researchers who who, like, express all this love for Tinker because it is a really great great product.
I think he was really happy to kinda get it out there in the world as, as well. Yeah. It's this is probably this is very much a digression, but, like, it's weird.
Someone passionate about API design that it took this long to find a good fine tuning API abstraction, which is effectively all he wanted. He was like like, guys, like, I don't wanna worry about all the infra. Like, I'm a researcher.
I I just want these four functions, and, like, it's it's kind of interesting. Yeah. Yeah.
Cool. Before the OpenAI comms team barges in the room I know.
So what feedback do you want from people on, like, the agent builder? For example, the thing I was surprised by was the if else blocks not being natural language and using the common expression language. I'm sure that's something already on your on your road map.
What are other things where you're kinda, like, at a fork that you will love more input on?
I think, like, one of the things that we spent a lot of time discussing was, like, whether we want kind of more of, like, deterministic workflows or more LM driven workflows. And so I think, like, getting feedback on that, honestly, people model existing workflow a lot of what we did was kind of work with our team on especially with engineers who are working with customers, like modeling the workflows that already exist in the agent builder, and what gaps exist, what types of nodes are really common and how can we add those in. I think that would be the most helpful feedback to get back.
use in that type of API. So, like, more modalities, for example.
Yeah. I mean, I think, like, for sure, like, more modalities. Like, we you know, I think kind of voice would be, is already something that a lot of people have talked to us about even today at Dev Day.
So I think modalities for sure, but also more like the logical nodes of what what can't be expressed today. Yeah. Well, you know, you're you're building a language.
Right? You have common expression language, which I never heard of prior to this. I thought this was this Python?
Is this JavaScript? And then there was a whole link in there. Was that a a big decision for you guys?
Was you know?
I think that was more just kind of like a way that we thought we could kind of represent a mix of, like, the variables and Yeah. I don't know, like, conditional statements.
Yeah. Yeah. The other thing I'll also mention is that you let once you so there's a trope in developer tooling where, like, anything that can be that can store state will eventually be used as a database, including DNS.
So so be prepared for your state store to become a database. I don't know if there's, like, any limits on that because people will be using it.
It's actually funny. Yeah. I I I I heard this quote before, and there's definitely some truth to it.
I I don't know if our stateful APIs have become a database. I guess quite yet, but, like, who knows? Like, you know Yeah.
I mean, conversation Well, you charge for it. You charge for assistance.
Know? Storage. Yeah.
The storage. Yeah. Yeah.
Right? So there's some limit on that. But, like Yeah.
But it's very cheap. It's like I remember we priced it, like I think if you wanted to kind of, like, dump all your data somewhere, I don't know.
like, transforming it all into this shape. It's like the It's useful. It's easy.
Best place to order. But yeah.
what we try and do. So yeah. How do you think about the MCP side?
So you have OpenAI first party connectors. You have third party preferred, I guess, servers, you will call them, and then you have open ended ones. Do you see the that part of registry like functionality expanding, or do you see most of it being user driven?
Auth is, like, the biggest thing. Like, if you add Gmail and Calendar and Drive, you have to, like, auth each of them separately.
on what what's the thinking there? Yeah.
like, companies to kind of manage what their, like, developers have access to, managing kind of the configurations around it. And I think in terms of, like, first party versus third party, like, we wanna support both of those. We have some direct integrations, and then, I don't know, anyone can kind of create MCP servers.
I think we wanna make that a lot easier too, like, establish kind of private for for companies to use those internally. So I think, like, just really excited about that ecosystem growing. Yeah.
I I think one of the coolest things observed too is just I I actually think we we as an industry are still trying to figure out the ideal shape of connectors. So, I mean, part of why I think the one p connectors exist too, like, we we end up storing quite a bit of state. It's like a lot of work for us.
But, like, by having a lot of state on our side, we call them sync connectors, we can actually end up doing a lot more creative stuff on our side when you're chatting with IGBT and using these connectors to to kind of boost the quality of of how you're using it. Right? Like, if you have all the data there, you can do all this, like, re ranking.
You can, like, do we can put in a vector store if you want and put it anywhere else. Whereas and then so there there's some inherent trade offs here where, like, you you put in a lot of work to get these, like, one p connectors working. But because you have the data, you can do a lot more and get higher quality.
But then but then the question is, oh my god. There's, like, such a long tail of other things, which is where the MCP and, like, the third party connectors come in. But then you have the trade off of, like, you're beholden to, like, the API shape of the MCP creator.
It might actually work well, it might not work well with with the models. And then what happens if it doesn't work well? Then you kind of have to, like, you know, you're kind of, like, at the mercy of this.
And MCP, by by the way, is, like, really great because it already does some layer of standardization, but my sense is there's still gonna be more evolving here. And I think, you know, we wanna support both of them because we see value in both right now, especially working with working with developers who wanna have kind of, like, all options kind of on the table here, but it will be interesting to see how see how this evolves over time.
Yeah. When I saw about three, four months ago when you launched the forum for, like, signing with IGBT interest, I think, to me, that's kinda like the vision where I log in and I have the MTPs tied in, and then I sign in with IGBT somewhere, and I can run these workflows in that app where I'm logging in. So, yeah, I think Sam, you know, said in an interview that this is ChatGPT as, like, your personal assistant.
So, I think this is, like, a great step in that direction. Yeah. I think there's a lot more to to go in that in that direction.
right, which is a different role in the s in the off ecosystem.
Yeah. It's interesting because so so direct answer is like no plans right now, of course. But I actually think we currently have some version of this, which is our partnership with Apple.
Because with Apple, you can actually sign in to your ChatGPT account, and some of that identity does carry with you into your iOS experience with Siri. Right? Like, if you if you I don't if you've actually used this the the Siri integration.
I I actually use it quite a bit. But if you sign into your TryGPPD account, the Siri integration will actually use your subscription status to decide what type of model to use when passes things over to ChatGPT. And so if you're, you know, just a free user, you get, you know, the free model.
GBT 5, which is I think what they have. Also recently announced the partnership with Kakao. Oh, yeah.
Kakao's another one. Yeah. Where, I think you could fit a similar thing where you can sign in with ChatGPT.
directly there. Yeah. Mean, Sam's been talking about it for a while.
It's a very compelling vision. We obviously wanna be very thoughtful with I mean kinda how we do it. Know, now you have a social network.
You have a developer platform.
very, very valuable. Yeah. Yeah.
Exactly. Okay. So and then on the other side of off is something I was really interested to look at, and I couldn't get a straight answer.
Is there some form of bring bring your own key for Agent Kit? Like, when I when I expose it to the wider worlds, obviously, like, I mean, by default, I'm paying for all the inference, but it'd be nice for that to have a limit. And then if you want more, you can bring your own key.
Yeah. I mean, we we don't have something like that yet.
Yeah. But I think, yeah, it's definitely an interesting area too. Yeah.
It doesn't do it out of the box today, but, you know, developers have been asking about it for forever. Like, it's a really cool concept because then as a developer, you especially a new developer, you don't need to bear the burden of of inference. Yeah.
someone has to bear the cost. Like, sometimes you wanna mess around with, like, the different levels of responsibility.
Yeah. I will say in general, like, if you kinda look at our road map, we we we engage a lot with developers. We kind of hear what is the are the pain points, and we try and build things that address it.
You know, ideally, we're prioritizing in a way that's that's helpful. But, yeah, we we've definitely heard from a good number of developers that, like, the cost is or, like, all of the, like, copy paste your key, like, solutions right now, which are, like, huge security ish like, hazards because developers don't wanna bear the burden of of inference. You know?
Hopefully, we make the cost cheaper. So It's a model skip getting cheaper. Yeah.
So, hopefully, you know, that that helps. But, but but what we realized is as we make it cheaper, you know, the demand for that goes up even more and you end up, you know, still spending quite a bit. But, yeah.
Yeah. Do you see this as mostly, like, an internal tools platform, though? Like, to me, like, you've been doing a big push on, like, the more forward deployed engineering things.
It's almost like, hey. We needed to build this for ourselves as we sell into these enterprises. Might as well open it up to everybody.
What drives drive building these tools? Like, you think of people building tools and then expose or mostly on the internal side? Yeah.
I mean so, like, I think our again, our first deployment is ChatKit, which is kind of one of the it's intended to be for external users. But I think one of the things that we also did see a lot as we were working with customers is that a lot of companies have actually built some version of an agent builder internally to kind of manage prompts internally, to manage templates that they're sharing across, know, the different developers that they have, maybe the different product areas. And we're seeing that kinda, like, over and over again as well and, really wanted to, like, build a platform so that this is not, you know, an area that every company needs to invest in and, like, rebuild from scratch, but that they can kind of have a place where they can manage that these templates, manage these prompts, and really focus on the parts of agent building that is more unique to their, like, business.
It is interesting too. Like, from a deployment perspective, it is like, it it has spanned both internal and external use cases. Right?
Like, kinda like these internal platforms, people use it for end like, data processing or something, which is an internal use case. But if you saw some of the demos today, like, there have been a huge number of companies that are trying to do this for external facing Mhmm. Use cases as well.
Customer service is one template. Customer service, the, like, ramp use case. We use this internally and externally.
Like, our customersupporthelp.open.com already powered on agent kit and then various other, like, internal Yeah.
Use cases as well. And one of the things that I actually think the team has done a really great job of so, like, Tyler, David, and Jiwon on the team, They built the the especially the the chat kit components. They built it to be, like, very consumer grade and and, like, very polished.
Like, you kinda look at the there's, like, a whole grid of, like, the different widgets and things that you could create there. Like, ideally, people see it and they're like they see it as, like, these very polished, like, consumer grade ready external facing things versus, like, you know, I think of internal tools and, like, the UI is always, you like, the last thing that people care about. But, like, you really, you know, push the team, and I think they did a really great job of making the ChatKit experience, like, really, really consumer grade.
And it should feel almost like ChatGPT or Yeah.
responsive designs and all of that. Yeah. I think your point on widgets is, like, definitely, like, really resonates, right, because ChatKit handles the chat UX, but we're also just building, like, really visual ways for you to represent, like, every action that you wanna take, and that is definitely, like, very high polished.
Yeah. And when working with customers, like, those have been the the most helpful customers for us to work with because, you know, when Ramp is thinking about, you know, how what what they want to publicly present to people, like, they have a pretty high bar, as they should, as as as well as, you know, all the other customers that have been iterating on it. And so that kind of feedback from our customers has really helped us uplevel the general product quality of of the launch that we had today as well.
Yeah.
Would you ever would you open source Checkit?
Talked about it. Uh-huh. We've talked about it.
There are a bunch of trade offs.
I think so so Checkit itself is, like, an embeddable iframe, and so I think the actual Iframe. Yeah. Right.
And so that helps us keep it like evergreen. Right? So if you are using Track Kit and we come up with new, I don't know, a new model that reasons in a different way, right, or kind of new modalities that you don't actually need to rebuild and, like, pull in new components to use it in the front end.
I think there's parts of, you know, widgets, for example, that is much more like a language and can definitely it's something that is easier to explore that for as well as kind of the the design system that we've built for ChatKit.
yeah, more evergreen Yeah. More evergreen experience that is a pretty opinionated Like, there'd be no point in being open source. Right.
Because Yeah. Yeah. You won't do it.
Then you don't get the benefits of it. You know, being Stripe alums, like Stripe Checkout, like, it's it's auto opt optimized for you to, like So I'm not a Stripe alum, but Christine is.
Stripe Checkout. It's yeah. So it's very similar philosophically.
Right? So Stripe, you know, can build elements and checkout, and not every business needs to rebuild, right, the pieces that are really common. And I think we see the same with chat.
We see chat being built over and over again, especially as we kind of come up with new, you know, modalities, like reasoning, everything.
yeah, to the developers. Does it feel I mean, I know WordPress is, like, a bad connotation in, a lot of circles, but to me, it almost feels like the WordPress equivalent of, like, chat is like, hey. This is, like, drop in thing, and then you have all these different widgets.
Do you see the widget becoming a big kinda, like, developer ecosystem where people share widget? Is that kinda like a first party thing? And then No.
What's like the MVP versus widget forest? No. Exactly.
it seems great for people that are, like, in between being technical and, like, not really being technical enough. Yeah. Yeah.
Yeah. I mean, I think that's a big part of building widgets. Right?
Like, it's real already kind of in the language that is very, consumer friendly. You can use in our widget builder already already. You can kind of use AI to create those widgets, and they look pretty good.
I don't know if you guys have gotten a chance to try that out yet, but definitely see kind of, I don't know, a four s If haven't haven't tried out the Widget Studio and and the demo, like Yeah. Yeah.
Apps as well. Yeah. Remember You got a custom domain, like widget.
studio. It's cool.
Actually, don't know how we got that. But Yeah. Everything's in chakka.
and then we have, like, the playground there so you can try out what chakka would look like with all the customizations. We have chakka.world, which is a fun site we built.
I've I was, like, spinning the globe for a while this morning. It was And then Kasha widget spinner. Kasha also, like, uploaded some of her Yeah.
Solar system stuff and Yeah. Yeah. Yeah.
All the demos as well. Yeah. And then that's where, like, the widget builder Yeah.
So so it's, like, it's really come together. Like, it's taken, like, almost more than a year to, like, come together and, like, build all this stuff, but it's coming together. It's, like Mhmm.
Really interesting. Yeah. It's something that we, like like You definitely planned all of this upfront.
Oh, yeah. Yeah. Yep.
We have the master plan from, you know, three years ago.
No. But, like, I think especially on this stuff, I think there was, like, an arc of a a general, like, you know, platform that we did want to kinda build around, and it takes a while to to build these things. Obviously, Codecs helped speed it up quite a bit now, but it yeah.
I will say it does seem great to kinda, like, start start to have all the pieces start fitting together. Yeah. I mean, you saw we launched Evals, and we got the fine tuning API for a while, and and we laid all the groundwork for for some of the stuff over the last year.
And we're hoping that we can eventually, you know, make it into this this this full featured platform that that that's helpful for people. I think you have.
you have since you did the codex mention, maybe a quick tip from each of you on codex power user tools or or tips.
The so there's there's actually a a funny one that one of the new grads has has, I think, like, taught our team in general. And I think this is, a a point for, like, just how, like, new grads and and younger, you know, generation people are actually more AI native. So one of them is to, like, really lean in to, like, like, push yourself to, like, trust the model to do more and more.
So, like, I feel like the way that I was using Codecs, and so for me, it's most usually for my my personal projects, they they don't let me touch the code anymore. But you give it, small tasks. So you're, like, you're you're, like, not really trusting it.
Like, I I view it as, like, this, like, intern that I, like, I really don't trust. But what a lot of the, like so we had an intern class this year. What a lot of the interns would do is just, like, full YOLO mode, like, trust it to, like, write the whole feature.
And it, like, it doesn't work work for worse. It, like, doesn't work sometimes, but, like, I don't know, like, 40% of the time, it's just, like, one shots it. I actually haven't tried this with, like, codec g b d five codecs.
I bet it I bet it probably, like, one shots it even more. But one tip that I'm, like, starting to, like I feel like undo this like like, relearn things here is to, like, really lean into, like, the AGI component of it and just, like, really let the model rip and, like, kinda trust it. Yeah.
Because a lot of times, they can actually do stuff that surprises me, and then I have to, like, readjust my priors. Whereas before, feel like I was in this, like, safe space of, like, I'm just treating this I'm giving this thing like a tiny bit of rope. Yeah.
And and because of that, I was kinda limiting myself with how effective I could be. Like, sure, but okay.
But, also, is there an etiquette around submitting effectively, you know, vibe coded PRs that someone else now has to review. Right? And it's like, it can be offensive.
Reviews now.
Okay. It actually reviews itself. Does Codex approve its own PRs a lot more than humans?
Oh, it doesn't doesn't get approved then, but it I was gonna say, I think, like, the Codex PR reviews are actually one of, like, the things that my team, like, very much relies on. I think they're very, very high quality reviews. Yeah.
On the Codex PR side, like, for the visual agents builder, we only started that probably less than two months ago. And that that wouldn't be possible without Codex. So I think there's definitely a lot of use of Codex internally, and it keeps getting better and better.
And so, yeah, I think people are just finding they can rely on it more and more, and it's not, you know, totally vibe coded. It's still, you know, checked and edited, but definitely as a kicking off point. And I think I've heard of people on my team, it's like on their way to work, they're like kicking off, like, five codex tasks because the bus takes thirty minutes.
Right? And you get to the office, and it kind of helps you orient yourself for the day. You're like, okay.
Now I know that the files, I have the rough sense. Like, maybe I don't even take that PR, and I actually just, like, still code it. But it helps you just context switch so much faster too and be able to, like, orient yourself in in a code base.
There are so many meetings nowadays where I have, like, one on ones with engineers, and I walk into the room. They're like, wait. Wait.
Wait. Give me a second. I gotta figure out my, like, OdeX thing.
I'm like, oh, sorry. Yeah. We're about to enter async zone.
Notes. Right? You're like, let me And they're, like, typing, like, okay.
Now we can start our one on one because now it's great. Yeah.
Cool. We're almost out of time. I wanted to leave a little bit of time for you to shout out the Service Health dashboard because I know you're passionate about it.
Oh, yeah. Well, tell people what it is and why why it matters. Yeah.
So this is a launch that we actually didn't, you know, it didn't get any stage time today, but it was actually something I'm really excited about. So we launched this this this thing called the service health dashboard.
You can now go into your usage or your settings account and see the health of your integration with our OpenAI API. This is scoped to your own org. Basically, if you have an integration that's running with us doing a bunch of tokens per minute or queries, it's now tracking each of those responses, looking at your token velocity, TPM that you're getting, the throughput, as well as the responses, the response codes.
And so you can see kind of like a real time personal SLO for your integration. The reason why I care a lot about this is obviously over the last year, we've spent a lot of time thinking about reliability. We had that really bad outage last December, you know, longest, like, three, four hours of my life, and then had to, you know, talk to a bunch of customers.
We haven't had one that bad since, you know, knock on wood. We've done a bunch of work. We have an Infrared team led by Venkat, and they've been working with Janna on our team, and they've just been doing so much good work to get reliability better.
And so is we actually again, knock on wood. We're we think we've got reliability in the spot where we're, like, comfortable kind of putting this out there and and kind of, like, letting people actually see their their SLO. And hopefully, you know, it's, you know, three, four, soon to be five nines.
But the reason why I cared a lot about it is because we spent so much time on it, and we feel confident enough to kinda have it behind the product now. Five nines is, like, two minutes of outage or something. Yeah.
Yeah. We're we're working we're working to get to to five nines. Yeah.
What is what is an extra nine take? It's it's exponentially more work. So, you know, and then but, like, we always we were you know, in the last couple weeks, we're talking about, like, hitting three nines and hitting three and a half nines and then hitting four nines.
But, yeah, it's it's exponentially more work. I could I could go for a while on the on the different different topics. But We'll have to do that in a in a follow-up.
I mean, mean, that's all that's the engineering side. Right? Yes.
Yes. Yes. Like, you're serving 6,000,000,000 tokens per minute.
We actually zoomed past that. Yeah. That's the that's the That's outdated.
Yeah. But, yeah, it's been crazy, though, the growth that we've seen.
Awesome. I know we're out of time. It's been a long day for both of you, so, we'll let you go, but thank you both for joining us.
Yeah. Yeah. Thanks for having us.
Thanks.
Thank you. That's How was that? That was great.
Okay.
We have the mic software. The thing I I I didn't want to say on the podcast was on the on the Tinker thing.
Shared via Hopper