Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI

Latent Space: The AI Engineer Podcast
28 July 2026 1h 9m
0:00 --:--
Episode Description
There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right.A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M million users combined (as we cover in the pod,

Summary

Akshay Nathan from OpenAI discusses the evolution and vision behind ChatGPT Work, highlighting its journey from Codex to a 'super app' for productivity that blurs the lines between personal and professional use. The episode covers the challenges of enterprise AI adoption, the importance of AI agents and artifacts, and the future of human-AI collaboration in product development and daily life.

Chapters

Full Circle: From No-Code to AIAkshay Nathan reflects on his career journey from no-code/low-code platforms like Walrus and Airtable to leading core product engineering for ChatGPT Work, seeing it as the ultimate no-code solution.
OpenAI's Mission and Enterprise LearningsAkshay discusses OpenAI's startup culture and consistent mission of bringing frontier intelligence to everyone, sharing insights from ChatGPT Enterprise about the diverse use cases and the challenge of teaching users how to leverage AI.
The Genesis of ChatGPT WorkThe decision to launch ChatGPT Work stemmed from the surprising adoption of Codex by non-developers at OpenAI, revealing a 'superpower' for a broader audience beyond just coders.
Product Positioning and Harness EvolutionAkshay explains ChatGPT Work's positioning for 'worky' or productivity-related tasks, including personal use, and details the shared underlying harness between Codex and ChatGPT Work, with UX differences for specific modes.
Model Guidance and ArtifactsThe discussion shifts to model selection, advising users to stick to the best default configuration, and highlights the significant improvements in artifact generation, such as editable Excel-like spreadsheets and hosted sites.
AI-Powered Research and CollaborationThe host shares a personal case study of using ChatGPT Work and 5.6 to generate a playable game site and research panel, illustrating the power of AI in auto-research and creating complex artifacts for collaboration.
Balancing Simplicity and CapabilityAkshay discusses the design challenge of balancing simplicity with the vast capabilities of AI products, aiming to show users what's possible without overwhelming them, and the growth strategy of 'show not tell'.
The Future of Productivity and Personal AIAkshay outlines the sequencing of AI adoption from developers to general knowledge workers and eventually to everyone for personal productivity tasks like meal planning, envisioning a future where AI agents are universally useful.
Power User Advice and AI in ReviewsAkshay advises power users to broaden their imagination about AI's capabilities and provide more context to models, giving an example of AI's surprising utility in gathering context for performance reviews, while emphasizing human oversight.
Memory Systems and OpenClaw InspirationThe conversation delves into ChatGPT's memory system, its improvements, and the inspiration drawn from projects like OpenClaw for persistent computer environments and personal automation, aiming for a unified, extensible agent experience.
Sub-Agents and Harness DesignAkshay discusses the design philosophy behind sub-agents, balancing the display of their parallel task execution with avoiding information overload, and the ongoing challenge of managing memory across diverse, long-running projects.
Building in the AI Era and Measuring ProgressAkshay reflects on how AI has dramatically changed product development, accelerating the idea-to-reality loop, blurring team roles, and shifting the focus from traditional proxies to measuring 'at bats' and distinguishing between 'motion' and 'progress'.

Topics

ChatGPT WorkCodex adoptionAI agentsEnterprise AIProductivity measurementAI in game designMemory systemsSub-agentsOpenAI cultureNo-code platformsModel capabilitiesArtifact generationHuman-AI collaborationProduct developmentPersonal automation

People

Akshay Nathan (guest) Vibhu (co-host) Gabriel Chua (mentioned) Samir (mentioned)
Key Concepts (16)
No-code/Low-code — The idea of enabling more people to build things without writing traditional code, which Akshay Nathan started his career with and sees ChatGPT Work as an evolution of.
Super App — A single application that integrates multiple services and functionalities, a term used to describe ChatGPT Work's ambition to be the ultimate no-code platform.
Frontier Intelligence — OpenAI's mission to build Artificial General Intelligence (AGI) and make it accessible to everyone, acknowledging that this vision will not be a linear progression.
One Size Fits All in Enterprise — The challenge in enterprise AI adoption where a single solution doesn't meet the diverse and specific use cases of different teams and departments, requiring tailored approaches.
Show Not Tell — A product strategy to demonstrate AI's capabilities through direct experience rather than explicit instruction, helping users discover new use cases organically.
Harness Engineering — The development and evolution of the underlying system or interface that allows users to interact with AI models, specifically comparing the classic ChatGPT harness with the newer Codex harness.
Artifacts — High-fidelity, editable outputs generated by AI agents, such as spreadsheets or hosted websites, which enable users to collaborate and iterate on complex tasks.
Multiplayer Collaboration with AI — The future potential for multiple users to collaborate directly with AI agents on shared artifacts and projects, rather than individuals acting as intermediaries.
Token Billionaires/Maxing — A concept referring to users who consume a vast number of AI tokens for complex tasks, often in personal projects or auto-research, pushing the limits of model capabilities.
Auto Research — Using AI to automate the research process, such as generating benchmarks, tuning hyperparameters, and creating research artifacts like interactive sites for analysis.
Productivity (OpenAI Definition) — OpenAI's mission for its productivity team is to enable people to do things they couldn't do before, giving them leverage in both knowledge work and personal life to free up time.
AI in Performance Reviews — Using AI agents for gathering context and highlighting achievements for performance reviews, acting as an 'agentic search' tool to assist managers, not to solely generate the review content.
Memory System (ChatGPT) — ChatGPT's ability to learn and retain context about a user across sessions, making interactions more personalized and feeling like 'their ChatGPT'.
Chronicle (Memory System) — An experimental feature that acts as an input source into ChatGPT's memory by learning from a user's computer activity, building deeper context and surfacing relevant insights proactively.
Motion vs. Progress — A distinction in measuring productivity where 'motion' refers to easily generated activity due to AI tools, while 'progress' requires deliberate focus on achieving specific, validated goals.
At Bats (Product Development) — A metric for product development teams to measure their ability to efficiently generate ideas, build, get feedback, react, validate hypotheses, and move to the next idea, emphasizing both quantity and quality of iterations.
References (6)
Walrus
Airtable company
Amazon company
Strata
Wealthfront company
OpenClaw project
Transcript (127 segments)
Speaker 1

Okay. We're here in the studio with Akshay from OpenAI. Welcome.

Thank you. And with our trusty cohost, Vibhu, we so you recently launched ChatGPT work. You lead core product engineering.

You It's been a long journey into into all this. I find it very interesting that you started with no code or low code with Walrus and Airtable. And to some extent, ChatGPT Work is kind of like the super app of super apps of, well, here is the ultimate no code.

You just write a prompt.

Speaker 2

Yeah. Yeah. It's it's it's funny how things come, like, full circle.

I mean, I think for a long time, my career I mean, I started my career working in consumer fintech, but then after that, like, there's this hypothesis that, you know, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more accessible way, then that would be truly magical. We were working on a startup. It's actually funny, like, before LLMs, before vision LLMs on how to do automated testing with AI.

And I was just kinda jank back then, but, you know, doing what we can. And then and then worked at Airtable for a while on, you know, the same thesis that, like, if can bring a database or the primitives behind a database to people, that would be really useful to them. But once, I think, LMs came onto the scene, it became clear that, like, this was, like, the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what's going on underneath the hood.

And so, like, I think this launch and, you know, a lot of the stuff that we've been up to is, like, the manifestation of that. How was stuff when you joined? So you joined OpenAI twenty twenty three.

Speaker 3

Now we've got, you know, so much more stuff.

Speaker 2

Codex app, ChatGPT for work. Have things changed? Actually, I think the more interesting thing is how things haven't changed.

Like, I guess, like, one, I I joined I remember when I joined, it was, like, 500 people. One thing I was worried about was, like, I was looking for something, you know, more early stage and, like, was it gonna feel start up enough? And I joined, and was like, good.

This feels even more start up y than I could ever imagine. And, like, that that really hasn't changed even till now. I mean, I think the, like, level of, like, bottoms up ambition and, like, the ability of anyone to, like, you know, do anything or have an idea and and and ship it is is really cool.

But on the, like, sort of mission side, I think what was really compelling to me is this mission of, you know, bringing frontier intelligence to everyone. Like, building AGI and then bringing it to everyone. And I think acknowledging back then that, like, that vision is gonna, you know, not be a linear progression.

Like, we're probably gonna, like, try different products and and have different things that that succeed and don't. But the vision has stayed the same, and the mission has stayed the same, and we're starting to see the pieces fall together, and and that that's really cool. You worked on enterprise.

Speaker 1

What a lot of people never touch chat chat GBT for enterprise, god. What is something that you learned from there that you're bringing into your work now?

Speaker 2

I think how there's no, like, one size fits all solution in enterprise. I remember in the early days of ChatGPT Enterprise, like, when we talked to customers and, like, everyone that was, like, when I think it was a year after ChatGPT was released, and everyone was so excited to bring, you know, AI into their enterprise. And, like, all there's all these teams stood up that were being stood up as, like, you know, the AI deployment team with, like, these enormous budgets.

And if you asked anyone, like, what were they excited about? Like, what were they excited about solving? Like, first, you'd get, like, you know, kinda like the the baseline answers of, yeah, we have all this context and data and all this stuff.

But then if you ask them, like, you know, what was, like, a discrete use case that, like, they want AI to enable in their in their workplace, you get such a different, like, variance, like, explosion of different types of answers. And it's interesting, like, you know, you using all these models and in these products, you you have this box and you can say anything to it, which is the magic. But it's on the flip side, it also means that, like, you don't know what to do with it.

And in enterprise, I think a big part of that is, like, actually meeting the users where they are, like, use case were they trying to solve, and then actually teaching them how they can use AI to, like, gain leverage there. Do you meaningfully differentiate that from forward deployed engineering? Or I I think there's like the the, like, go to market side of it.

Yeah. And then there's like the product side of it. Think With neo you need to see more of the product side.

Yeah. And I think, like, however good we get at FDE Motion, like, I think at the end of the day, if we have a user who's, like, looking at their computer or looking at their phone, like, it's our job in the product to, like, be enabling them and showing them where to go. So we're really excited about that.

Do you think there's been changes, you know, over the past three years of adoption? So there have been, you know, step function changes, you have reasoning models and whatnot.

Speaker 3

the same problems of enterprise has black box, don't know what to do with it, or have things changed?

Speaker 2

I mean, we're seeing now that, like, there's this huge uptake. Right? Everyone's are extremely excited about it.

It feels like, you know, many peep like, hundreds of millions of people are using ChatGPT. They understand, like, how generally to work with AI. But then, like, every time, like, a new capability gets unlocked.

So now, like, we're seeing with agents, like, there is probably a contingent of, like, early adopters still who, you know, truly get it, who are like, you know, you can do anything. You just have to make sure the right context is there. It's connected with the right tools, and then you're supervising it, but, like, anything is is is possible.

But then there's, like, this, like, 10 x or a 100 x bigger market, or, like, they don't yet get that or they don't yet see that. And so I think that's the next stage here. So I guess to answer your question, like, I think the adoption is there and and growing fast, but I think the opportunity is, like, far, far bigger than that.

That's where we wanna play, especially with ChatuchPety work.

Speaker 1

Yeah. Well, let's let's skip ahead to ChatuchPety work. Only only, like, a month ago or so announced, What was the sort of decision process that led into it?

Speaker 2

merging of the super app. Is is that what we're officially calling it? You deprecated the browser as well.

Just, I guess, summarize your last, like, couple of months of working on this thing. Yeah. And it feels like forever now, but I guess it's only been a few months.

I think maybe the one one, like, impetus that, like, is most alien is when we release Codex or even internally add Codex. Like, it was really surprising to us. I think we recently put out some stats on this, that there's this, like, real inflection of, like, adoption among non developers at OpenAI.

And I, you know, through this product development process, like, would go to, like, these UXR sessions to talk to people internally. And the thing that stuck out to me is, like, one, like, you know, you you go talk to, like, strategic finance or marketing or whatever, they're all using Codex for, you know, their use cases. That that per that part's cool.

But the thing that really stuck out to me is how proud people were that they were using Codex. Like, how like It's like I'm not supposed to be using it by the It it was that. It was like that they were, you know, early to this, like, new thing, but it was also this thing of, like, they felt like they had a superpower.

Right? And what we recognized then is that, like, the power of Codex, power of agents, like, we already have this massive distribution base of people who have, you know, come to know and love ChatGPT. Like, how do we show that to them?

Like, how do we bring it to them? Which is like a hard product problem, and it's like a tricky tricky thing. Alright.

There's many ways you can go about it. And so that's what we call the merge and the super app over time, and and and ultimately launch it in ChatGPT work is how do we do that. But it came from that initial realization that, like, the the power was not only for developers, like, much much earlier than probably even we thought.

Like, it could be extended to to everyone.

Speaker 3

products differently? So, like, who is it for? Right?

So Codex sorted out even CLI, then app. Now there's a merge of ChachiPT Codex and ChachiPT Work. So is it the opening for the average user, for enterprise, for work?

How how do you position it?

Speaker 2

related things, for lack of a better word. Right? I think productivity is like a is is actually what, like, the pillar that I that I support.

Like, that's the name of the team. And the reason for that, the reason we call it productivity and not, like, you know, enterprise or or, like, work or something like that is because there's also personal productivity. Right?

And, like, think Charjibouti work is I've seen people do things in their personal lives that you wouldn't classify as like work technically, but like these agents are, you know, super capable for. Like one one recent example that someone posted about on our Slack is like someone has like a missed package. Like, they didn't receive it, and then they got like the picture of it, you know, from Amazon or whoever the courier was, and they like asked how JBT work to like find out where that package is.

And like the agent, you know, is extremely tenacious and like like took the image and like looked at a bunch of like listings around neighborhood. I figured out exactly the apartment complex in which the package was. I gave them some information.

And so, like, I think there's all these things that, like, you, you know, worky or productivity related things. I think that's what we want the product to be. You asked about Kodex.

I think we think Codex is, you know, a durable brand. But we have a principle that, like, the user you know, we don't want a user to get stuck in a tab or an experience where they don't get the power of the product. And so, like, basically, everything that you can do, you know, in the Codex portion of the product on on desktop, can do chat to your work and vice versa.

But we made some opinionated product decisions on, like, you know, how much of the git state, if you're in a git repo, do we wanna expose to the end user? Or how much do we wanna make the the experience of seeing the agents thinking, like, diff forward so that you're you get exposed to the diffs out of the and then, like, on the safety side, like, how do we wanna think about, like, sandboxing and making sure that we have the right defaults in one one state versus the other? So there's, like, some opinions that go behind that, but we do we do want we don't want the user to need to choose which experience they're in.

That is a good goal for AGI. Right?

Speaker 1

to choose what version of AGI they want. They just want the AGI to decide for them. Can I get an answer or like, it's not super clear to me?

Is the Codex harness and the ChatGPT work harness the same? Is it just UI affordances, or are they actually prompt level or even even deeper differences? So the harness is the same.

The harness is shared.

Speaker 2

On in both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plug ins or computer use or artifacts. You get that power regardless of which experience you're in. On the UX side, there's opinionated takes that we have when you're in codex mode, what the UX should be how the UX should behave, and some stuff around the sandbox like I mentioned.

But the underlying harness and capabilities should be the same. I think I'm just kinda curious. Maybe we can is there a query that we can run that would look different in the the two modes?

Yeah. I tried to create, like to ask it to create, like, a retirement calculator spreadsheet or something in in both in both modes. And then in codex mode, you might have to be in a in a repo for this, but you'll see, like, the diffs of, like, the the sheet that it's creating and stuff like that and the file edits.

Speaker 1

But in Ruck, you won't be able to see that. I think that's that's super clear. And then also the other thing I wanted to dive into was your the productivity team.

What else is there? First of all, you know, what what what are the top level teams other than productivity? Isn't productivity everything?

Speaker 2

You know, we have a team focused on ChatGPT, like the the core chat experience for consumer, which is like, you know, not I think all productivity, like there's people are using ChatGPT every day for search to, you know, figure out how to write messages to loved ones, to think about how to, like, learn a new topic, etcetera. And so there's so much more inside to create images. There's so much more in chat that, you know, the hundreds of millions of users are using that, you know, obviously, that that warrants, like, a a very dedicated effort.

And there's teams focused on enterprise and infrastructure and API and stuff like that as well.

Speaker 3

Will bring Yeah. So I have them both running. Yeah.

This is work. There's a codex version here. I picked 5.

6 SOL, so this will take a while. Uh-huh. I think I think we'll just keep it in the background, and, know, as as they finish, we'll look into some of the differences.

Yeah.

Speaker 2

That it assumes it assumes Git. Yeah. Exactly.

Yeah. Like, Dynamic Island assumes that you're in a Git repo.

Speaker 1

And you might miss some stuff because some of it is, like, in the actual chain of thought. What what were changes and how we display that? But Is is there an unintuitive like, is there a thing that you wanted to ship and then you got feedback and you're like, no.

Let's not do it. Like, what's the thinking behind that?

Speaker 2

In Chateaubrii Work? Yeah. I think one direction we could have gone with this is, like, keeping the experiences, like, completely separate.

So it's like, why Different apps. Exactly. Like, different apps or even in the same app, like, different completely different experiences.

Like, why merge it all? Like, what is, you know, Codex, obviously, people love? Like, why why bring these products together?

And I think the intuition here is that, like, all of our jobs are, like, changing dramatically with AI. Like, for, like, every few months, like, I feel like I wake up and I'm, like, doing a completely different thing than I was doing a few months ago. And my my hypothesis here is that alright.

So our hypothesis is that, like, part of what we're we're building in this this technology is giving people leverage, like, you know, the things maybe it's the more mundane parts of your job or or parts that, like, if you were able to automate, you'd able to share more ideas faster, or whatever, like, you're able to do now. And because of that, like, that might actually blur the lines between someone who's like only writing code, or creating strategy docs, or, you know, planning events, or helping with marketing, or doing podcasts, or whatever. Right?

And so like, these things are gonna get blurred over time. And so, like, trying to draw a hard boundary based on, like, the who you are is is gonna be is gonna be tough. And, we should enable users to choose, but we shouldn't box them in.

And so a lot of the work that went in here, like, know, keeping the primitives the same. Like, for example, plugins are, like, unified across this product and ChatGPT in the cloud was because of that. It's this this thesis that, like, eventually things are gonna come together.

And and we don't wanna be like, wanna be prescriptive about when to be in neither experience, but we don't wanna box anyone in. I wonder if there's users who are very tuned to the old ChatGPT harness that is effectively now replaced by the the Codex harness. I can't imagine what that was, but maybe they're more the more conversational side.

Speaker 1

Can you con compare and contrast the the two harnesses because only you've seen it?

Speaker 2

Yeah. I I mean, I think ChatGPT, the the existing harness, like, still exists today. It, like, exists in this app.

The classic. Right? The You just start a new chat and you don't go under at work.

Right? Yeah.

Speaker 3

another. Yeah. But I guess, you know, on instance.

Speaker 1

Yeah. So this one's not gonna code, or it's gonna be in line. It's not in line in a sandbox.

Speaker 2

Actually we try to push you to to to go to work if you're creating a spreadsheet. And this way, this is a router decision. Sorry?

It's a router decision? This is the decision that, you know, the model is making, and then, like, you know, it sees that you're able to or you're trying to do something that would be better served in work mode. But I I think your question was, like, what what are the advantages of, like, the the chat, like ChatGPT, chat harness?

Speaker 1

I wanna basically do an oral history of harness engineering. Right? The ChatGPT harness lasted us from, let's call it the o one era until now, and now it's being replaced by the Codex harness effectively.

And they're they're overlapping somewhat, but I'm curious what changed if, you know, if there is.

Speaker 2

My perspective on this is, like, there's there's there's all there's sort of like a constant process of, like, divergence, convergence, divergence, convergence. And in chat, like, many of the use cases I was talking about before, like, you know, search or learning, I think we're we're really optimizing for latency and optimizing for personality and, like, different things that over time, like, the product the reason people love ChatGPT is because we've been optimizing for those things and working on them for so long. Codex, what we learned was that, like, if you give the agent access to this infinitely flexible environment as a computer, it can do really, really powerful things.

And so when we think about, like, okay, well, for for knowledge work, like, what is which mode should we choose? It was like it felt more natural to us to bring that to this, like, computer environment and, you know, maybe abstract some of the details of this computer away from users who might not be used to that, but, like, give them that same power. But ultimately, I think that we want the power in in all places.

Right? We wanna meet people where they are. So I'm sure there'll be work down the road in order to to get things to be equivalently capable in in all scenarios.

Speaker 3

what we've been focusing on the product on historically and what we're focusing on now. I think alongside that, outside of just harness and when to use codecs, JTP, or work, there's also the new models you've released. Right?

Any guidance there? So people love to min max what to use, like, use Terra on high reasoning versus for this, you you wanna use SOLE here, ignore There's 32 options. Yeah.

Yeah. Yeah.

Speaker 2

what's what's the advice? Right? Well, mean, I think before the advice, like, first thing is like, none of this would be possible without these models.

Like, I think you asked earlier, like, you know, what was like the inspiration for work? Like, you know, early on, like, I I I mentioned like what we were seeing with Codex, but that was also because the the models were getting infinitely more capable. That's happening again.

I think it's like another step function jump now. And to answer the question on device, like, we want this default to be the best possible like, we wanna be opinionated about the default. And so we've we've chosen a default that we think is gonna be the best for everyone.

And, you know, we have, for power users, options under the hood. We could one could argue that there might be too many right now, and we're, you know, working on simplifying it. But you can extend, you know, the reasoning level, and you can change between the different model classes if you need to.

But the default should be the best for for most most use cases. So my advice to most people would be to stick to that. And then, you know, if you reach a situation in which you think that you could you wanna try a different configuration if you're not seeing either the the efficiency on the on the cost side or or the the quality on the intelligence side, then you can change the defaults and see if you can get something better.

But but we think the the default should be good enough. I have a I I'm just gonna run something by you since you you have way more experience than me.

Speaker 1

but with goal. With the idea that the goal basically augments the reasoning effort, but with more terminations and turns. Mhmm.

Is that a good way to think about it? As opposed to Soul Ultra or Soul, you know, extra high. Yeah.

It's hard to say because Yeah. It's like an interaction effect. Exactly.

Speaker 2

on, you know, for you as an individual, like, how do you like to collaborate with the models? Like, how many of those, like, terminations, as you call them, do you want where, you know, you can steer or make sure that it's doing the right thing? I think generally, people should try whatever works for them.

I think that, like, using Ultra or the, like, multi agent setups are best for, like, when you have, like, tasks that are either incredibly complicated, like, open expirations or very paralyzable. I think even for tasks using goal, I think, is best for for tasks that you know that you'll be able to make consistent progress in a way that's verifiable over time. But I think for most tasks, they actually don't fall into either of those buckets.

And so, like, at least when they're starting.

Speaker 1

and then seeing, like, where you wanna go from there. Right. You guys worked on a slider, which actually is super helpful for reducing the amount of panic.

Yeah. Yeah. It's nice on mobile, at least.

There's a nice slider here. I haven't tried it. So you you you have you have the advanced view there, but if you click advanced view yeah?

Yeah. Oh. Just a nice slider.

Yeah. Very very pretty, very colorful.

Speaker 2

Yeah. That idea was here here was, like, reduced it to, one dimension, even though there's multiple dimensions. Right?

Try to project it onto a single dimension for the user. Yeah. Like, you know, something from that represents, like, know, speed and efficiency on one side Yeah.

And then, like, sort of, like, quality and thoroughness on the other side.

Speaker 1

I am just puzzled that it uses Soul so much. Like, the lower No. I think spider, if I'm not mistaken, is Terra.

Oh, it is. Yeah. It's e personal.

Nice. So they they preset Terra to only be the light one. I see.

But, like, I think a lot of people actually would more people should use Terra. One, because Seoul keeps running out of capacity.

Speaker 3

I'm the reason. You know? Here's ten minutes of our retirement calculator.

Oh, that's the Excel thing we're doing. Oh my god. Look This is work, and then Codex is still cooking, so we'll get back into it.

I think it'll be interesting to actually see the thought process, the reasoning. And also, you know, I guess this is eight minutes on work. Codex is still cooking.

Speaker 1

Yeah. And by the way, I so I've do you know Gabriel Chua? He's part of the OpenAI Singapore team.

He showed me this, and I was, like, pretty shocked that this looks like Excel. Yep. It it edits Excel files.

You never paid an Excel license. Right? Like but somehow this this is, like, kind of workable, and it's agentic Excel.

Speaker 2

Yeah. I mean, one of the big, like, pushes that we made for this launch was, like, artifacts. Right?

Like, both on the on the model side, like, I think if you compare this with with 5.5 and and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of these artifacts, and then also on the product side.

The UX side is also crazy. Like, hosted sites and whatnot, no no longer needing to host your own little web page, like Oh, I have a story about that.

Speaker 1

a separate thing? I'll need to take the the the visuals here, but we'll we'll we'll cut to that later. Was there co training, I guess, because you are moving making this big move, and you launched five six on the same day as ChatGeePsy work, was there influence between the model training teams and the harness teams, or do they did the launch dates just happen to line up with the same day?

Speaker 2

heavily with the research teams. And, I mean, I think that's like one of the most magical parts of the job, like the the most fun parts of the job. But, yeah, I mean, just just using artifacts as an example, like, know, a lot of what you're seeing, like, underneath the hood, there's a lot of work that went into making sure that, like, you know, we had the right infra to be able to train the models to get better at this.

And then on the product side, like, had the right experience for users to be able to collaborate with the model on an artifact like this. In fact, like, whole viewer, like, the intuition here is that, like, you know, it's not necessarily that you wouldn't need an Excel license. This is stage one.

Right? Like, is probably not what you meant when you're like making a retirement calculator. You wanna iterate.

And like, when you when you're seeing it and if this thing is high fidelity to like what you'd actually see in in or what your your coworkers would see if you were to send this to Sean. Like, that that, I think, makes it so easier and makes you trust the the the product in in terms of iteration.

Speaker 3

multi team collaboration with artifacts? Any any things you guys think about? Share it.

Right? Yeah. Yeah.

It's it's something that, you know, we're actively thinking about.

Speaker 2

One thing that, you know, we've noticed internally without talking too much about the road map is that, like, I there's many times when someone will ping me about something, and I will ask Tragic Beauty Work the question, and then I'll ping them back the answer. And then I'll be thinking would be, you know, the three of us are just all on one hosted. Exactly.

And I'll think about, like, you know, like, was I required in this loop? Or and then maybe it was, you know, rephrase, like, what they were asking or pulled from certain context or whatever. But, like, you know, when I gave them back the answer, that process was also lossy.

Right? Like, I gave them just, like, my interpretation of what Chachibiji work cooked up. But, like, underneath the hood, there's so much context, like, in the rollout and stuff that that could be interesting.

So, like, the answer was preemptively respond to every inbound request?

Speaker 3

No. And it's just, like, literally, like, this is what I do sometimes with my job. I know you copy paste, and then you you know, you're just a message forwarding service from AI to AI.

But I think it's interesting. Right?

Speaker 1

oftentimes people don't realize until they try or someone shows you, and then you're like, oh, okay. Okay. I see.

Yeah. I think it's also there's also like a light security issue where, like, basically, you're the permissions layer.

Speaker 2

yes, I could query everything that you query, and I could get an automated response, but maybe I'm not supposed to see it. Yeah. And that Yeah.

There's no way I would know because I'm not supposed to know what I don't know. Especially as, like, you know, with ChaiDB works, asking you to connect your plugins and, you know, it's pulling from your local files and stuff like that. Like, the the amount of context that the agent has access to is, like, deeply personal, and, like, that's something that, like, we need to preserve.

So that'll be definitely a challenge.

Speaker 1

Excel, there's PowerPoint, there's Docs, you know, the grand trio of work. What other formats of work do you do you think about? You know, like, obviously, you worked on Airtable.

Is there a future where there's, like, OpenAI Airtable? Like, know, like, what what what does that look like if if you ever ended up doing it? It's a really good question.

Speaker 2

I mean, one that you didn't bring up was Sites, and I think that was There's a part of this one one side of Sites that I think people commonly talk about, especially on Twitter and stuff, or X, of, like, you know, this sort of, like, prototyping tool. And actually, like, we saw that happen with this launch even. The the model slider that you guys were referencing earlier, like, that was developed almost fully in a site.

Like, you know, the the collaboration between design and engineering and product on that was, like, on a site where we play with, you know, the the the affordance and and and figure out how it feels and and all of that. But the other aspect that that I think is a little bit less talked about is, like, sites as, like, an an artifact for for knowledge work. I was actually talking to someone the other day who was on, like, our our corporate finance team, and, like, they're mentioning how, like, now when they have these reports that they're they're working on as a team month to month, historically, things were in in slide decks and in spreadsheets.

And now they're just in Sites. Like, Sites is the mechanism that they collaborate across the team. And the reason is because it's, I think, it's, somewhat higher bandwidth.

Like, you know, at these tools, like PowerPoint and Excel are, like, infinitely flexible, but at some point, you reach the boundary of, like, either as a human, you may not know how to use some feature or something, or the product itself doesn't support it. But with a site, you kinda do anything. You ask for anything, and you can get that.

Once people see that magic, I think it's been really valuable. Yeah. Let me show you my my case study.

This involves all the hot topics, including ChatGripT work, but also 5.

Speaker 1

token billionaires, and token maxing, and sites, and auto research. I'm a fan of this game called Strata. It's basically it's like a little board game that you that you play with physical blocks that come on top of it like that.

So over the weekend, I I I took, like, 30 photos and just threw into ChatGPT. 1,700,000,000 tokens later, out comes this site with a fully playable thing with with with three d block placement and everything because it requires physical blocks, and I needed friends to train on it so they can get better so I can play against them. But also, I could also do things like train an AI on it, and that that's the That's your auto research.

That gets into auto research. So you want to train your own AIs and then make sure they self play against each other. I need to set both AIs.

So this is AI versus AI, and they're they're gonna self play. Obviously, that the AI start out bad, and then you want to define a loss function and and get good. I wasn't gonna supervise all this.

I was at I was at Dallas San Mateo attending a conference. What I ended up doing was auto researching in in on this and creating benchmarks, and that there was just way too many parameters for me to to read. So I started asking it for a site, and it's created this this this this lab panel.

Where is there is there a is there a shortcut for a site that is created?

Speaker 2

You should be able to go in the sidebar to sites, top of the sidebar. The left sidebar. This one?

Oh, left? Yeah. I just scroll all the way to the top.

Oh. Oh, it says sites. Yeah.

Oh, there you go.

Speaker 1

Yeah. Oh. So it create it creates the sites.

I don't I don't think this is a it is exactly what I what I wanted, but let me let me show you what it popped up. Right? Like, I think as a as a research artifact, it is very important to communicate exactly what is being done.

Outputs this this thing which I eventually started publishing. So I I moved it off of Sites because I wanted more database and infrastructure than Sites afforded me, but this is this is, like, a research output that you can start to mess with and, like, try to think about, like, what hyperparameters are you tuning for training your AIs. And, like, I was trying to make, like, scaling laws and everything and doing all sorts of, like, game optimization stuff.

And the fact that you can just kinda throw this up as a research artifact, like, I no longer need to read ChatGPT output. I read site output. But then there's also a huge sprawl.

Like, look at how long this thing is. There's so many numbers. It is pretty overwhelming.

So then I have to start pruning from there.

Speaker 3

transition from markdown Yeah. Activity that you're putting out to yeah. You're putting out a whole functional site.

Think markdown just isn't that optimal for people to read. Right? Might as well just write HTML website.

And I don't know, I think you can do a lot with customizing this, right? You have your skills that explain what you want. Like I noticed they're quite verbose.

I don't need a lot of this information. It's very verbose. So, and then the nice thing of having a site side by side is, you know, you just iterate on what you want and what you don't.

Right? Yeah. I I don't know if any that triggers any stories for you of how it's run internally.

Am I doing this right? Yeah.

Speaker 2

is now becoming a site. And, like, with a site, you because it's just HTML, you can, like, it's infinitely flexible. And so, you know, if you if you want to give more prominence to a certain thing that, like, in the slide deck would, you know, feel like it was buried, like, you can do that.

You can have it be like the hero image. Right? And so I think that, like, people are starting to see that.

There's obviously more work to be done to make these things, like, much more easier easy to collaborate on. You mentioned that they're very they're long and verbose, could be broken up. I'm sure that there's something Super long.

Yeah. Yeah.

Speaker 1

that's, like, much more flexible than than what they ever had before. I think your job also comes becomes kinda meta. You're not designing the products, you're designing a product to make products.

Speaker 2

And I'm curious how you manage that. I I think one thing that we've been like, when we look at the UX, like, that that we've been thinking a lot about is how can we balance, like, simplicity with capability? Like, if we if we're designing a product, like you said, that, like, is is made to make up build other things.

Right? You can build so many different things. But we can't put that all in front of you because you'll get overwhelmed.

Yes. And so we had similar problem or or similar challenges even chat with ChatGPT. But especially now, like, when there's so much that can be done, I think the balance that we're constantly trying to strike is, like, how can we give the user enough of a UI surface where, you know, they can be expressive, they can they can tell the the agent what they need, they can verify that it's using the right tools, it's pulling from the right sources, etcetera.

But then it gets out of the way. And then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is gonna be like, how do they discover the next use case?

And the next one after that, if they really want to, you know, to to be super powered by AI.

Speaker 3

Yeah. It's interesting. I feel like everyone also just has a different way to do it.

Right? I made a similar version of this, same game. I didn't take any pictures of board or rule game.

I threw an at goal eighteen minutes fifty three seconds later. A lot of tokens later, I've got a similar version. Obviously, not with all the auto research and whatnot, but, you know, you gotta do all the latest trends.

And, yeah, I did it with did it with Codex, not work, but it's interesting. Right? Yeah.

Speaker 1

And this is obviously GPT image generating the

Speaker 3

pro avatars. Very good for game design. Like, a lot of game designers were, like, really into GPT image.

I will say, like, the the broader takeaway probably is the reason that we do this is more so just to test the tools. Right? Like, this was also a test for 5.

6 came out. I had done the game on 5.5.

Right? The ability for me to no longer need it to I had defeated the rules. It's it's a pretty niche game.

It couldn't find how to do this on its own. Oh, yeah. 5.

6 It is auto distribution. That's why I was also very keen on testing the 5.6 capability.

But, you know, this is just as as work comes out, as new things come out, these are just our side ways to test things. Right? Yeah.

It's some kind of private eval, I guess. Yeah. Is not all this private.

But also valuable, because now you can send this to your friends. And I mean, I learned about this game through seeing this. It's a hard game.

He's very good.

Speaker 1

It's good to when no one is competing with you. But, yes, it's a classic RL problem of, like, self play, bootstrapping your game AI. Yeah.

You see how easily work becomes personal, and personal becomes work?

Speaker 2

which I I imagine is a growth strategy. Yeah. The the show not tell is a big piece that, you know, I think we've we're not still not fully cracked of, like, you know, showing people all the the things that they can do with the product versus, like, trying to teach that to them through, like, you know, articles or onboarding or whatever.

Yeah. Submitting them in the moment.

Speaker 1

right, where your job is to show. And then you're like, what do you mean? You you don't you don't need actually, your job is to tell.

Mhmm. And then and then but the product people are like, well, we don't need you if our product is intuitive enough.

Speaker 2

Yeah. I mean, that's the magic of the the models. So you can tailor the telling or the showing to, like, specifically what the user needs, like, what what they care about, what they've done in the past, exactly where they are on the adoption journey.

So I think that's, like, gonna be a super big opportunity.

Speaker 3

custom showing. Right? People have different use cases.

As much as you said you don't wanna segment different people into different buckets. Right? It's also not that hard too for people that are in different categories.

But the question, I guess, is you said your team is more broadly on what was the term you used? Productivity? Productivity.

Yeah. Productivity. Which is now work, basically.

Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different than ChatGPT codex or work?

Is there is there more that the mass isn't targeting?

Speaker 2

I see it as like a sequencing. Like, you know, the the vision is, like, bring useful agents to everyone. We started with, like, developers.

Like, developers historically are, like, early adopters. They're willing to put up with more friction, set things up, etcetera. Like, that's where, you know, Codec started.

I think the next opportunity is, like, sort of what we call general knowledge work, you know, all the other functions around developers. I think when when you go from developers to this segment, like, there's inherent challenges, obviously, with, like, you know, this this show not tell thing that we're talking about, making the product more understandable, bringing in new capabilities that matter more for for this cohort, that matter for developers, things like artifacts, things like computer use, etcetera. And then I think, like, the the same learnings, like, similarly how we took the learnings from developers and brought it to, you know, general knowledge work.

Next The stage will be like taking the learnings from general knowledge work and bringing it to everyone no matter what they're doing in their lives. And we're already seeing that a little bit. Like, this this game example that you have is, you know, something that's like on the border of like fun and personal life to to, you know, your your professional life.

I use ChatTPG work full time at home for everything, like, for for whatever I'm doing. I used it the other day to come up with a meal plan and, like, you know, save that on on the, like, computer environment that it has and something that I can continue going back to. Like, is everyone doing that yet?

Probably not because the thing says work on it. But eventually, you know, we wanna get people there. Tragedy beauty life.

Yeah. Exactly. Tragedy beauty cooking.

But I think there's a lot of there's a lot of opportunity there. But I I see it as, you know, we're we're built we built the foundation in in in software engineering, and we're gonna take the same learnings that we take from software engineering to knowledge work, knowledge work to to everyone. Do have any power user advice?

Speaker 3

there's a group of people that will live it, use it for everything, stay on it twenty four seven. Yeah. And then there's a bit of a gap between that crew and people that, you know, okay, I use it for work, I use it occasionally, sometimes I pipe questions.

Any advice, any learnings, anything you recommend, or just, you know, takeaways that you found that help bridge that gap?

Speaker 2

that it really helps to broaden your imagination of what's possible. And this has been a learning even for me. Like, you know, the technology has progressed so fast that, you know, it was something that, like, even three months ago, was like, no way no way that models can do this.

Like, now it's like, wow. It's like, it actually can. Like Give an example.

We're going through right now there are, like, review cycle internally. And, you know, people always talked about this as like kind of a a thing that the the models are good at. Like, you know, there's a cliche of like, okay, like, no one wants to be writing reviews, and, like, we just use AI to do it.

But I mean, and then In all seriousness. Evaluate it as well. Yeah.

Exactly. In all seriousness before, it was like just like slot basically. And, like, I I think it was helpful, but, you know, not super productive.

Now I've found out, like, the model can do a much, much better job of me, especially in this environment of, pulling context on, like, what people are up to, how they've like the things that they've done to make a difference, highlighting, like, you know, wins that they've had that, like, I might may not even have seen, you know, has access to, like, everything. Right? Like, the code, like, you know, things that they've caught, reviews, Slack, everything.

And so it's, like, incredibly powerful in that domain. And, like, just, like, six months ago was the last time we did this cycle. Like, I didn't even I tried using it, but it was not at all helpful.

And this time, it's been, like, incredibly helpful. And, like, so I think continuing to push the the frontier of imagination of what's possible, even if you tried something before, I think, is maybe the my biggest piece of advice. The other, I guess, thing is, like, the more the you more put in, especially in this environment where, like, you know, the model has access to to everything on your on your computer or in Chargebee's work, like, you can create, you know, artifacts over time and save them in your library, like, the model will continue having access to those.

The Like, more information you give it about whatever domain you're in, whether it's your life or your work, the more valuable it becomes. And it'll become both valuable in ways that might surprise you. It might pull from context in a way that may be proactive and that you might not even have thought about, but it needs to have access to those to that those tools or that context first.

Speaker 1

that's a very sensitive thing. And you're you're a founder, you've managed people, you've hired people. As manager myself, I'm very reticent to put out any LLM generated things, especially when it comes to people, because it feels like you don't care.

Mhmm. Presumably at OpenAI, people are obviously more open to being basically rated by GBT.

Speaker 2

unofficial rules around this? Like, what's the etiquette? Oh, I mean, I think the etiquette is that, like, I wouldn't ever write something via, like, only solely via AI and, like, present it as, like, a review for someone.

What I was talking about is more like gathering context. Yeah. That's the place where it's So it's just search.

It's agentic search. Yeah. It's like agentic search, but, you know, that you can tailor and steer much more capably than you could before.

And because like the the thing is it's all there's sort of a flywheel happening. Right? Because of Codex, people are able to do, and because of Chargebee Orr, people are able to do so much more now than ever before.

And if you're able to do so much more, it's easy to miss things as well.

Speaker 1

you know, where where we can be helpful. I think the the the thing, like, obviously, I I run a small company, so easy to search. But at the scale of OpenAI, with the amount of messages that you guys put in Slack, do you think that it misses things?

Speaker 3

Probably, but I think that I also miss things. Like, it doesn't matter. I think sometimes.

It needs to be human level. It's all relative. Right?

Yeah. Sometimes it's nice when it finds things you wouldn't. Right?

Like, now, my codec system prompts, they're set up in such a way that every project I have has a separate notes MD, and it just writes learnings to there. And then the the global one can pull from all these. So sometimes it'll be like, oh, there's this project you did like four months ago.

Here's a note that we had, and it randomly pulls it back into context. I would never do. I haven't thought about it.

And I'm like, okay. This is quite superhuman. Right?

Like, stuff that would and, you know, it'll save, like, hours on chunking of stuff or find something that's already been done. I'm like, as much as it might miss stuff, I would too, but it's very useful when it finds stuff. And I have like a very, you know, non super engineered solution to this.

It's just markdown files that get pulled whenever they want. Yeah. I actually have a funny anecdote about this.

Speaker 2

gearing up to this launch, you know, the team has been, you know, really cooking on it for for for a couple months. And over that time, like, there's so much conversation chatter going on in Slack and Docs and elsewhere. And one one of the members of the team set up this scheduled tasks, like automation, to, like, look at everything that's going on and, like, come up with the best memes, and then post it in one of our shared channels.

And, like, there are two cool things about this. Like, the first is, I think the models are, you know, over time, like, actually starting to become, like Funny. Funny.

Nice. Whereas, like, you know, a year ago, like, that was not at all the case. The second is that it was what you were saying.

Like, they find things that in surprising ways that you may not have not have thought of and, like, create connections that you may not have thought of. And that really helps with like the meme generation because then you can see something that, you know, genuinely surprises you and then and is funny in that way. So, yeah.

Speaker 1

use of this the the technology, but it it does it doesn't cover this, like, this capability that's emerging, which is just like defined information that you otherwise would not know of. Talking about the the launch, I I think I have pretty much said this is the most successful launch in a long time. I think even more successful personally than five point o, and you're announcing 10,000,000 users.

Does it feel different?

Speaker 2

You've been through a lot of launches. I think it feels like a culmination. Well, I I think two things.

One, it feels like a culmination, like I was mentioning earlier, like this, like, vision mission that we've been on for a long time. Like I said, we saw the magic of Codex internally, and then we're, like, extremely excited to bring this to many more people and to see it working, to, like, see us reach, you know, the distribution goal I mean, numbers that you mentioned. Like, I think that's, like, huge and and and super exciting.

The flip side of that is, like, there's so much more to do too. Like, that's also really exciting. Like, you know, ChatGPT as a whole, like, the the this product that, you know, everyone almost equates to AI and, like, loves, you know, has hundreds of millions of users.

And so, like, 10,000,000 is really cool, but, like, we we need to get this to everyone. Like, we need everyone to feel this magic. And so that's the next step from here.

But, yeah, I think extremely pumped about how how it's going so far and and the opportunities.

Speaker 1

Awesome. I did want to also because I've I've I've been tracking the the number closely. It transitioned at some point from just Codex users to Codex plus ChatGPT work, obviously, because it's the same harness.

The whole point is that you don't you can you can't count them separately.

Speaker 2

ChatGPT users? What is it just jump to 1,000,000,000 right away? Like, isn't that the default on ChatGPT or no?

We don't default you into ChatGPT work if you're on ChatGPT. If you're free. Yeah.

It's also only available to to paid users right now. And and then I think there's, like, a process of, you know, educating users of what is the value of this product, having them try learning from their feedback, and making it better over time. But, I mean, the goal is to, you know, get as many of of the people who who love ChatGrBT today to, like, feel the feel the power of ChatGrBT work.

But I think it'll be a journey. Yeah. Codex will still be alive as a brand for the foreseeable future.

Yeah. And we'll just toggle between them as as as needed for UI stuff. Yeah.

I think it's an even stronger point than that. Like, I think we fully intend to like, you know, treat developer like, developers have been know, a core market for us for so long. And, like, there's there's so much more that we can do to make Codex great specifically for software development, and we'll continue to do that.

This doesn't take away from that at all. If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff to creating an artifact or, you know Yeah. Doing a search for refactor.

Speaker 1

terminology leaks to the nontenical user. Like, do they have to learn to say artifacts if I want artifacts?

Speaker 2

Or, you know It's funny. Like, we call artifacts internally because that's what the teams call it internally. Like, no one says that.

No one calls it an artifact.

Speaker 1

like, describe things whatever they're used to. Right? So if, you know, Chatty Work is good at creating slides, they'll say Chatty Beauty Work is good at creating slides, and that's actually what we want.

One big another I mean, it's July 2026. One big thing that also happens in for OpenAI was OpenClaw, and that's, I think, a lot of people's first time really maxing an agent for personal stuff, but also crossing over to work in in a some same way. As far as as far as I understand, OpenCLaw is still independent.

But did you go through your own OpenCLaw moments?

Speaker 2

or back? Whatever. I think there's a lot of inspiration.

I I did go through my own OpenCLaw moment.

Speaker 1

Yeah. Tell the story.

Speaker 2

my my wife, like, set up an OpenClaw to, like, try to manage everything in our house. Not that there's, like, a ton, but it was, like, actually quite useful. We gave it a calendar and started, you know, creating events for us and stuff.

At some point, the the laptop they were running on had died, and then never got a chance to to pick it back up. But there's a lot of inspiration there. Like, you know, in Charjibouti work in web and mobile, like, you you get access to, like, persistent computer environment where, you know, you can store files and those files stay around between sessions.

And the idea is to be able to enable use cases like this. One of the members of our team actually uses ChatGBD work for what they used OpenClaw from before, and I feel like it has, like, completely transitioned, which is, like, workout planning and, like, meal tracking, which, again, it's like a work y thing. Right?

It's, like, not work necessarily, but it's, like, in personal productivity space. But it has all the same primitives. So it has scheduled tasks.

It has the ability to store files on a file system, it has the ability to like reference those things over time. And so you start to see the same types of use cases emerge, which has been really cool.

Speaker 1

ChatHTT Work completely replaces OpenCLOA?

Speaker 2

Obviously, they're independent. So Yeah. I mean, I'm I'm not close to it, so I can't speak to Yeah.

The OpenCLORE roadmap, but I I don't think so. I think that there's gonna be you know, there's always a need for, like, this, like, incredible, like, open source technology that that that team has built. And I think that we can draw inspiration in the product and, you know, ChatGeBT I think many more people have, like, heard about and used ChatGeBT than have you have used OpenCloud.

And if we can take the magic from OpenCloud and bring it to them, I think that'll be a success. I think that, like, one thing on the ChatGeBT work side that we feel strongly about is that, like, the core experience is that you come to this product, and you have a conversation, start a session, whatever you wanna call it, with this agent. And the magic of the product is that you can do anything in that moment.

And we would like to create a product where you don't have to click a button or to go to a different place, whatever. And you can get whatever functionality exists in, you know, your your finances app or or any other product, like, in this in this one place. And so that's the goal.

It's like, it we want an extensible system with plug ins where you can connect to the tools that you need in order to be able to accomplish, like, a financial task where you can, you know if if you're doing, like, science work, like, we have an ability to, like, extend the system and such that you can, like, write the tech and it and performs well. There'll always be, like, products that we support that are best in class at those things, but we want as much of the magic as possible in that core experience. Yeah.

Do you think that you can do everything you used to do with Wealthfront in ChatuchPity Finance? I actually tried it. I mean, like, ChatuchPity doesn't yet custody cash and and assets for me.

So so that part, no, not yet. But I I mean, there was, like, a whole component of, like, retirement planning and sort of, like, financial planning and budgeting and stuff that that we were looking into when I was there. And, like, with the finances plug in, like, that's all possible in terms of today.

So I feel like at least that component's replaced for me. I haven't really plugged it in yet.

Speaker 1

I was somewhat scared to look at the answer.

Speaker 2

Like, that's honestly, like, the same reason for health and finances. Like, oh my god. Don't know.

It's really good. I mean, it's it's really cool how I mean, we were talking about, like, the agentic search aspect a little bit earlier, but, like, it's really cool how, like, you know, in in conventional UX, like, if the more power you wanna give to a user, the more, like, knobs and bells and whistles you need to add. Like, for like these finance and budgeting apps, like, there's always like a bunch of the different filters and like search bars and stuff like that.

But like now, like, with the right connectivity to the right data, you can have whatever you want. Like, you can ask any question you want Mhmm.

Speaker 3

And into that box and get the answer, and I think that's super powerful. I think it's also nice to just have it centralized in one space, right? You have different health apps.

I have one for a smart scale, a watch, all these different things.

Speaker 1

co locate it. Which is, you know, part of the whole thing of OpenCLI. Right?

Like, that that you would have personal OS, which presumably ChatGPT wants to become. I do think that just relying on, like, just in time pooling of data for, let's say, via MCP, CLI, API, whatever you whatever you do, still not enough. Like, I, like, I come from a bit of a data engineering background.

Like, you still want, like, a data warehouse or some kind of caching or semantic layer.

Speaker 2

do you already have that? I can't speak to, like, all the details on how everything works, but I think it depends on the access pattern. Right?

Like, if you want an answer immediately, then, yes, it's very difficult to do that. You need to pull from all of these sources. But a lot of the, like, use cases that we wanna enable in ChatGitBuddy work aren't necessarily something that you need immediately.

It's more like a task that you want the agent to go and do, and that's gonna take a certain amount of time. And, you know, with things like programmatic tool calling and stuff now, like, some of that time and sub agents and stuff like some of that is also parallelizable. And so it's possible I I think it's very possible that there's a the ceiling on what can be done, you know, with MCPs and, like, calling out to these certified services has has been raised Yeah.

Substantially.

Speaker 1

So we're really excited about that. You mentioned some agents. I I I gotta double click on that.

Ultra is a new mode. You have special affordances in ChatGPT itself to show off the the agents. Can't really do much with them, to be honest.

I just just just watch. What have been what have been your experiences?

Speaker 2

issues that you would call out to other builders building with sub agents? I think it's sort of goes back to the balance that I was raising earlier about, like, you know, showing builders the power of the tool, but also creating enough of an abstraction to not overwhelm them. I think sub agents, the thing that we wanted to show is that you can take a task that, you know, has many parallel tracks or is is complicated in a way that, you know, sub agents can handle, and this product is for you.

Like, the the model can can accomplish those goals or try to accomplish those goals. And so, like, that's the point of, like, showing them in the product, and and and that's where we've we've gone with the design.

Speaker 1

with information. And so this is like the deliberate trade off that we made for now. I mean, you do display quite a lot of transcripts.

Right. Right.

Speaker 3

more than that? No. No.

It's Some some people could want more. So I'm one of those people that will basically throw a lot of stuff at goal, pretty much every goal, I'll tell it to use sub agents. Seems redundant.

Right? But every time I'm like, okay. Use sub agents where possible.

And I have a lot of people, a lot of friends that recommend and do the same. Whereas I'll sometimes talk to people that are like, okay, this is where I want you to use sub agents for this sub task. And I'm sure they would appreciate seeing into how they're being used.

For me, it's primarily like two things, right? One is net time efficiency. So span out across sub agents.

Two is probably cost. Right? Right.

Don't use big expensive model. Offload to a lot of smaller, cheaper models. And some people want that level of control.

So if you have repetition in what you're doing. Right? Say I want something built where I want it to consistently do this every day, I might wanna go in and fine tune sub agents here, sub engines there.

So you can see both, but I think if I'm not mistaken, it's hidden by default. There's a dropdown that goes a lot where I'm like, okay, just gonna keep, You you know can change the model that they use? I know I tell them to be steered.

I'll say my I know Anthropic offers this in Cloud Code. You can tell Fable to use Sonnet or Opus to use Sonnet as sub agent. So pretty trivial thing.

You tell it to span out sub agents with Sonnet, it's cheaper, faster. I would assume if it's not there, it could be built there. But I think there's a side of There's too many toggles.

It's not a toggle, actually. It's just utility check. The the way I do it is prompt it.

Right? And I think this is something that gets abstracted unless it's something you built for repetition. Right?

So if I'm building something, say that's a podcast prep, right, research into people, do a very, very deep extensive research, that I might wanna configure to cheaper, faster model just for web search. Right? I can see a world in which you want both.

I think the default is actually pretty good right now where it's hidden, but you can drop down and get some more info into what's done. I know people talked a lot about it on five point six's launch. This thing loves to use a lot of sub agents and causes the ChatGPT app to just crash because it's so processor heavy.

For what sort of that's not my experience. Yeah. I mean, you know, I I I haven't had a crash from sub agents.

I I haven't either. We both have big laptops. But I I know I know people brought it up.

There was a topic of discussion that we didn't see the same, but it is another vibe eval. Right? People are like, okay.

The amount of sub agents sold this morning is crazy. And I'm like, I think this is okay. I think it's good, but just stuff people bring up.

Speaker 2

as opinionated about, like, who is Ultra for? And, like, when when should they be using it? And since then, we made some changes to, like, you know, require you to turn it on and and and find it in the advanced setting because that's what it is for.

It's for, like, power users who understand what's gonna happen, because it also, you know, depending on your use case, can can use more of your your limits as well. Yes. So that's where I think a lot of the the feedback was coming from.

It's okay. Reset the limits. Always reset the limits.

Well, you know, today we're resetting because of this. I wanna change topics to one last piece of the harness memory.

Speaker 1

A lot of people are commenting on memory recently. Chai Gypsy's new memory system used to suck. It's not very good.

And then this guy also basically the same thing.

Speaker 2

And Samir, who you presumably work with Mhmm. Talking about memory. What can you say there?

I think that, you know, Samir and the team have made a ton of and and then the research teams have made a ton of updates and and improvements over time. I think when I talk to friends, family members about what they love about ChatGPT, like, the fact that it knows them, that they feel like their ChatGPT is is their ChatGPT, think, comes up probably Yeah. Number one.

And ChatGPT work in the in the cloud. Like, by default, all conversations, like, inherit from your ChatGPT memory, so you'll know they'll know context about you, and they'll also be able to write back to this memory. With it like a like a like a small text write, like you tell me when you're writing.

Right? Is it No. It's part of the same like memory v three system that that we we launched.

Yeah. Dreaming v Yeah. So I think that's been really powerful because, you know, going from Chattypedia to Chattypedia work feels like an extension of what I've already been doing with the product for sometimes many years.

So that's been awesome.

Speaker 1

the the improvements here. Is there so it's it's basically a retrieval problem. Right?

Like, are you retrieving the right things? Are you over focusing on the wrong things?

Speaker 2

more false positive or false negative, you know, if if that makes sense? Like, what's the bigger problem? So I don't work on memory directly, so it's hard to say what the bigger problem is with, certainty.

But I I think you're right. I think that, like, you know, that the there's two sides of it. It's like, you know, making sure it knows things about you, but then also having the EQ to, like, bring those things up at the right moments proactively or surprising you in ways that are positive, not negative.

Yeah. So I think it's a very challenging problem, but something that I I think we feel very there's a huge opportunity to get right, which is like why we made, like, big investments in it.

Speaker 3

managing memory across different projects, collaboration, and whatnot, how do you see the side of what's separate from the harness? Right? So if I have four threads on one project, any learnings on how to build memory systems there?

For background as well, I guess, steer it a bit is when you do chat style applications, I'd say you have a lot of one offs. Right? Yeah.

When you switch to work, it might be something you're doing for a month, something you do a lot. Right? Now, as I add more sessions, there's a lot more than just single threaded.

Right? And there there might be memory there.

Speaker 2

I mean, I think first I challenge that like the depth of the memory or the, like, value of it is, like, fundamentally different across chat and work. Like, it it is true that, like, you know, there are a lot of, like, shorter sessions on chat, but I think, you know, the ChatGPT, the product has had, like, a ton of longevity in you know, as long as this this technology has been around, and and people use it for worky, like, productivity related things already today. And so I think we found that there's a lot of value.

I mean, I found this with my personal usage, like, all these one offs add up over time into something like quite durable and like quite a good representation of who I am. I know like from time to time something will go viral on on Axie Bell, like, know, Chad Jubilee telling you everything it knows about you. And people are always surprised like how how deep that is.

The fun roast me, you know? Exactly. So like, I think like the that's all to say that like, I think there's a lot of depth there in the existing, you know, Chatty Beauty product.

And so that's why I think we think it's valuable to bring into the work product. But the other reason I brought that up is because I think, like, hopefully, we can use some of the same fundamental primitives and systems to extend memory here as well. And I know this is something that the the team that focuses on this is, like, working working through right now.

I wanted to bring up one element of memory, which I honestly don't really use much, and I'm curious if you do.

Speaker 1

which was is up on screen right now. It's kind of a super memory, or, like, what what is it?

Speaker 2

I think the idea is that, like, it can learn from, you know, how you're using your computer and, like, it's another input source into memory. And I think it's, you know, experimental right now and something that, like, isn't default off. But I'd recommend that you try it.

I think that it's, like, quite interesting how it it goes back to a conversation we were having earlier on, like, you know, you were asking, like, does it can ChatGPT miss things? Like, does it you know, on on Slack, when when it's searching, does it miss things? Cause there's such a volume of stuff.

Right? And, like, it's I'd you can ask the same question about, like, everything that you're doing on your computer. Like, is it gonna you know, everything that you're doing is gonna capture the intent and stuff like that?

Probably not. But, like, it probably will find things that you might not know about. And then if it can surface those to you in relevant times in proactive ways, like when you're doing tasks, and I I found, Elise, that it can be quite helpful.

So it's worth trying.

Speaker 1

longer term.

Speaker 2

Yeah. Exactly. Like, insights and it builds context that that makes that can make you more productive on certain tasks, but it's it's it's hard to describe without feeling it.

Speaker 3

I will say you can feel it pretty well. Like, the idea of what they're saying here. Right?

Just check through my memories or check through my logs and add skills. Yeah. Pretty underrated.

Right? That's automations.

Speaker 2

You you can repeat that using a cron job. Checking through your memories and creating skills? Yeah.

But I think the creation of the memories from Chronicle itself is like what's different.

Speaker 1

memories because you have Chronicle on. It's there. I don't use it much, but maybe I just I need more examples.

Speaker 2

I I imagine you guys use a lot of it internally, so I'm always fishing for use cases. Yeah. I would just try turning it on and then like It just auto works?

Like it Yeah. And seeing like where where it might start helping you, think you'd be surprised. Yeah.

Amazing.

Speaker 1

I think that was about it in terms of like the the overall coverage of ChatGPT work. I think there's been a lot of, like, good progress and discussion on building and all these things. There's a lot of, like, ex founders in the in the community and in OpenAI as well.

Do you think that things have changed a lot? I guess, like, your overall reflection of building pre AI and post AI?

Speaker 2

I mean, I think things have changed a ton. I think it's it's like super exciting to see how quickly you can go to from idea to something real today. Yeah.

Whereas, like, even before, like, I think, you know, five, ten years ago, like, it it was fast if you were scrappy and, you know, like, rolling to build the the minimal viable thing. But, like, now the extent of what you can build is, like, much, much broader. And I think that also like, what we've seen internally building is, like, that gives you an opportunity to validate much more quickly, to talk to users, to talk to to internal documenters, etcetera, and, like, make sure you're on the right track.

And, like, that loop, I think, has been has become more close than ever before, and that's like a win for product development. I think it's a win for for consumers and users too because ideally, that means they're getting much more better much better products out the gate. Does it mean your teams are smaller?

I think there's much more to do now. So I think people can accomplish more individually or in a small team than they were that would require more people than before. But there's at the same time, there's also more to do.

So I think the teams are much more ambitious. You seen any changes in scopes of roles and building teams and how we used to have teams, say, a few years ago versus what ideal teams look like now?

Speaker 1

designer, etcetera. Like Yeah. I wanna bring up this quote.

There will be only four jobs left in tech. There's AI slop cannon, the people who just, like, throw burn a bunch of tokens. And then there is there is SRE, who people who are more responsible.

There's grown ups who sell things, and then there's hot people.

Speaker 2

This is an interesting take. I think my my suspicion is that there's everything everyone will be like t shaped in a way, and that, like, AI will enable everyone to become a generalist. Like, you know, things that like, I I never would be able to, like, come up with a design before.

And, like, even now, I don't have it's, like, the visual taste required, but I can iterate on something with the help of AI. But then people will have a specialty, and that's, like, the the I guess, the straight line in the t or the the upward line in the t. And so, like, you can have a specialty that you're interested in.

With the help of AI, you can go deeper and become better at over time, but then you'll also be a generalist. And so with that foundation, the way you can accomplish is, like, almost limitless. What are you bottlenecked by in terms of specialties?

Like, do you need more designers? Do you need more slop cannons? Do you need more hot people?

I think the bottleneck some becomes like sort of like ideas and tastes, I guess. I think because anyone can can build now, I think it really is the era of, like, bottoms up ambition.

Speaker 3

you know, the amount of ideas and amount of things that you're doing at any given time. Do you think models help solve that? Models?

Yeah. I mean, I have the example of, like, I have a front end design skill that's like they give me four drastically different examples of what this looks like. Sure.

It burns a lot of tokens, but, you know and then I'll mostly just condense down. Okay. Like this part.

I like this part. Let's draw these together. And it's like yeah, I had a vision, but, like, I don't know.

I would say that the one automation that I would love to work and it doesn't work is bring me new ideas.

Speaker 1

Right?

Speaker 2

Somehow, LLMs are just not it. One interesting part of that idea is, like, they're not, like, in a vacuum. It's like not they they usually come from somewhere.

And, like, you know, in in product development, like, they're coming from talking to users or reacting to, you know, friction that you're seeing or feedback, building on some foundation that you already have planned out before, whatever. And so I think that's where, like I think there there will always be value in in these, like, generalists that we talked about, like, you know, closing that loop and and and having coming up with those ideas that are grounded in in that feedback or talking to users or whatever it is. Cool.

You work on a you lead the productivity team.

Speaker 1

How do you define productivity?

Speaker 2

I think our mission is to make it possible for people to do things that they weren't able to do before. And right now, we're thinking about it from the the perspective of knowledge work. And so when I look at knowledge work, I think about people are no longer siloed by their roles.

They're no longer siloed by maybe the the background or training that they have. Like, no matter what function you're in, you can suddenly build things. You can suddenly get access to data that you otherwise might not be able to interpret, etcetera.

And then I think that extends to your personal life where we want to give you leverage at the end of the day. Like, we want the models and the product to be able to give you leverage so that you can, you know, create time for yourself to do the things that you love. Does that also translate to a way to measure productivity?

Speaker 1

Like, what is Neil, how do measure leverage?

Speaker 2

I think we haven't figured this out yet. Part of the reason is it's so diverse. Everyone has different goals.

And really, the true measurement is like their ability to achieve that goal. Did we help you or did we not? Yeah.

And it's very difficult without knowing what that goal is up front and also tailoring your current individual. And the thumbs up and thumbs down from ChatGeePee doesn't give you anything. Right.

I mean, you don't know if they're thumbs down in the the content of the answer, the vibe of it, to help them with their goal. Yeah. I think that's difficult.

But it's something that I think we we will need to figure out, and the industry at large will need to figure out because that's how we measure success if this is what we're for.

Speaker 3

and how you measure it? Basically, said there's a lot more work that can be done, a lot more scope.

Speaker 2

Has it changed? I think it was always true that what you really wanted to measure is like, was your team, was the individual, were you personally able to hit the goal, or are you closer to hitting whatever your goal is. Right?

But I think previously we used proxies for this. So like, you know, code commands or code lines or whatever.

Speaker 1

Story points. Yeah. Exactly.

Story points. And like They're coming back, by the way.

Speaker 2

May maybe. I mean, but but that is for a part of the change. And like, I think with with AI now, those proxies starting to fall apart.

Like, you you know, and the number of tokens you use or the number of pull requests you make or, like, no longer, like, maybe as hyper correlated with them that this virtual team able to hit the goal, or are they on track to hit their goals? So I think we'll need to come up with with new measurements.

Speaker 1

For the managers listening, give them one thing to try.

Speaker 2

I think for me, what's important is, like, at bats. Are we as a team building the muscle to have not just quantity of at bats, but quality? Like, are we able to go all the way from, like, generating an idea, building it out, getting the feedback, reacting to that feedback, actually validating or invalidating the hypothesis, going on to the next idea?

Are we able to do that really efficiently? Like, that goes to, like, you know, the actual, like, code that's being written or the designs that are being made or the specs that are being written, whatever, but also the culture of the team. Like, do we have the humility and and then and are able to, like, go through that process many, many times and stay motivated and excited throughout that?

So that's the thing that, like, I think is important now, especially when we're on the frontier of this technology and, like, there's so much to build, there's so much to do. That's probably the most most important thing that we look at.

Speaker 3

around measuring productivity? What should teamwork on? I feel like there's a lot of, okay.

We added a lot of LMs. We have dashboards for this and that, but not much has changed. Right?

Speaker 1

is the trap. Yes.

Speaker 3

And, you know, the the broader source of the question is for for the managers and teams building, you know, how how should they approach this? I think maybe the trap is, like, conflating motion and progress.

Speaker 2

think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be, like, very prescriptive and deliberate about, like, what you're actually trying to achieve. And it goes back to our question of measurement.

Right? Like, you were when we were talking about, like, can we, OpenAI, like, figure out how to measure productivity for our for our users? That's that's a very hard problem because of the diversity.

Speaker 1

looks like for you and for your team. And if you don't have that, then it's very easy to conflate these two things. I think at bats is a really great thing.

I'm I'm I'm really glad. I I like the discussion between motion and progress. I think that's a quote that we're gonna feature on the write up.

You've been very generous with your time. Thank you so much, and congrats on 10,000,000. Yeah.

Thank you for having me. Next one at 100 in two months. For sure, two Two weeks.

Thank you.

Shared via Hopper