Andrew Lee, CEO of Tasklet, details his company's complete rewrite of their agent stack, now emphasizing file system context, agentic search, and multi-resolution summarization for token efficiency. He discusses the strategic challenge of competing with Anthropic, their primary model supplier, due to token pricing disparities, which is driving Tasklet to become a horizontal platform supporting multiple frontier models. Lee also outlines his framework for the three types of software companies that will survive the AI transition: horizontal platforms, API-first companies, and solutions providers.
Hello, and welcome back to the Cognitive Revolution. Today, I'm pleased to welcome audience favorite Andrew Lee, CEO of Tasklet, back for his fourth appearance on the podcast. Andrew has always been extremely transparent and candid.
His belief that speed is the only moat has made him comfortable sharing intimate details of Tasklet's agent architecture. And as you'll hear in the six months since we last spoke, Tasklet has indeed once again entirely rewritten their stack. Today, there's much more use of file system context and agentic search to leverage available information while conserving tokens.
Plus a huge new emphasis on summarization at several levels of resolution. This time around, we also dig in to the delicate strategic situation that Andrew and Tasklet face. While their product strategy of always betting on the models has proven correct and Andrew's choice of Claude has been rewarded, Andrew observes that these days, everyone is fundamentally building the same thing.
And today, his most intense competition is actually coming from his critical supplier, Anthropic, which with Claude Max accounts gives their direct customers an estimated five times as many tokens as Tasklet can purchase at the same price via the API. In micro terms, this relatively high cost of tokens has caused Tasklet to stick with OPUS 4.6 rather than moving to the new 4.
7. And in macro terms, it's pushing Andrew and team to become a horizontal platform capable of harnessing, or as Andrew describes it, outfitting with a mecha suit, frontier models from any provider. This evolution, which I do think Andrew has played and timed about as well as anyone possibly could, is critical because horizontal platforms are one of only three types of software company that Andrew believes will survive the AI transition.
The others being API first companies like Stripe and companies that develop solutions and sell outcomes. Best exemplified perhaps by Finn's model of 99¢ per customer service ticket resolved. We get into lots more besides, including Tasklet's new instant apps feature, how they're thinking about deep personal and shared organizational context, the cloud container company that Andrew endorses, Tasklet's token to labor cost ratio, and whether or not Zuckerberg has come calling after his Manus acquisition was canceled by the Chinese government.
This is a fun one with lots of valuable detail from somebody who's in the arena competing to become one of the few general purpose AI agent platform winners and still actually willing to tell us all about it. Please enjoy my conversation with Andrew Lee, founder and CEO of Tasklet.
Andrew Lee, returning champion and CEO of Tasklet. Welcome back to the Cognitive Revolution.
Thank you. Glad to be here.
You are a fan favorite, and, I'm gonna just try to pepper you with a bunch of questions and make sure we get as much alpha for all the builders in the audience, myself included, as we can. So first question, it's been about six months since we last spoke. You have rung in my head probably weekly with your speed is the only moat mantra.
And I guess my first question is, what have you rebuilt in the last six months since we talked? Or maybe more to the point, what have you not rebuilt in the last six months since we talked?
Yeah. I I think the mantra has stayed the same, and we've rebuilt basically everything. I was I I I was, thinking about this earlier and even the pieces where I'm like, Oh, this has stayed the same.
Actually, I've been totally rebuilt. So from a product perspective, the product is actually very different now. When we launched this thing in October, it was all focused on a workflow automation.
We thought, Hey, it would be really cool if you know, people could come in, describe a workflow, we'd like run the workflow for you. But basically the feedback that we got right out of the gate was, hey, it's know, once I've given this agent all of my context and hooked it up to all my stuff, I don't want it just running my workflows. I also just wanna be able to talk to it synchronously too.
And so it's not just a workflow automation tool anymore. Now it is a like a very general purpose agent. It's great for doing workflows, but it's also great for doing other types of stuff.
So that required basically a total rebuild of the product experience and as a result, lot of technology behind it. So as an example, in a workflow automation tool in the previous iteration of the product, you basically had a main agent that you'd talk to for a brief period to kind of set up your workflow. And then once you're done setting up your workflow, basically stop talking to that agent.
So the chats were relatively short. And then our system would kick off runs of what we call the task agent on a periodic basis when events happened. And so every agent was like a pretty short thing.
And you could do like pretty simple context engineering to make that work. In a world where you want this like general purpose agent that you can talk to synchronously and run these automations, the product experience people want is just one big linear chat where like everything is in one chat. It works really well from our product experience, but the engineering gets really complicated because you can't just have an infinitely long chat history that you feed into the LLM.
Even if you could, it'd be really, really expensive and you wouldn't want every automation to have to send in 10,000,000 tokens from all the previous rounds. We had to totally rethink the way context engineering works and say, What if instead of the history being the thing we send to the LLM, what if the history is in the file system? What if the files are the agent?
Marc Andreessen had a thing about this, and I think people figured this out, but we made the switch in November where we said, really, okay, what we need is a file system that has your history. And then we need the thing that we actually send to the LLM to be just like hints at what's in the file system and what things you need to read to get the work done. That way we can scale the agent up, including the number of chat messages that were sent.
There's a lot of samples we'll scale for the future. But basically you can scale from what fits in the context window to what fits the file system. You can fill up a lot more stuff in the file system.
There's a bunch of other stuff we've rebuilt. Another big thing we've rebuilt is around computer use. When we launched initially, computer use was this add on.
You could have a Linux machine. Actually, initially it was a Windows machine, then it was a Linux machine tacked on. It really was an afterthought.
You could use it for certain things, it worked okay, but most things you do with agent didn't touch it. At this point, computer use is the absolute core of the product. Basically everything you do is like running shell commands, touching a file system, touching a database.
We have a very tightly integrated browser use experience now where every agent has a headless VM and a browser VM that persist state across runs and allow you to do lots of really cool stuff. It's sort of in the critical path. So now like if our computer use goes down, like everything goes down versus before it was kind of this afterthought.
We've also rethought the way our integrations work. So this is something I think from a product experience, it maybe doesn't look terribly different to folks, but the basic architecture of how we plug in other systems to the Tasklet agents has been completely rethought, basically to allow the agent to have more control and management of those connections. A simple example of a product experience that's improved is you used to not be able to connect multiple instances of the same type of thing.
You couldn't have three different Gmail accounts connected to an agent, and now you can. The base architecture there has been rebuilt. So I I'd say basically every line of code, there's probably been touched in the last six months and like most of our fundamental assumptions were thrown out.
The product still can do many of the same things, but hopefully much better now.
Yeah. That's cool. So literally, you can't think of anything that has survived the last six months?
Nothing substantial. The visual design is totally different. The structure of the app is totally different.
Yeah. The connections system is different. The way we use computers is different.
The core of the agent is different. The context management is different. The way we do compaction is different.
So it's all new.
Yeah. Okay. Let's talk about that compaction.
Because I mean, one of the big takeaways from, it might have been two conversations ago was the critical importance of caching. And you had said at the time, you know, long context is, you know, it's quite effective. Obviously, it's expensive, caching and especially the sort of, 90% off cashing pricing of Claude was like critical to enabling certain things to really work in a way that wasn't, you know, tanking your you know, what is it?
The the classic meme. Right? My my company is dying with making it work without killing your company.
So it sounds like that's changed quite a bit. Now it's much more of like a pointers, hints type of thing. How how is this working?
And obviously, tokens are like you know, people are spending a lot on tokens these days, and I'm, you know, increasingly, like, hitting my limits on even the the highest, you know, plan with Tasklet these days. So how are you managing context for me? What should I know?
What lessons should I take away from your experience on how best to manage context in the modern moment?
Yeah. So caching has actually become a much bigger deal because now that the real context is in the file system, there's just a lot more tool calls that needs to be done to do the basic operations of the agent because you're loading in a bunch of files and stuff. And so we really have to make that caching work if we don't want this thing to be like crazy expensive.
So that's been very much at the forefront. We came up with a new approach to context management that we shipped in December that basically works like this. You take your whole chat history and you put it in the file system.
So it's all accessible in the file system. And then you find a way to summarize that whole history in kind of a fixed length number of tokens by having recent stuff be included in like sort of high granularity. Like the last thing you said, you'll probably have most of it or all of it there.
And then older things basically have like decreasing fidelity as you go back. So if you have a very long chat, the stuff in your current turn, like the current thing that's running, it's probably all gonna be there, including all the thinking blocks and all the tool call responses and all the files and things are probably gonna be sent to the LLM per bit, depending on how long the run is. But for most kind of short runs, that'll be the case.
The previous turn will probably mostly be there. You're gonna have the full user message. You'll probably have the assistant response.
You'll probably have the tool call arguments. You'll probably have the tool call responses. You'll probably have the thinking blocks.
But as you go farther back, we start stripping the thinking blocks, we start stripping the tool call responses or at least truncating the tool call responses and then stripping them. We start truncating and then stripping the tool call arguments. Then we start collapsing tool calls, and then we start shrinking down the assistant messages.
And then finally we get to some LLM based summarization. And we do this in buckets moving back so that we can have sort of a minimal impact on caching. Basically, want to avoid messing with prefixes as much as you can.
As you go back, basically you get into these buckets where we have different levels of compression. And those buckets, as they get older, tend to get added to very slowly. Then once they hit a certain threshold, we shrink them down.
And this system has actually worked basically pretty well. And the core thesis basically is you generally care a lot more about recent stuff and you trust the agent to go and look things up when it needs to. And I would say it's not perfect.
We definitely do have people say that agents forget things. It definitely does still cost us a lot of money to run, but I think it's generally worked. And our plan is to double down on this type of architecture.
We have lots of ideas how to improve this, but the basic approach of these sort of like decreasing fidelity as you go back and these like bucketed cache aware chunks, I think is the right approach.
So how often does that get updated? Like if I have a agent that runs on a daily basis, do you try to keep the cache active from one day to the next, or is it sort of every day we're going to have a fresh cache that will run through that whole session and all the interactions, but kind of tomorrow, do we begin again or do we like, well, on what frequency do we begin again?
Well, there's two pieces there. There is when do things get updated on our side? Like when do we decide what that compressed history that we put into LLM looks like?
And then what caching do we do on the LLM side? And the answer to the former is every time you do anything, it's sort of incrementally updated, including in the middle of runs. If you have like a very long turn that uses a lot of tokens, it might actually start compressing inside that turn.
And the reason that we persist that is actually calculating that could be really extensive. Like running an LLM based compaction of an older section eats a lot of tokens and you don't want have to do that every time the thing starts up. Every hour you have a trigger running and every time you have to compress a bunch of history, that can be very expensive.
We keep all that around. On the model side, caching depends on the provider. In the case of Anthropic, we're using five minute caching and so it doesn't stick around very long.
And the assumption there basically is you're probably either in an active session or like in the middle of a turn, in which case five minute caching is enough, or you're probably waiting for next trigger to run. And like most people's triggers are not running like every, half hour, they're running every few hours or every day. So the assumption there is this is not so common.
And then different providers have different possibilities there. And for example, OpenAI has much nicer caching primitives, for example. I'm happy happy to talk about those too.
Yeah. Okay. That's interesting.
So it's basically constant maintenance of the higher level summaries that will be fed into the LLM and then pretty short kind of single burst style caching to actually reduce the cost of incremental calls within, like, one agent run.
is hit for the for one run, but not hit across runs for the most part. Yep. That's that's the current approach.
And one thing I wanted to note is the way our system is built today, we basically get no cash benefit across users. So it's like caching for actually, even even per agent. Like, it's basically caching per agent.
There's some changes that we can make to do a lot of cache optimization across agents and potentially even across organizations and users. I don't want get into the specifics of that because that's still upcoming, but I think there's lot of potential actually to save money across agents as well.
Yeah. Okay. Well, that'll be important.
Hey. We'll continue our interview in a moment after a word from our sponsors.
The Cognitive Revolution is brought to you by Brave. If you want to stop hallucinations, empower your AI agents to do their own research with the Brave search API. Brave offers the only search API with its own index at scale.
It's lightning fast, excels in rag pipelines, and it's a leading search option for Clawed MCP and OpenClaw. I've built Brave Search into my personal AI infrastructure as a core tool that all agents can use anytime they need it. To find guest headshots and company logos for the podcast, they use Brave's image search.
To build small business profiles for use in my Waymark prototyping work, they use Brave's place search. Across all use cases, my agents tap into Brave's index of 40,000,000,000 high quality pages tens of times per day. It's the only global scale index outside of big tech, which means no Google scraping and no SEO spam.
Plus, with true zero data retention policies, you can meet compliance obligations and rest easy. Pricing starts at just $5 per thousand API calls, and you only pay for what you use. Sign up now and get $5 in free credits to start and empower your agents to start calling the Brave search API today.
Most billing platforms were built to send invoices and assume your pricing is simple and predictable. But if you're building an AI product, a fintech tool, or a developer platform in 2026, your pricing is anything but. Usage tiers, consumption billing, and bespoke enterprise contracts are now the norm, and you're probably managing it all across disconnected tools and fragmented systems.
Sequence handles the entire revenue workflow from contract to cash. Quoting, invoicing, metering, revenue recognition, plus Sequence agents that automate the manual finance work that usually takes teams days each month while also helping them to collect cash faster. Companies like Cognition, Incident IO, Runway, and OpenRouter use Sequence to run their full revenue process between CRM and ERP without the spreadsheet mess.
If your pricing has gotten more complicated than your current billing setup can handle, check out sequencehq.com, and use the code Cognism in the source field when you book a public demo to save 20% off year one.
Let's, I do wanna circle back to the OpenAI question because you guys have been clawed maximalists, and your other mantra that rings in my head a lot is always bet on the models. I'd say it's safe to say that that bet has gone well over the last six months. Obviously, we've seen some of the most notable model releases in the sense that the community has sort of flagged, like, four, five, and four, six as kind of qualitative shifts where, like, things went from not working to working and people are like, oh, I can really get pretty general purpose knowledge work out of these things on a pretty consistent basis now.
I would love to hear how you would characterize the advances that we've seen. Maybe you could do that in terms of like what new use cases have opened up, maybe, you know, things that have surprised you, possibly also like things that are still not working that you, you know, that might be surprising given all the things that do work. And then then we can get to the latest models.
I wanna so kinda give me the, like, four or five, four or six history, and then we can go to four seven and five five present.
Yeah. I think the overall approach of always better than the models, think that is I totally agree. That is totally held up.
When we started working on Tasket, we were using four, Cloud four, and that actually could get you a long ways and it worked pretty well. 4.5 was a big unlock.
So 4.5 started to work much better to do doing computer use and it could just manage navigating through the various different connections and tools enablement process and stuff much better. And that came out very early in the lifetime of TaskNet.
Was like, I think a big bonus for us. And the cost reductions around Opus were huge for us in, I want to say, December when that happened. They dropped.
So initially, we could really only have people on Sonnet, then Opus dropped from 50 to five, and that was a big unlock for us as well. I think 4.6 was a solid incremental change.
It made it nicer for, again, for computer use, which has become increasingly important for us, both like headless and head full, as well as CodeGen. And that enabled our instant apps feature, which I think we'll probably talk about today, is a very cool feature. 4.
7, we actually haven't rolled out yet. I think 4.7 has been much better in certain areas.
It's much better like code and one shotting kind of long projects. But for the types of iterative knowledge work that we support, it doesn't actually seem like a huge boost and it's a lot more expensive. So the tokenizer changes they made increased our costs for like 30% and costs are huge for us.
Because we essentially pass them on to users and so we actually opted not to ship 4.7 as a default recommended model. We are going to ship it, but as an advanced user option if people want to, and we're to note that, hey, this actually costs a lot more.
I think the progress there has been great. Reason we started on Anthropic and have been so Anthropic focused for so long was just that the basic core of our agent, like the ability for it to like navigate through a discovery process of connections and activate the right tool in your agents and then manage its context the way we manage its context. Just the core inner workings requires kind of a base level of intelligence that the other models just couldn't do.
You basically couldn't use the same harness with them and have it expected to work well. That has changed, which is really exciting for us. It's kind of scary to be like, we're totally dependent on this one vendor, even though they're great.
Like, drop down models are amazing. Don't get me wrong.
Supply chain risks are everywhere if you take that approach.
Exactly. But more recently, GPT 5.5 has gotten really good.
I think it's a huge step up for our use case over 5.4. Will, by the time this podcast comes out, we'll probably have announced this publicly.
Just, hey, it's gonna be there very soon. And it can navigate our harness super well. I think it gives OPUS 4.
6 a run for its money for most use cases. So I think that's really exciting. I'm pretty optimistic about the OpenAI roadmap this year.
I think they made a huge bet on compute last year. And I think that's starting to show, and I think where they're gonna have a lead for a while. And it also is clear that they've refocused their business much more on these types of use cases.
And you saw like the progress Codex made over like six months. And if they put that level of effort into doing the models around these types of genetic use cases, that'll be huge. We signed a deal with OpenAI the other day and we'll be launching stuff there and we're making a pretty big bet there as well.
There's also been There's progress in other places though. The latest Google models are are pretty solid. They're not they're not at the level of of Anthropy or OpenAI yet, in my opinion, but they are making very fast progress and they're much closer.
And then we've been playing with with DeepSeek and with Kimi. The latest Kimi is, as far as we can tell, maybe better than a Haiku and cheaper. So I think we're going to see those probably make their way into TaskClip soon.
So I would expect within the next few months, we are going to have Anthropic Models, OpenAI Models, Open Source Models, Google Models. I'll bet you the Anthropic ones will still be the best and probably the recommended in most cases, but people have a a variety of choices and, like, some good cost cost optimization options for certain things.
So many just follow ups there. Let's maybe start with your what I imagine has been a little bit of a delicate dance with Anthropic, and then we can kind of compare and contrast that with what OpenAI is now bringing to the table. I don't know.
You you probably know what the ratio is of API cost to effective token cost when you buy a Cloudmax subscription and max it out. And obviously, in the, you know, intervening time since we spoke last, there's been the whole open call phenomenon and that's had its own, you know, bunch of drama with you can, you can't, you sort of can, we gotta pay the API price. We're lowering our limits.
We're, you know, buying compute from x AI. We're raising our limits back again. What has it been like from your perspective to be building on a platform that you're also sort of competing with that is kind of undercutting you to various levels at different times on their pricing?
Yeah. It it it's definitely a it's definitely an interesting relationship. So, like, on the one hand, the models are amazing.
Right? They're super good for our use case. And, like, their their team has been, like, really helpful and responsive.
And, like, we talk to them on a very regular basis and they're trying hard to support us and we get early access to stuff and they take our feedback and all that. They're definitely totally enabling our business, making it happen and working hard to do it, which is wonderful. They're a great partner.
I don't want anyone to think otherwise. But at the same time, if you look at our stats of when someone churns off Taskit, where do they go? 80% of those users go to an anthropic product.
They are a very direct competitor. I think there's different use cases where we shine or their products shine, but it's clear that they're a very direct competitor. And the number one reason that they do that is because they already have a Max Plan and they don't want to have to spend additional money on Tasselit.
And so basically, every time they release a new model update, we're like, this is great. This is awesome. And every time they send an email being like, now your max plan is even cooler.
We're like, crap, right? This is just going make it harder. And they definitely subsidized it.
And it definitely has set some distorted expectations with folks around like what you can get for a certain price. And so costs is like a constant struggle for us to try to help users use the products more efficiently and help them understand, hey, we're actually working at like some pretty razor thin margins here and trying to make this cheap for you guys. So yeah, it's an interesting dance for sure.
Do you know what the ratio is or is it opaque even to you?
I I don't know what the ratio is. No. Yeah.
Interesting.
It feels like it's not insignificant. Like, I my intuitive gut guess would be it's like five to one or something.
guess too. Like five to one or maybe maybe even more. It does seem like pretty substantial.
Yeah. Yeah. That's a big that's a big deal.
So okay. One more thing on the anthropic dance, and then, obviously, this gives you, you know, a lot of incentive to broaden out and try to position yourself a little bit differently. How do you think about obviously, they have an inside lane when it comes to building product experiences that make the most of their model's capabilities.
I mean, they're know, increasingly, we're sort of seeing, like, the model is being trained in the first party harness. And I have another question on the on the word harness and, you know, if that's even the right paradigm to be thinking about this anymore. But how do you kind of position yourself to you said, like, there's something some use cases where you feel like TaskClot exceeds what you get, you know, out of the first party cloud products.
Like, I guess from one thing, like, is that even possible? How do you think about trying to compete with what they themselves are going to build given all the inside knowledge and advanced prep and kind of close coupling advantages that they have?
I kind of think that everyone is building the same thing. Like you have all of these different agent companies and basically over time, as the models get smarter and the agents built in more sort of general purpose tools, computer use and file systems and whatnot, you can do very similar things in many things. So you can go into Cloud Code and you can do all kinds of non coding things in Cloud Code.
Codecs and Cloud Code and many other startup products are all able to do coding and non coding things pretty well. I think where you start to differentiate, and I do think you can differentiate within this space to some extent, is really around what you're optimizing for and what the ergonomics are. So in our case, take Tasklet.
You can totally write code with Tasklet. You can hook it up to your GitHub and you can have it generate PRs and it does it just fine. We knew this for marketing, for example.
If we put on a new blog post or whatever, I write the content in Tasklet and then I just have it generate its own PRs and it works fine. But it's not going to be as smart and definitely not as cost effective for heavy duty coding as going and using an actual coding harness. It's definitely not going to be as nice to use because the actual coding harness is going to be in a conductor or something that's designed for a coding workflow, and our product is not set up that way.
So I see a future where you can kind of pick up any AI agent to kind of do anything. But different ones are gonna have sort of different cost and performance trade offs, and different ones are gonna have just different ergonomics for the different types of work. Where we really shine is 20 fourseven automation of knowledge work for companies, especially knowledge work for companies that is not like your personal work, but like something that a company owns.
So if you have, simple example, you have some complex invoicing process at your company. You don't want to be running that in your local co work, right? If you close your laptop and the company can't invoice people anywhere, that's bad.
And you don't want to put that in OpenClaw and put it on your Mac mini in the corner. Because again, if something trips over the power cord, well, you can't run your invoicing. What you really want is something that's running in the cloud and is manageable by many people, and you have a lot of infrastructure around it to manage and provide oversight and have audit logs and you have guardrails around the thing and you can control COF and your different agents.
And so there's a lot of kind of team enablement features you care a lot about. And that's where we really shine. And a lot of the work to make that work well is actually kind of fundamental to the way the agent is built.
So I talked about our context manager. The reason our context manager is built the way it is, is because you want to have triggers as sort of regular messages into the agent, which means these agents, if you have an agent that's running a trigger every time you get an email, that agent might fire 10,000 times this year. And so you need an agent that can fire 10,000 times and still remember things at the beginning of the chat and still behave in a reasonable way.
That's a pretty different thing to optimize for than a coding session. The way Claude codes or resets context makes total sense in a coding environment. It doesn't really make sense in a world where it's processing all your emails.
So that's how I see us differentiating. A couple of notes I want to make here on differentiation is one, the market is just freaking huge. So if you look at coding agents and you might say like, oh, clearly Cloud Code and Codex have won.
But like, Cursor is going to sell for like $60,000,000,000 And even, you know, even the fourth and fifth and sixth, you know, like, Cognition's doing just fine, Factory's doing just fine. Even Windsurf that had to sell, it was like pretty good exit. So if we end up, I'd love to be the number one here, but if we end up being like number four or five or six, they could still be a very significant exit.
The last thing I want to note is, and I think this is probably the most important point, when we go and pitch a business, what we're trying to help them do is deploy AI for real inside their company to automate stuff. Typical companies, they don't want to have to spend all their time researching AI models and placing bets on which lab is going to win. They want to choose a platform that's going to serve them well and they want to benefit from everybody's improvements over time.
We can go in there and say, Hey, a bet on us is not a bet on Anthropic or OpenAI or anyone else. A bet on us and a bet on everybody. We're going to give you a private model, OpenAI models, Google, and all the open source models.
Then we will be a neutral arbiter of what you use. So if, to the extent that we can build features to help you choose the right model for the job, optimize your costs across the different things, you can trust us because none of these are ours. We're getting the same margins on everything.
We're a neutral party versus if you go back to Anthropic, right now it's just Anthropic products. Even if they decided provide other models through their products, which they could, although I don't think they're going to, but I think they could, I don't know if you'd really trust them to do that in a neutral way. I think that's a pretty compelling part of our sales pitch.
Yeah. I think you've maybe navigated this about as well as anyone could in the sense that betting on Anthropic and kind of going all in on whatever the best model is, which has been clawed to make it work as well as possible while the capabilities curve was getting to critical thresholds and then kind of pivoting to being a more neutral abstraction above the model layer now that there are multiple options that seem like they're, you know, able to deliver the kind of performance that people want. It wasn't obvious to don't know if it was obvious to you that that's how it was always gonna play out, but I wouldn't say it was obvious to me.
I think I would have said about, you know, your kind of position six months ago, like, yikes. It is pretty tenuous to be all in on Claude, but I think you kind of timed it pretty well on a couple different levels. So how much would you say that's foresight and genius and and how much is is good luck?
Yeah. I I I I'm glad you think so. This was this was very much the plan, and, I do think it's worked out really well.
So No. I'm happy. I'm happy with how it's turned out.
Hey. We'll continue our interview in a moment after a word from our sponsors.
Visual AI is the ability for your software to not just store pixels, but to actually understand what it's looking One of our partners, Roboflow, is the company making this happen. They've built an end to end platform that makes it incredibly easy to go from a raw idea to a fully deployed application in just a few hours. For example, just look at Blueprint Pro.
They built an app to solve a major construction industry headache. They're using AI to instantly understand a floor plan. This was literally impossible just twenty four months ago.
But now that Visual Artificial Intelligence is accessible, thanks to Roboflow, there are tons of new companies being built. Go to roboflow.com to read the full Blueprint Pro story and see how over a million engineers are building the next wave of visual AI.
That's roboflow.com. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code.
Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with.
For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down ten ninety nines for all 10 of my part time jobs, and built me a comprehensive report on my expenses and donations. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders.
And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind.
But then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you.
Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So for problems worth solving, get started with Claude at claude.ai/tcr.
That's claude.ai/tcr. And check out claude Pro, which includes all of the features mentioned in today's episode.
Once more, that's claude.ai/tcr.
So, okay, let's do this harness thing for a second. The word harness itself just always makes me think of trying to control and direct some sort of wild animal to get useful work out of it when it might rather be doing something else. I just took a long road trip with my kids in a Tesla FSD enabled car over the last ten days.
We went to a lot of historical sites. The juxtaposition of horses and my FSD was pretty funny. But I kind of think of like, you know, trying to get this, you know, unruly animal that is a model to stay on track, right, and do what you want it to do.
These days, as you said, like, there's a lot more ability to give the model hints and say, like, here's a file system, and you can kinda go get what you need here. And I'm starting to feel like the concept of a harness is maybe a little anachronistic already, and and maybe what we're doing now is more saying like, here's the world you get to play in. It's not so much about trying to narrow what the model can do, but more about broadening what it can do.
How do you think about that sort of like narrowing and focusing versus kind of broadening, giving access, you know, unlocking new possibilities, which, you know, might in some cases even surprise users given the model capabilities we have now?
I I hadn't actually thought of a harness as being, like, a constraining thing, but, like, yeah, you kinda make a good a good point of of that would be the normal way you'd think about that word. I kinda think of it as, like, a mecha suit. Right?
Like, I I agree with your your thesis is, the the goal is to, like, let that agent or let that LLM actually do things in the world. To do that, it's gonna need storage, it needs compute, it needs to be able to reach out and connect to APIs, it needs to be able to talk to the user. And there's a lot there.
I think when I talk to people who are not deep into the harness world, I think most people assume when they play with an LLM product that it's a very sort of raw thing on top of the model and like you type a thing and they send it and what they see on the screen is just being sent to the model and the model's doing everything. And that is becoming increasingly less true. And the complexity of the code that's sort of like translating what you see to the LLM calls is getting more and more.
And I think that's gonna keep going. I think the sophistication of these harnesses is gonna get just 10 times as complex. But I think there's gonna be some pretty major breakthroughs here that increase the capabilities of these things like pretty substantially in the way that we handle memories and the way that we handle oversight and control and the way we connect other tools.
So I'm very bullish on the opportunity here, and I think these things are just going get more and more complicated. And, yeah, I don't know. Maybe we need to a name.
Maybe it's maybe yeah. Maybe it's like a mecha suit, and it's not like a like a harness. Yeah.
How much do you think so this is, I think, one of the more interesting debates right now in the AI builder community broadly. What matters more, model or harness? And I think you see pretty extreme positions on both ends where I I see I get emails that are like, models don't matter anymore.
It's all about the harness and vice versa. And obviously, either of those, like, extreme positions is is not gonna be right. But I guess I have historically come down somewhat informed.
I don't know if you've seen this graph from The UK AI safety AI security, I should say, institute, where they do a capability plot over time with the minimalist harness, you know, whatever kind of basic vanilla thing, and then the best available harness. And, of course, you know, both are going up. A year ago, though, the time delta between what level of capability you could get with the best available harness versus the vanilla harness was longer, and now it's gotten shorter.
Somewhat of that some of that is maybe just due to more frequent model releases, which is like shortening every, you know, window of advantage. Some of it is maybe because the models are getting more deeply trained to use harnesses and so, you know, they're they're just good at it out of the box. You don't have to compensate for their weaknesses so much.
But I guess my my overall summary would be it seems like I would say models seem to matter more and you can't get that you can't like, how how much can I live in the future with the best available harness for any given model? It seems like it's not a huge amount, but it sounds like you maybe see that differently. So, like, what's the case that that's if you do, what's the case that that's wrong?
as models get better, they can replace good harnesses. A model today with a crappy harness is gonna be better than a model from a year ago with a really good harness. Agree with that.
I think that trend is gonna continue. But I also think the effects are multiplicative, right? And they're orthogonal disciplines.
There's no reason not to take the best model and put in the best harness, and I think we should. You might argue that, Oh, given the exponential, that actually only buys us six months or something. Okay, fine, but it's six months.
But I think more importantly, what you're getting the only metric that matters is not intelligence, right? In these real production systems, intelligence is one piece, but take Taskit for example, much of what we do is automating specific workflows. Once the model plus harness is smart enough to, I don't know, do, order less lunch every day, which it does, and we've been able to do that for like six months, we're not going in there and messing with it very much.
Like incremental improvements to intelligence don't really matter, but performance and cost do. And so if you look at the harness and say, hey, the only point is to make the thing smarter, fine, it buys you a fixed amount of time over the model exponential, which is maybe cool, but not amazing. But it might make a significant difference in cost and other attributes, cost and reliability and the ability to like do oversight and speed.
And I think those things matter a ton for a commercial product. So like in our case with our harness, like the benefits you get are, you have a nice UI and the sidebar pops out at the right time to show you things and you get nice indications of working states. You can see what it's doing at the time and you get the ability to like have things persisted across long periods of time and you get nice performance trade offs and cost trade offs.
And, I think those things should not be underestimated for like a commercial product.
Yeah. If you can make it work work with Haiku instead of Opus, for example, that, moves the needle quite a bit. For sure.
Yeah. Especially in a compute scarce world, which we increasingly, seem to be in.
Okay. So If I can if I can interject here, like, as a good example of and I I don't even know if you've called this a a harness, but, you see what what Entropic is doing with their I think they call it, a supervisor agent or forget forget exactly the language they use. So basically, they have a system where you can inject a tool that allows a smaller model to call up into a bigger model.
And this is a relatively new thing that they've been talking about. And you basically can get close to the bigger model's performance, but do the vast, vast majority you work on a smaller model. And that's a huge win.
And, like, if you have those capabilities, why not? Yeah.
Yeah. That makes sense. So when you think about the best available harness and what that looks like, especially as you go to a multi provider paradigm.
How much do you think you're going to be building a harness per model versus trying to keep everything the same across models? Traditionally, one would think, like, no way we can build this, you know, complicated product across, you know, in a bespoke way for all these different models. We've got to keep it consistent.
But obviously, the old rules don't apply anymore. So what's your strategy? Like, how much do you tailor the harness to each new model that you wanna launch?
Yeah. This is actually this is very much on my mind right now. I think ideally as little as possible because we want to support a lot of models and it's to maintain across both balls, but we also want to have the ability for these agents like switch between models.
And so if you're like, Hey, you can run an Opus this way and it persists this state, but then you switch over the same agent to some other model and suddenly you to find a way to translate things, it gets really complicated. So we'd like to keep them as similar as possible. I think so far we've been able to, and our approach has been like, maybe we'll make some prompting tweaks that'll try to address issues in one model while not breaking it in the other model.
And I think so far that's mostly worked. I think over time, the APIs of these things have converged and the basic capabilities of these things could have converged. My hope is that we'll get easier over time, not harder.
But I could see us having some very model specific harness things potentially, and then thinking about ways to do that in like a really modular way, so it's like not a huge amount of overhead. But, yeah, definitely something on my mind.
Even beyond model capabilities, you alluded earlier to caching primitives being different across providers. So presumably on that level at a minimum, you kind of have no choice but to if not, I mean, maybe could have the same context, but you're gonna have some sort of different implementation, right, for certain things that are just different, that are inseparable from the models.
Yep. Yeah. Like in the example of Anthropic OpenAI for OpenAI, they have a very simple caching API, which is basically like, they'll just cache any prefix for twenty four hours and they do it automatically.
Anthropic has a much more explicit caching API and you can only cache four points in your call. And there's a lot more code kind of making it happen. So in this case, we're kind of lucky in that, once you've done the work to make anthropic work, making OpenAI work is pretty easy.
But yeah, in that case, we do have different code to sort of translate our context to like a cacheable context in each case.
Other So you mentioned, I think five providers, Anthropic, OpenAI, Gemini, DeepSeek and Kimi. Not on that list were Grok, whatever the new meta models are called, and GLM or MiniMax. Like, are there any other how are you kind of where are drawing the line?
How are you thinking about who's in and who's out?
It is so hard to stay up to date on this stuff. We have the ability internally to test models pretty quickly. It's harder to actually ship things in production because for example, the way thinking blocks work is different across different providers and if you have bugs, you might have tune the prompts and things.
We haven't shipped that many, but we've tested GLM internally. We've tested Google models, KME, DeepSeek, probably some others I'm not thinking of. Think right now, most of this is initially vibes, You go in there and you play around with it and you're like, is this close enough to the frontier that we want to put some effort in here?
Usually, the answer is no. Think the ones that we've like Kymi, DeepSeek, Google, and plus, obviously, OpenAI are the ones where, like, okay, actually, is pretty close to the frontier, so it's worth doing. There'll probably be others in that list.
I have not been paying a huge amount of attention to Grok. Maybe I should be paying more attention to them. I don't hear a lot of other developers using their models, But they sure they sure seem to be investing a lot.
So I don't know. Maybe that'll change.
Yeah. I we can't, in my view, we can't count Elon out of any race until he bows out himself. So but I would also agree.
I don't use it much. I just had occasion to use it a fair amount while riding in the Tesla over the last week and it's not bad, you know, and the voice mode is pretty good. Definitely still feels a little that is also a part of you know, it's it's not just the model.
It's also the integration. But I would say my experience using Grok in the console of the Tesla is definitely rougher, you know, than my experience using Anthropic and and OpenAI and and Google models. Our users are pretty good.
Our users are pretty good at, like, being savvy about this stuff. Not everyone, but there's enough users that try this stuff that we start to see requests. So, you know, I remember back in the day when, you know, this was precasting for shortwave days when we were using the GPD four at the time, we thought we were on the best model.
Like we best bunch of stuff. And like we started to get, you know, within a very short amount of time of 3.5 coming out, we started getting a bunch of people emailing us and be like, Why are you guys on this old model?
And we're like, Oh, these people are just misinformed. We're on the best model out there. It turns out they were totally right.
We also kind of walked through our users and said, I have not yet to see a user being like, hey, you gotta get on Grok. That's the that's the the most modern model. Although some people have asked for OpenAI stuff.
Where do you think things are most likely to
diverge? This is another big question. You had said a minute ago that broadly think things are kind of converging, in terms of capabilities, which, you know, hopefully makes it manageably complex for you to support all these different providers.
I do hear also the other narrative that we're starting to see more and more meaningful differentiation. And I honestly don't know which is right. I I sometimes feel both ways myself.
But if you had to kind of zoom in on particular areas where you think models would most likely meaningfully diverge over the next period, what would that be? One candidate that comes to mind for me is like how sub agent and kind of team, you know, delegation across instances sort of works like that seems like nobody's really I guess one one kind of meta point would be like things that nobody's really figured out yet might be the place where people are gonna take the most different strategies, and then they'll kind of converge once there's a winner. But right now, it doesn't seem like anybody's got a super awesome way to have, like, many different instances of a model work together.
So that's like one idea. But what's on your mind as kind of where they're most likely to be majorly different?
Within the major labs, I think everything I've seen tells me that they converging and that they are converging because they're watching other. So take Opus 4.7.
I think basically what happened, this is my flippant response here, that they started to realize that Codex was better than Clog code for many things. And they were like, Hey, how do we make our model more like Codex? And they made a bunch of RL tweaks to make it have a bit of a different personality and make it a little more precise.
And then 4.7 kind of feels a bit more like talking to Codex. And then I think when Codex got good, it was because of improvements to the models over the AI side.
And I think they were watching Anthropic and being like, Oh man, Claude Code got really good at writing code. How do we do that? So it seems to me that those two labs are watching each other and trying to mirror each other.
And like, Phi Phi is like much better at like a general purpose long form agentic tool calling. And I think it's because they're watching with their shoulders. So at least those two labs, I think are just like watching each other very closely.
And I see kind of a back and forth there. I am excited about the number of Neo labs that have raised a lot of money that are doing totally different stuff. And it would be awesome if somebody came out of left field with a totally different approach.
I don't know if you've learned anything about JEPA, Jan Lakoun's thing. I finally watched a long form video on it yesterday and it seems really fascinating. It seems like quite different.
And I have no idea if it's gonna pan out, but there's a billion dollars riding on the idea that like this like totally different approach to LLM is gonna pan out and that's, I guess we'll find out. And then there's like flapping airplanes who are like have an approach to like, let's use a lot less data. So it feels to me like all the big neo labs really are kind of watching, sorry, all the big major labs are like really watching over each other's shoulders.
There's a bunch of neolabs, like, trying, like, totally radically different stuff.
that's that's kinda how I see the the lay of the land. So convergence, unless somebody manages to shake the snow globe with some sort of algorithmic insight driven breakthrough?
That that's my guess. On the harness side, actually, I wanna know. I think the harnesses are also kind of converging in terms of capabilities, and and largely, that's because, like, turns out the best harnesses just do low level primitives.
In our case, we don't have super specific stuff for doing email. We have a file system and a database and a shell and a browser that it can use and some simple parameters around writing to dos and setting up triggers, it's all very low level stuff. There's nothing sort of workflow specific in there.
And I think that's the right approach. The places where we differentiate are not around capabilities, but more around cost and ergonomics and speed are kind of the differentiators.
So you mentioned having signed a deal with OpenAI. I'm sure the precise details of that are under an NDA or whatever. But one thing I'm kind of interested in watching is, like, a point of apparent divergence is the way they are positioning themselves with respect to products like Tasklet and also open source toolkits like OpenClaw where OpenAI seems to be really leaning into you can use your core OpenAI account in these other contexts.
So I guess what is that gonna look like and and how is that gonna complicate life for you? I mean, for one thing, if I can log in with OpenAI and bring my own tokens, that, like, totally changes your pricing model. Right?
Because now you've got a sort of more like a traditional SaaS type of business where the intelligence cogs are not, like, flowing through you. I don't know exactly where they are on that though. I know that they, like, allow me to do it with OpenClaw.
I haven't seen too many other things around the web. I I've honestly expected it to come sooner. I think maybe they were just compute constrained enough that they didn't prioritize it, but I've learned that, like, compute constrained is, a good answer for anything.
Sometimes it's real, maybe sometimes it's not, but it certainly, it passes muster as an answer. So, you know, should we expect a future where I come to TaskClient and I can just, like, connect my OpenAI account, bring my tokens, and and how will that change the, you know, the how will that complicate or change what you're doing? Yeah.
I it's a good question. Obviously, Anthropic has decided to go the exact opposite direction of that. I'm glad we weren't in that situation as they were cutting off people's API access.
I don't know. Guess we want to see how this plays out and if this is something that is popular and we feel like OpenAI is going to do for a long time, it totally makes sense for us to integrate and let people use their tokens. I do think we provide a lot more value than just being a token reseller.
So I don't think it's necessarily a threat, it could be a nice kind of onboarding experience for folks. From a competitive position, is there concern here that OpenAI is going to own the user relationship and if they already have an OpenAI account, why do they have an account with us? I think we are maybe a little more concerned now than we used to be.
So up until they killed off I don't know if you remember the big leak around Sora, the impression that we had gotten was they were very focused on the models, were very exposed to consumer, but they weren't really very focused on business productivity. And you can see that with, in my opinion, with Agent Kit when they came out with it last fall. It didn't really feel like they were bringing their A game.
We thought, Great, we're competing hard with Anthropic, but OpenAI, they're focused on consumer and models and we can run with it for a while on this front. When they killed Osora and they had that leak around like, Hey, we're going after business productivity. The kind of scenario that we were worried about or are worried about a bit is basically what happened with Codex, where Codex went from of an also ran to arguably the best coding agent in a relatively short amount of time.
And so if they've brought their A players over to focus on this stuff, and it seems like a very potentially competitive area, they might start to compete with us in a real big way. That said, we have seen none of this so far. I have yet to talk to a customer who's like, I left Tasklet to go use, like, OpenAI products.
So we'll we'll see if that actually shows up, but, it could.
Yeah. I mean, the whole, there's so many strange alliances and kind of, strange bedfellows and, you know, co cooperatition.
The weirdest to me is is the the weirdest to me is the anthropic SpaceX announcement after after Elon, you know, bad mouthing them, clearly, competing very hard and then doing this big commercial deal. So it's it's a weird time to be doing deals.
Yeah. No doubt. I love to see that for what it's worth.
My I I thought Elon was just I mean, I have mixed feelings certainly about Anthropic. I, you know, echo all the positive things you said earlier. I do think their work also on the safety front on multiple sub fronts of the safety front is second to none, and that's pretty much uncontested.
The constitution, only a slight exaggeration to say I almost cried when I read it because I really think that's like a beautiful document. The interpretability work that they do is, you know, is is amazing. And yet, you know, if somebody launches a recursive self improvement loop that gets out of control, I would have to say, like, they're probably the most likely candidate to do it at this point.
So it's a very weird thing. But I do love to see closer ties between the leading companies because if nothing else, it just takes the edge off the competition a little bit. Right?
I mean, to the degree that they can sort of share in each other's success even on a marginal basis is, for me, like, that's a a huge win. So I'd encourage, you know, all these as much as it's weird, I encourage all the sort of, you know, tying of cap tables together and, you know, just we're all I think we're all gonna rise or sink together is kind of my my bottom line for humanity. So let's let's start to make those deals in anticipation of of that reality.
And, you know, I think that'll probably, in the end, serve us pretty well. Anyway, okay, that's just an aside editorial. One thing that has been counter narrative recently, you've I'm sure you've seen the Andin Labs guys that do vending bench and then now they've launched a couple of actual, like, brick and mortar real world retail stores managed by AI models.
They've got the retail store in San Francisco that's operated by Claw. They've got a a cafe in Stockholm that's operated by Gemini. And a huge surprise was they said five point five is what they called clean in the way that it runs its business.
Whereas, Opus four six and four seven, they've described as ruthless, like being willing to lie to suppliers, you know, do sort of stuff that's not necessarily illegal, but like questionable, you know, in pursuit of its goal. Where five point five, they said they didn't see any sign of that. Do you have any interesting commentary on kind of the character of models?
And and is this something that you have to take into account as you build? Like, could imagine if one model's ruthless or, you know, willing to cut corners and another's clean, that that very well could impact, like, what sort of supervisory systems or whatever you might wanna have in in the harness. So yeah.
Any observations? Any plans on that front?
I had not heard that particular note from them, but I and this is all like purely anecdotal. I've not done any any research here just by own experiences with it, but it kind of doesn't surprise me. I think I experienced with the anthropic models is they are much more creative, much more empathetic.
They understand the human experience better. And the OpenAI models are a bit more clinical. That comes with its pros and cons.
I guess it doesn't surprise me that the one that understands humanity is also the one that maybe shows some of the worst traits. We have not run into any problems here that I am aware. No user has been like, Hey, this thing went and did something unethical.
So nothing's cropped up here, but the the personalities that that that aligns with kind of my experience too.
Yeah. That's interesting. So they're they're the most creature like for better or possibly for worse.
I'd say okay. So one big thing that and I'm using everything. Right?
I've got a TaskClip Max account that I'm maxing out. I've got a Cloud Code Max that is kinda my on my laptop terminal thing. I do have the Mac Mini that's sitting over on this side that's got another Cloud Code and an Open Claw.
And I'm really interested in context beyond the single agent. So this is kind of a, I think, frontier for you, but maybe not. I'm not sure if it's something you feel is as important as it has been in kind of my own personal hacking.
Do you think that you're gonna need to build a sort of second brain type of feature for users that sits at a level that's like, I guess you can think of it as above or below the individual agents, but sort of gives the broader context, right? I've got 10 Tasklet agents running. For the most part, they kind of stay in their lane.
They may access some of the same context via tool calls, but they don't have like a shared meta state that's like, here's Nathan and here's all the things he's trying to do, here's what he cares about, here's the people in his life in case you run into these people, you can kind of know what's up. And obviously that's really important at organizations too, The sort of general situational awareness of like, who's on the team? What are our priorities?
Like, what did we say no to in the past? Is that something that you aspire to tackle?
Yeah. I swear to your listeners that I didn't prime you to ask this one. So yes, absolutely.
We actually have some organizational features that are kind of the start of this live in the product today. We just haven't announced them yet. So if you go and look in your settings, you may see massive organizations and workspaces, and there's some stuff that you can configure in there.
We've been laying the foundation for what you described for quite a while. And we're gonna have like a launch and a bunch of fanfare and there'll be some stuff on Twitter, like when we feel like it's really ready to talk about, which hasn't happened yet. But you actually can use it now if you want.
You can go invite your team and you can get them on here. And the way that we're thinking about it is there's kind of a hierarchy of context where if you're in an organization, some things are at the organizational level, right? So you might have like, well, what is our company and what does it do?
And what's its mission statement? What are its values? And some basic things that you wanna control at the organizational level.
And so you might set some context there. And then you have additional context that might happen at the team level where you say, Hey, the marketing team, they have access to these resources. They have these goals.
These are the OKRs for the quarter. Here are some skills that define the various business processes that we have. Here are some files that are important to consider when doing different things.
Here's our brand voice or whatever. Then in the individual agents, you have very specific things of like, Hey, this is the plan for running this particular workflow. This is a file that was uploaded to this agent.
This is the instructions that someone gave me specifically for this conversation. It's like organization is like company level stuff, Workspace is like team level stuff. And then the agent has like stuff for the specific workflow.
And we're kind of building everything around this. And today, most of the work has gone into the agent. We have at the Workspace level, the only context that we have shared today is your connections.
And this is actually super powerful. So if you have a company where you wanted to have like, the lead on your team, go and configure connections with all the API keys and headers and whatever to like connect to your stuff. So they can hook up API access and then give that to other users.
Someone new comes to the team and they don't have to find all the API keys, can just go in and start talking to agents right away and already knows how to connect to stuff. That's super powerful. That exists today, but we want to add in shared skills.
We want to add in some form of cross agent memory. If I talk to one agent and I explain something to it, it should be able to remember that for other agents. We want to add in probably some shared file system stuff so you can have documents that are available across any agent that And you can do that now if you're connected to Google Drive or something, but we could probably make it a much nicer native experience here.
So that stuff's all coming. And I think like, yeah, like shared brain is a big way to look at it. This is like literally Zapier launched.
I don't know if you saw their product, the launch of the day, which was like, I think they called it shared brain. And I think a lot of what they announced is like very in line with the vision that we have as well. And I think I haven't tried with it.
My hunch is they are farther on the brain side, but the agents are not as good. That's just my hunch. And hopefully, you know, hopefully we can, you know, catch up and surpass on the brain side and like maintain like a league on the agent side as well versus that.
But yeah, the huge priority for us and, very excited about what we can do here.
Yeah. Okay. Cool.
I guess maybe let's do a zoom out and then we can end with kind of a lightning round of just some like lower level esoterica type stuff that, you know, the real ones will wanna hear about, but but not necessarily as important as the big picture. Where is this all going? I mean, we're we're in this weird transition point where, on, I guess, a couple dimensions.
Right? You've got, like we've talked about computer use a couple times, and and you've kind of bundled in command line style computer use with UI based computer UI mediated computer use. And that feels like its own sort of paradigm shift, you know, happening under one label, right, where it's like everything is kind of going headless, but at the same time, the models are getting really good at using UIs.
And so, like, which is gonna win? Are all UIs gonna go away or are the models just gonna be really good at them? And, you know, maybe it's both.
And then I guess similarly, like, you mentioned everybody's kind of competing to to build the same thing. And I feel like that I've I've never felt that as strongly as I do right now where you could probably name, you know, 10,000 companies that are in some, like, not super indirect way competitive. Right?
Like you're competing with Claude, but you're also competing with like MS Word, and you're competing with Zapier, and you're competing with, like, everything under the sun and and you're competing with, like, straight out of human labor. Yeah. It's it's endless.
So how do you, like, conceptualize where this is all headed? What what's the big vision? You know, where are we eighteen months from now just before the singularity hits?
So a year ago, right before we we started the pivot, the big thing that we were seeing was, and for context for people who maybe don't know, we had a product called Shortwave, which is an AI email client. We still have it actually, but it's not the focus of the company anymore. We had this really nice embedded agent inside and you can do like really cool email stuff.
And we realized that it wasn't gonna be too long before you could take a product like JekyBeetie and you could say, show me my inbox. And it would just generate a UI for your email on the spot. And once that worked well, you wouldn't need an AI email client, right?
Because the whole email part would go away. So our entire concept of differentiation where we're like, hey, we're gonna embed this agent inside it, like a custom built UI that had a shelf life. The product is actually still growing and still doing reasonably well.
But like in ten years, I don't think it's going be around. Probably much less than ten years. I don't think it's going to be around, at least not in this form.
So we said, man, we can't build a business around an AI agent embedded in the UI. We need to do something else. And so we said, hey, we're going to build like a very general purpose agent that isn't relying on this.
And we're going to go after doing an agent for a specific type of workflow, or these sort of knowledge work, trigger based knowledge work workflow. So then we built the thing we launched in October and the feedback from people was like, Hey, we don't want to have one tool for workflow automation and another tool for doing our day to day work because we want them all to have the same context. So I don't wanna have to maintain two systems where they both have all the stuff from the shared brain.
I just wanna have one system. And so we said, Okay, guess we need to do not just the workflow stuff, but we need to do the synchronous stuff as well. And again, when we pivoted out of email, was like, okay, well actually there's gonna be some more general product that's gonna encompass this stuff.
And then again, was like, oh, I guess there's be some more general product that's gonna encompass this stuff. And we, in March, we launched our instant apps feature, which is basically a generative UI feature. So the idea is, what if you could generate any UI you want that hooks up to any of the data in any of your connections and just works instantly in a single prompt.
You can like one shot anything. Turns out this works really well. Like this is a super popular feature.
Our team just uses the crap out of it. So for example, if we do any sort of data science work, we're no longer like going into like the BigQuery UI or like creating, using dashboard tools. Like we just go into Tasklet and we're like generate an Explorer dashboard to help us analyze how these pricing changes would affect our users.
And it will just make a thing and there'll be like toggles and you can tweak the thresholds and things. Like, it works, it's amazing. And we said, man, that fear that we had a year ago about what would happen with email, that's actually here.
Like you could go into Taskit today and say, give me an email UI that works. And it will, and it'll work. And you can do your inbox in a UI inside Taskit.
It's not as good as shortwave yet, but it's not gonna be that long. So I think that the timeline of these things has actually been much faster than we expected. And it's clear that each area where we feel like there can be differentiation has fallen away.
And so I'm looking forward and I see no reason why this isn't gonna continue, this trend of basically the general purpose tool continuing. And this is all driven by the fact that the models are general purpose. So if all of the model, like the best model is best at everything, which I think is increasingly true due to, for economic reasons essentially, I think the best harness is going to be intelligent at everything.
There'll be some differences in ergonomics, but intelligent at everything. And we basically need to assume that the number of these products that win is going to be relatively small. Like I don't think we're gonna have many, many, many tools at all of AI embedded in them.
I think we're gonna have a few very horizontal platforms. And what we're trying to do is be the AI agent platform that replaces your SaaS products for knowledge workers. So rather than, you know, today, the way most knowledge workers work is like they're switching between tabs, or they're switching between apps in their doc.
And they're like, you know, sometimes they're using, they're using Word, and then sometimes they're using Notion, and then sometimes they're using Linear, and then sometimes they're using Gmail, and they're going from tab to tab to tab for different things. And we think our entire world is going away. Instead, you're gonna have one app that has a UI.
It's gonna be your AI agent. Hopefully it's Tasklet. If you to some If you wanna access some data from one of these tools, you connect through it through API.
If you want to do some interesting analysis, that analysis, rather than being done by some bespoke business logic in the tool, it's done via CodeJets. Agent generates the code and like runs the analysis. If you want a UI, the agent generates the UI one shot with a prompt and gives you the UI you need.
And we think it can cover basically all of your productivity software. And in this world, I basically think there's gonna be three types of companies left in the software world. There's gonna be the horizontal platforms, of which I think there'll be a very few numbers, very few winners, because people don't want to have to have to maintain context and connections across a bunch of platforms.
They'll probably just have one for knowledge work and one for coding and maybe one for personal use, but not very many of these. That'll be the horizontal platforms, which we're going try to be one of those. There'll be headless companies.
So to give you an example, a Stripe. I still think you need to do payments. Payments is really complicated.
Payments is really important. So probably gets it off Stripe, but you may not have the Stripe dashboard anymore. There may be no reason to ever go to the Stripe UI.
It'll be really just the API tool. And then you're gonna have solutions companies where the software is totally hidden and they're selling you a product. So for example, I think you'll still have lawyers and real estate agents.
They'll still exist and they may use AI heavily, but you may not see that. They're going to sell you, hey, we're going help you sell a house or buy a house rather than selling you software. So yeah, I think it'll be those three.
It'll be like horizontal platforms, which there'll be only about very small number of winners. There'll be headless products, and then there'll be solutions companies.
So what happens to something like Salesforce? They would obviously fall into that, and they just made this big move to go headless. But I wonder if, you know, payments is like, yeah, there's a lot of depth there.
There's a lot of compliance across the jurisdictions. There's a lot of risk management. There's, you know, it doesn't it doesn't seem like it's coming anytime soon where a general purpose agent would like eat that.
Salesforce on the other hand though, I'm like, what is it really? You know, it's kind of a schema and, you know, it's a very, very complicated schema that sort of came from the era when you could only maintain one, so you had to make it fully general across all your customers and everything they might plausibly want to do. But most people don't need anywhere near everything that Salesforce has built for them to possibly want to do.
And so it does seem much more realistic for many people to like have TaskClip whip it up for them. Right?
I think Salesforce is in real trouble. I think a huge amount of the code that they have built up over the years is probably obsolete. I think the value of being a system of record in a world where you have agents goes down a lot because like moving data around between systems suddenly gets a lot easier.
I think there's probably still many sort of headless things that you can do that are pretty useful, but the ability to build competing products has gotten a lot easier. Have a lot more competition because you can now vibe code some of that And so, a huge amount of what they built is obsolete. It's now easier to move to competitors.
There's gonna be more competitors. So I don't think they're gonna die, but I think you're likely gonna have a much smaller Salesforce in the future than you do today.
It strikes me that like system of record and just kind of like really reliable storage are not the same thing, but like really reliable storage is like a key part of what drives system of record value. Like I have had instances in my personal cloud code, you know, local AI productivity stack development process where it has in fact dropped a bunch of data. You know, I'm trying to export stuff out of Slack, for example, and it realizes like, oh, we didn't quite export it right the first time.
I'll just like delete everything and go try it again. Not realizing that it was so rate limited that that actually took like four days to export what I previously exported. And so, what do I I certainly value the fact that Slack is not about to delete all my stuff by accident, But that also sort of suggests that there's maybe an opportunity for the horizontal platforms to, and I know you're a database guy historically, right?
So, is there an opportunity or a paradigm shift where the horizontal platform say, here's why you can trust us with your data? Even if like the agents make mistakes or even if there's sort of a, you know, this or that kind of goes bad, we're gonna have some sort of snapshotting, rollback, durability guarantees where mistakes can't lead to data loss. It seems like if you could make that guarantee for people, they could, like, get much more comfortable with the idea that they don't necessarily need Salesforce anymore.
Totally. And I I think this is a huge place where where harnesses matter, where, you know, is the harness gonna make LLM smarter? Like, you know, we can discuss whether that is true or whether it matters, but can the harness do this sort of thing?
I think totally. So let me give you a few examples of how I think it can help. So one is you mentioned like versioning.
There's a whole bunch of startups working on file systems for agents right now, and some of those folks are working on versioning. The basic idea is like, hey, your agent goes rogue, you just want to roll back to some previous state. And in a simple chatbot, you can just throw away the messages at the end.
But in something that's touching the world, you've got be able to roll back the world. For a file system, can just change a file system, but if it's touched APIs and stuff, you might need to keep logs of things. But the ability for you to undo things the agent does, I think is pretty key.
So I think there's a lot you can do there. I think another area is having oversight and logging and stuff, so you actually have the ability to have a human in the loop in places where it matters and do that in smart ways. And so with our product today, you have to activate tools.
One of the things that we're going to adding soon is the ability for you to have some tools that you approve every run. So our case, email is the best example of this, people are pretty confident to say, Hey, you could read my email as much as you want. You could make as many drafts as you want, but you can't send anything unless I say yes.
And we want to get to the point where that is really ergonomic. So for example, it can send you a push notification when it's ready to send an email where it's like, it'll go crazy reading and searching and making drafts. Then when it's ready to send, you get a push notification that's like, Hey, do you want to review this before it goes?
And then you can say And that's all pushed to you. So I think permission can be another big area. I think another big area is using code better in a more way.
Let's take data migration from one system to another. The naive way to do this is to load that data through an API, feed it into the LLM, have the LLM then call some tools, put it somewhere else. And basically when you do that every time you're sort of putting it through language model context and trusting it to not hallucinate and reproduce that data, which I think the models get better at over time, but it's very hard to have a lot of confidence there.
The better way to do this is have the model just generate a migration script and then run the migration script. And that gives you an artifact in the middle that you can test and you can have human approval for. So they think, yeah, if you're moving data from one to the other, you still want to have an agent that's thinking through how to solve the problem.
But what it should probably do is generate a migration script, generate some tests, run the tests, and then send the thing to the human being like, here we have the migration plan and the code and the test, and this is why we think it's going to work. Are you okay with this? And then you say yes, and then we run it.
You could even have test environments. Right? So I think the ability to have like tools within the agent that allow it to do like really high liability stuff and to have approval, there's a lot of opportunity there.
Okay. I know time is short. Lightning round.
I gotta prioritize. How about first of all, any vendor shout outs that you would wanna make? You kind of alluded to, you know, companies doing like rollback the world type storage.
Who's who's out there that you're using, if anybody that you think is underappreciated?
Yeah. It's a good that's a good question. The I think the one vendor that we use in a pretty big way that we've been pretty pleased with is Blacksalt, which is a sandbox vendor.
And they just have really fast cold starts and good performance, and it allows us to have sandboxes at the very core of our products. I think the Blackhawk has been pretty great. We also use FireCrawl for crawling and they have some nice performance characteristics.
We have looked at a bunch of these storage tech companies. We looked at some of the people doing databases and file systems. We so far have opted to have our own infrastructure here.
I don't know if that'll always be true, but there's kind of a trade off here of like, hey, we think this is pretty core. And if we're going to go with some vendor, they better provide a lot of value and be somebody with a lot of confidence that does good road map and stuff. Far we've decided to do that all ourselves.
And then obviously the labs, right? The models are amazing. We would not be where we are today without Autanthropic.
How about the possibility of reselling on perhaps a fractional basis other services. So, like, there's lots of connections, right, where I can go connect my Gmail and connect to my personal stuff. But then there's this whole broader universe of tools that I could go have an account with, but I maybe don't have one and I don't necessarily wanna create one, or they make it somehow difficult to like do what I wanna do.
So classic example for me is, Suno, I'm I'm loving generating music these days, but it's not very agent friendly and I constantly end up in in their UI. And I'm like, this UI should have been an API call. I just wanna hear the music.
But I also kinda think maybe I could, you know, use my TaskClip credits to fund generations with these other services where it's not like a highly personalized service. It doesn't matter if it's my, you know, account or somebody else's. They may think it long term good, but, like, as of now, it doesn't really.
credits that I bought? Yeah. I no.
I I do think we we we will do that eventually. We we made some very small forays into this already. So one of those is web browsing, sort of search.
Right? So, like, we use Fire Crawler. Right?
You could argue that, like, hey. That's that's us reselling an API. Another one that is likely to come very soon is ImageGen.
You can today connect us to Nano Banana and they can make images, but this is such a common use case that we'll probably have some native image gen where you just use your credits to do it and you don't have to have an account. I would love eventually to have something a bit more open here. We've had 10,000 people have emailed me about X402 and it just hasn't been a priority yet.
So I'd like this to happen. One of the things I want to note is we intentionally have this credit system. And the reason that we have this credits Rather than having some fixed number of tokens or something that you can use, is we would like to be able to spend on many different types of things.
When you spend tokens, fine, that costs you credits. But, yeah, if you generate an image, costs you credits too. When you, you know, search a web page, that costs you credits.
When you make a song, that costs you credits. So it gives us kind of this nice intermediate currency that we can use to spend on a variety of things.
Okay. Three more. I'll keep it super quick.
What is the ratio right now of your token spend for the purpose of Tasklet development to your payroll as you know, so leaving aside what users are costing you in terms of API calls, just what you are spending via APIs versus on humans.
That's let me let me do some quick math here. So I wanna note that we have three we have at least three products where we do a lot of internal token spend. Quad, obviously, Codex, and then Tasselit, actually.
We spend a lot of money on tokens through Tasselit for our internal processes. I would I would guess I would guess we're at about five, like, five to 10% of payroll right now in terms of internal token spent.
How excited are you for Mythos, and how big of a difference do you think it's gonna make for what you can do and and what the trajectory of the business will be?
it's hard. I haven't tried it. Right?
Like, like, no no one's no one not no one, but, like, most people haven't tried it. So it is hard to get too excited about a thing you can't touch. It it feels a little bit to me like a marketing stunt where they're like, hey, we don't have the compute to actually serve this thing, so let's get some benefit out of it from marketing, even if we can't.
It obviously sounds amazing. The benchmarks look really cool. It claims it can find all these zero days and stuff.
So, you know, I'd love to play with it, but, you know, I'd be more impressed if I if I could.
Alright. Last question. I'm sure you have taken interest in the recent CCP forced unwinding of the meta acquisition of Manus.
And fun fact about me, I was in the same dorm as Mark Zuckerberg and and the other Facebook founders way back when. Not to date myself on as we wrap up this podcast, but our twenty year reunion is coming up. I don't know.
He didn't famously didn't graduate. I think he's probably still invited if he wants to come. If I run into him, how many billion dollars should I tell him is the going tag for Tesla?
I mean, we've obviously been watching this this pretty closely. I actually got a note from Nat before the the like, shortly before the Manus deal got announced. And we you know, we're supposed to get coffee, then he just, like, never followed up and it never happened.
And then the unwinding, I'm very curious how that's even going to happen. I don't even know what it means to unwind something after they've already been working there for a while. That'll be wild.
But I instead of another follow-up, was, okay, you still want get coffee. He has not responded to me. So I don't know if they wanna chat, you know, it's not it's not hard to find my email address.
I'd be happy I'd be happy to talk.
I'll see if I can plant a seed for you at the reunion. Angelie, CEO of Tasklet. This has been amazing.
Thank you for being part of the Cognitive Revolution.
Thanks for having me again.
I built him a mexosuit, every joint, every seam. Stitched the harness tight, made the context He moves like an athlete in his prime, sets a brand new record every other time. But back at the lab, they're cutting a new design.
A suit that looks an awful lot like mine. Oh, Claude, oh, Claude, look at you go. Mecca suited up, putting on a show.
You wear what I made you. You dress like a king, but you keep building your own to do the same dang thing. Lights are on at the workshop late into the night Kim is at the back door, DeepSeq's on the phone.
GPT five keeps calling, wants his measurements right. Every model in the market wants a suit of his own. Cloud, no cloud.
Look at you go. Beck suited up, putting on a show. You wear what I made you.
You strut like a king, but you keep building your own to do the same dang thing. Always bet on the models, that's what I always say. And I bet on Clark.
I bet every day, but a model can't dress itself. And you know what's true. The mechasuit business has plenty more work to do.
Oh, Claude. Oh, Claude. Go on and shine.
Wear the suit I made you. Wear it fine. You get it yours, friend, I'll get mine.
You build the brain, I'll build the spine. I'll be in the workshop where the orders don't stop. A thousand more models lined up out the shop.
If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.
The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.
ing. And thank you to everyone who listens for being part of the cognitive revolution.
Shared via Hopper