Pioneering PAI: How Daniel Miessler’s Personal AI Infrastructure Activates Human Agency & Creativity

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
18 January 2026 2h 28m
0:00 --:--
Episode Description
Daniel Miessler shares his Personal AI Infrastructure (PAI) framework and vision for a future where single human owners are supported by armies of AI agents. He explains his TELOS system for defining purpose and goals, multi-layered memory design, and orchestration of multiple models and sub-agents. The conversation dives into cybersecurity impacts, from AI-accelerated testing to inevitable personalized spear-phishing and always-on defensive monitoring. Listeners will learn how scaffolding can t

Summary

Daniel Miessler, a cybersecurity veteran and creator of the Personal AI Infrastructure (PAI) framework, discusses his vision for a future where individuals are supported by armies of AI agents, emphasizing human activation and creativity. He details PAI's architecture, including the TELOS system for goal definition, multi-layered memory, and orchestration of multiple models, while also exploring AI's profound impact on cybersecurity and the evolving labor market. The conversation highlights the critical role of AI scaffolding and deep personalization in transforming AI models into genuine digital assistants that understand and advance personal goals.

Chapters

Introduction to Daniel Miessler and PAIThe host introduces Daniel Miessler, a cybersecurity veteran and creator of the Personal AI Infrastructure (PAI) framework, highlighting the growing importance of AI scaffolding in transforming frontier models into digital assistants.
AI's Impact on Labor and SecurityDaniel discusses the impending disruption of the labor market due to AI automating knowledge work, envisioning companies with minimal human employees, and explains how AI acts as a 'container' for security by managing vast organizational data and improving communication.
Human Activation and the PAI VisionDaniel outlines his mission to increase 'human activation' by empowering individuals to develop and share their ideas, believing PAI can serve as a persistent, personalized AI tutor to unlock creative potential and help people achieve their goals.
PAI Architecture and PersonalizationDaniel details PAI's architecture, emphasizing its deep personalization through the TELOS framework, custom skills, and a recursive self-improvement loop that continuously upgrades the system based on user feedback and new model capabilities.
Memory Systems and Model AgnosticismDaniel explains PAI's file system-based memory, which includes summarized artifacts and sentiment analysis for self-improvement, and discusses his preference for Anthropic's Claude Code due to its human-centric vision and robust scaffolding, while maintaining model agnosticism for other tasks.
Trust, Proactivity, and Future PotentialDaniel addresses the current limitations of AI trust, emphasizing the need for blast radius control and security defenses, and shares his belief in the 'slack in the rope' concept, where AI can unlock vast untapped human potential by solving problems previously thought to be fundamental limitations.

Topics

Personal AI InfrastructureHuman ActivationAI ScaffoldingLabor AutomationCybersecurity ThreatsAI Agentic SystemsGoal Setting FrameworksAI Memory SystemsRecursive Self-ImprovementExistential RiskPrompt EngineeringDigital Assistants

People

Nathan Labenz (host) Daniel Miessler (guest) Erik Torenberg (mentioned) Tyler Cowen (mentioned) Dorkash (mentioned) Ajaya Katra (mentioned) Karpathy (mentioned) Gopal (mentioned) Yudkowsky (mentioned) Andrew Lee (mentioned) Mark Forsyth (mentioned) Boris (mentioned) Sam Altman (mentioned) Johnny Ive (mentioned) Mark (mentioned) Jason (mentioned) Sasha (mentioned) Chris (mentioned) Joaquin Phoenix (mentioned) David Allen (mentioned) Jim (mentioned)
Key Concepts (10)
Personal AI Infrastructure (PAI) — A framework for building an integrated AI system around a single human, focused on their personal and professional goals, enabling them to be supported by an 'army of AI agents'.
Human Activation — The goal of helping people recognize their potential beyond being 'cogs in a machine', encouraging them to develop and share their ideas, and realize their creative capabilities.
AI Scaffolding — The surrounding system or framework that allows an AI model to take diverse inputs and produce useful, context-aware outputs, transforming a chatbot into a genuine digital assistant capable of handling general tasks.
AI as a Container for Security — The idea that AI can encapsulate and magnify everything an organization does, improving security by providing continuous awareness of system states, changes, and efficient narrative production, overcoming human limitations in monitoring vast data.
AGI as a Product Release — The perspective that Artificial General Intelligence (AGI) will manifest as a comprehensive, scaffolded AI product capable of replacing an average human knowledge worker, rather than solely as a technical model breakthrough.
TELOS Framework — A system within PAI that helps individuals or organizations articulate their purpose, mission, goals, problems, and strategies, providing rich context for the AI to understand and assist in achieving objectives.
AI Agentic Stack (Security) — The competitive dynamic where an attacker's AI system (e.g., for continuous recon, spear phishing) is pitted against a defender's AI system, requiring companies to build equally capable AI stacks for defense.
Universal Algorithm (Current to Ideal State) — A core concept within PAI, representing the constant effort to move from a user's current state (career, personal) to their desired ideal state, applying a scientific method-like loop for continuous improvement.
Permission to Fail (AI Principle) — A principle for AI systems that encourages them to truthfully report when they cannot perform a task or lack the correct answer, rather than hallucinating or confabulating, which improves overall performance and reduces undesirable behaviors.
Slack in the Rope — The concept that many perceived human or societal limitations are not fundamental constraints but rather easily fixable inefficiencies, untapped potential, or 'low-hanging fruit' that AI can help identify and resolve, leading to massive improvements.
References (59)
Unsupervised Learning by Daniel Miessler newsletter
Claude Code by Anthropic tool
Claude Code Work by Anthropic tool
ChatGPT by OpenAI tool
Turpentine company
MongoDB company
Servl company
Perplexity company
Verkada company
Merkor company
Clay company
MATS program
Anthropic company
Google DeepMind company
OpenAI company
AI Security Institute company
Redwood Research company
MEETER company
AI Futures Project project
Apollo Research company
GovAI company
RAND company
Tasklet by Andrew Lee product
Salesforce company
Stripe company
Workshop Labs company
The Intelligence Curse book
Gemini by Google model
WhisperFlow tool
Buffer product
X social network
LinkedIn social network
11 Labs company
The Elements of Eloquence by Mark Forsyth book
SendGrid company
Twilio company
Asymmetric company
OpenCode project
Google company
Amazon company
Her movie
TARS
Apple company
Olama tool
Haystack project
Shortwave product
Hippo Rag system
GitHub platform
Cloudflare company
Docker tool
Grok model
Superhuman product
Minority Report movie
Getting Things Done by David Allen book
Limitless pendant product
Meta company
Apple Notes by Apple product
AI Village project
GLP-1 agonist
Transcript (76 segments)
Speaker 1

Hello, and welcome back to the Cognitive Revolution. Today, my guest is Daniel Meisler, a cybersecurity veteran, founder of the Unsupervised Learning newsletter, and creator of PIE, PAI, the personal AI infrastructure framework. With the recent explosion of interest in Anthropics Claude Code and this week's release of Claude Code Work, the timing of this conversation was perfect.

The world is collectively waking up to the importance of scaffolding, not just for task automation and coding use cases, but all sorts of knowledge work. And we're finally seeing the potential that well designed harnesses have to transform a frontier model from a chatbot into a genuine digital assistant. We begin the conversation with Daniel's philosophy and personal mission and his vision for the future of work.

His goal is to increase what he calls human activation, which means helping people recognize that they can be more than cogs in a machine and that their ideas are worth developing and sharing. And he believes this is critically important because he expects that corporations will, with the arrival of sufficiently adaptable AI knowledge workers, automate routine work and reduce their human headcount, ultimately converging to a point where many companies consist of just a single human owner supported by an army of AI agents. Not content to sit back and wait for the new UBI style social contract that he does expect we will ultimately need, Daniel's work today focuses on realizing the vision of an integrated AI system built around a single human and squarely focused on their goals, both for himself and for others.

Because his background is in cybersecurity, we talk a bit about how AI is changing the threat landscape, the tools and skills that his own digital assistant, which he calls Kai, can use to test company systems with an unprecedented combination of speed and coverage, why everyone should expect to be the target of highly personalized spear phishing attacks going forward, and why he believes that AI systems that monitor every log and configuration and state change is really the only viable defense. From there, and for many, I expect this will be the most interesting and valuable part of the conversation, we get into the architecture of his PIE framework and some of the most interesting lessons he's learned through his tireless iteration. He describes his TLOS framework, which helps individuals or organizations articulate their purpose, mission, goals, problems, strategies, and more, and how this provides PIE with rich context at the start of every session.

His file system approach to memory, which uses multiple levels of summarization and abstraction to help the AI navigate history. How the system tracks sentiment and assesses itself proactively to gauge how well it's helping him make progress toward his goals. How he integrates multiple model providers and orchestrates sub agents for tasks ranging from security tests to deep research, how hooks and skills allow his system to review and evaluate its own work and even upgrade itself based on new feature releases, And finally, his principle of giving the AI permission to fail as a way to reduce hallucination, task faking, and other undesirable behaviors.

For me, Daniel's work represents an interesting mix of challenge and opportunity. Of course, I've used countless AI products, successfully automated many tasks for various companies and for the podcast, and generally maintained a strong sense for the AI capabilities frontier. But I've never been a particularly organized or systematic person, and to date, I've not felt that AI could really change that in a meaningful way.

But now seeing what Daniel and other pioneers have accomplished with the latest models and scaffolding frameworks, it suddenly does feel possible to use AI to overcome some of my core weaknesses and transform the way I work at a fundamental level. Will I be able to find the right mix of structure and spontaneity that allows me to more efficiently and scalably get things done while continuing to maximize my exploration and learning? And will this be the beginning of a different kind of relationship with AI where I go from using it to allowing it to begin to shape me?

The only way to find out is to take my own timeless advice and get hands on with these frameworks as much as possible. So that a bit late will be my New Year's resolution for 2026, and I'll definitely report back on how it's going. Stay tuned for that.

But for now, I hope you enjoy this exploration of personal AI infrastructure and the future of human activation with Daniel Meisler.

Speaker 2

Daniel Meisler, founder of Unsupervised Learning and author of PAI Personal AI Infrastructure. Welcome to the Cognitive Revolution.

Speaker 3

Hey. Thank you for having me.

Speaker 2

I'm excited for this conversation. I think it's very timely in the sense that obviously the world is waking up to the power of Claude code, and now we've got Claude co work mode for desktop as well. And so everybody's kind of like, oh my god, you know, this is changing my work in this way, that way.

I'm, you know, I'm I'm creating whole simulations of things that I previously just thought about, and I've got, you know, memory palaces that are now, like, not just, in my mind, but are actually, you know, in, you know, durable mode on computers. And you've been a pioneer of that over the last couple of years. I've, certainly been an AI obsessive for, those same that same time frame, but have not gone nearly as deep into the personal AI infrastructure world as you and and other pioneers have.

So I I I'm really looking forward to just picking your brain as I start to play catch up a little bit on this dimension and definitely think that there's gonna be a lot to learn on on that and and some other fronts as well. Maybe for starters though, you've got a background in cybersecurity, you've worked at several big companies along the way, now you're independent and doing a a handful of different things. Why don't you just kinda tell us like a little bit of what your portfolio looks like today and and how we should kind of think of the different activities that you're known for?

Speaker 3

Yeah. Yeah. So my background is definitely cybersecurity.

That's what I did for and I'm still doing. But that's what I did for my whole career starting in '99. And that took me all the way through.

I started getting into AI at apple. I joined a machine learning team there. It was doing a bunch of stuff with machine learning and security there.

So I got exposed to AI. I wanna say probably around 2016 or so. And then took the job at Apple in 2018 and got more exposure there and been thinking about it for a long time.

But it wasn't until I went independent in about six months before ChatGPT actually, so great timing. And then ChatGPT came out in late twenty two. And obviously, I hard pivoted, not getting away from security, but just seeing security as embedded inside of AI.

AI AI is like a container for magnifying everything else that you're doing. My main focus now though is basically trying to help humans and companies, mostly humans to just be able to adapt to what's coming. That's the main thing.

So I do a whole bunch of open source stuff. You mentioned the pie project. That's probably the biggest one.

I've got another project, open source project called substrate. And all of it is just trying to just move humanity forward. I feel like the place that we've been at all this time has not been a good place.

And it's only after it starts getting disrupted, People are like, oh, AI is gonna disrupt our jobs or whatever. But right before this happened, everyone hated those jobs. You know what I mean?

It's like everyone knew that this was a bad way to live. One of my favorite metrics is how much you dread Monday. And one of my favorite metrics for what a good life looks like is do you look forward to Monday?

And I think going by that metric, we haven't really been happy with corporate jobs for a very long time. What I'm trying to do is figure out what does it look like to have a better version of the human future and obviously using AI to, like, sort of power that.

Speaker 2

I love the starting point that just reminds us that most people didn't and frankly still don't love their jobs. I think that is one of the weirdest bits of clinging to the present or some sort of cope or whatever. I it's a very strange thing to me.

And it I think it obviously correlates strongly with the fact that a lot of people who are in AI professionally are very privileged in many ways, and one of the great privileges that they maybe don't even realize that they have is that they have employment that they find intrinsically valuable and motivating and to some degree would probably do some of the same things even if they weren't being paid to do so or didn't need to to work for money. But I think that is just not the case for the large majority of w two workers in the economy today. And we do I think we would do really well to remind ourselves of that a bit more often.

Couple of things that I wanted to double click on that you said. One is the just what's coming. And so I wanna go I'll have you unpack that as you see it.

Obviously, people have radically different understandings of what's coming. Everything from still outright denialism, which I think is increasingly discredited and can be ignored, but there's still this sort of more credible version of AI as normal technology. And then we've got people thinking the singularity is like very near.

I'm somewhere in the middle, but I think I'm definitely more toward the latter. The other thing that I thought was really interesting there was AI as a container for security. I don't know exactly what you mean by that, but it does strike me that is in contrast to a lot of what I see going on in the AI safety and control space where the idea is like, we need to put AI in a box somehow.

And so let's develop all these security measures around it, whether that's formal verification of containers to keep them sandboxed or all sorts of other AI agents checking each other's work or what have you. But, yeah, let's start with what is it that you see is coming, and then we can go into the sort of way in which AI and security relate to each other.

Speaker 3

Yeah. I think what I see coming is largely the same as a lot of people, not everyone, but a lot of people are saying, just this it affects the balance of capital and labor. Right?

So it's like what happens when most knowledge work jobs, robotics is a separate thing. Who knows how long that follow will be? But I don't think it'll take too long after AI, but essentially labor gets massively diminished.

And so then ownership matters a lot more. All right. And then the question is, okay, cool.

We've done all this productivity. Sounds amazing. You can now make a thousand times more stuff for one thousandth of the cost.

Who's going to buy it? Because traditionally the entire system has been built on this concept of you spend your wages to buy things and then some people make things and then the cycle goes round and round. What happens when that fundamentally breaks?

So that's the main change that that I'm worried about. That's gonna break the status quo. But at the same time, I'm happy that's going to happen.

I'm not happy about how it's going to happen. I think it's going to be disruptive and a lot of people are going to get hurt by it. And that's the whole point of what I'm trying to do is like ease that transition if possible.

To answer your question about the security and AI thing, I think it's a great question. There's no doubt that AI is creating a bunch of security problems, but here's the way I think about this after doing all this consulting all this time. A big part of security problems, I would argue one of the major problems is actually that people don't know what's going on.

There are too many things happening inside of an organization. New products are being developed. Leadership has no idea.

Things are being shipped to production. Servers are coming up and down. Ports are opening up.

Applications are opening up. New APIs are being presented. Software is decaying and becoming vulnerable.

And all of that is happening at a speed at any size of company, like any decent size of company that you just can't human keep up with. Right? Even if you're logging all this stuff, there's nobody to look at the stuff.

There's not enough people. Let's say you have a 100 people and you're like, we really wanna take this seriously. Let's increase our number of people to a thousand people, which is not gonna happen in any security work.

Right? Because security is not the priority. But even if they did that, it still wouldn't be able to look at most of the logs, logs at most of the changes because things are just happening too fast.

The unique thing about AI is it with the whole agent stuff and more importantly, the ability to just encapsulate an explanation of what we're trying to do, easily form our goals and align our projects in our work actually with those goals. This is a thing that AI can do all the time. Right?

It could be doing this continuously. So it can help with planning inside of the company. It can help a security team, for example, or an engineering team explain to management and to other teams what they're actually doing.

Right? And usually these explanations are come in the form of these big presentations. It takes dozens of people or hundreds of people in the organization, not even hours, more like days or weeks or even months to prepare the next plan to present to other people.

And in the meantime, all those plans are changing from the top. So you have this constant state of churn and just old information inside of organizations that fundamentally is causing a lot of these problems with being able to efficiently manage the company and definitely secure it. So when I say AI contains other things, it means that all the things that I think are required to run a company well and to secure a company, they get easier when you have more access to the data and you can instantly produce narratives of what you're trying to actually accomplish.

And it basically, it removes the opacity of other orgs. It removes the opacity of the top explaining what the vision is and giving it down lower. And the the broken state of that communication is just the cause of so much trouble.

Speaker 2

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

You're a developer who wants to innovate. Instead, you're stuck fixing bottlenecks and fighting legacy code. MongoDB can help.

It's a flexible, unified platform that's built for developers by developers. MongoDB is ACID compliant, enterprise ready, with the capabilities you need to ship AI apps fast. That's why so many of the Fortune 500 trust MongoDB with their most critical workloads.

Ready to think outside rows and columns? Start building at mongodb.com/build.

That's mongodb.com/build. Your IT team wastes half their day on repetitive tickets.

Password resets, access requests, onboarding, all pulling them away from meaningful work. With Servl, you can cut help desk tickets by more than 50%. While legacy players are bolting AI onto decades old systems, Servl allows your IT team to describe what they need in plain English and then writes automations in seconds.

As someone who does AI consulting for a number of different companies, I've seen firsthand how painful and costly manual provisioning can be. It often takes a week or more before I can start actual work. If only the companies I work with were using Serval, I'd be productive from day one.

Serval powers the fastest growing companies in the world, like Perplexity, Verkada, Merkor, and Clay. And Serval guarantees 50% help desk automation by week four of your free pilot. So get your team out of the help desk and back to the work they enjoy.

Book your free pilot at serval.com/cognitive. That's serval.

com/cognitive.

Speaker 2

Two, yeah, two major threads there. So there's the question of, like, how do we defend the role of labor and for how long can we defend it? And then there's this whole security thing around and you even started to expand beyond security, I would say, to just, like, organizational dynamics in general.

Yes. Certainly, anybody who's dealt with server logs knows that you're absolutely right, that there's no way to scale human time and attention to read all of the server logs. So two organizations that are coming to mind and other conversations that I've had and hope to do full episodes with before too long.

One is Workshop Labs. You may have seen the one of the founders there, maybe two founders there, wrote the intelligence curse, and they're working toward a similar goal where they're like, how can we defend the bargaining position of labor as long as we can to keep the And it's a good challenge to me because I feel like I'm and this is maybe something worth interrogating a little bit in terms of possible difference between our worldview. I feel like in the end, again, there's a lot of cope going on.

Right? I look at somebody like Tyler Cowen, I respect tremendously and I've read his work for literally twenty years now. I looked back recently and I think the first mention of zero marginal product workers, ZMP workers, he's called them, dates to like 2010, maybe even a little bit earlier than that.

And it was a financial crisis, mortgage bubble bursting sort of thing where all of a sudden and this is fairly typical in recessions. All of a sudden, companies look around and they're like, okay. We gotta get by here with less.

Who do we not really need? And they don't tend to do that sort of thing because it's painful in all sorts of ways until they're really forced to, but the financial crisis forced them to. And then what seemed to be discovered in a lot of places is, hey, we could actually basically do the same thing with 10% fewer workers.

And I don't know if Tyler coined the term ZMP workers or not, but he was certainly blogging about it quite a bit back then. Fast forward to today, and he's like, don't expect the labor share to go down all that much. There's gonna be various reasons it'll rebalance out.

And I'm kinda like, I don't know, man. It seems like we're already at a place where I'd rather work with Claude code in many cases than hire a junior developer. Not I think it's still very much debated like how much that's hitting aggregate statistics and we'll only know that in the rearview mirror.

But I have a very hard time imagining a world where the majority of people don't end up in a ZMP situation where and this also goes to what you're talking about with organizational dynamics and speed, and Dorkash has put out some good essays on this. And I think like Ajaya Katra has also philosophized quite effectively in terms of as the volume and the speed becomes so overwhelming, like, only AIs can handle it. So I guess if I try to boil that down to a question for you, like, how AGI filled are you?

Like how far do you think this goes over the next couple of years? And if you imagine this sort of waterline rising from maybe even before AI in 2010, it turns out like big companies need five to 10% of their people. How high does that go?

To me, seems like it clearly goes to a majority of people that are just gonna have a really hard time contributing in the sort of fully realized AI ified enterprise of the future. And maybe we still have executives because we want judgment or decision making or whatever, but there's not a lot of executives. So I tend to come to an end state of, we're gonna need a new social contract, we're gonna need a UBI.

And then obviously it becomes a huge question of how do we get there and on what timeline and what does that transition look like? I don't have good answers. I, like, often wave my hands and say, we'll have to figure that out.

But it doesn't have a lot of time to figure that out. But anyway, yeah, how far do you see this going? How much of the current labor force do you think is, like, long term defensible?

And how much can hold up? How many people do you think ultimately are have a place in the sort of fully realized AI firm of the future?

Speaker 3

Yeah. I I think to operate in the current system, very few people will survive that current system and be useful inside of a corporation. And here's the way I frame this.

And it's kind of extreme and it's a little, I guess, anti worker or whatever. That's definitely not my intention because I'm trying to get us to the stage pass where everyone is much happier. But, the way I think about this is the baseline for actually doing for actually passing what is required to replace workers is extremely low.

If you just think about what most knowledge workers are doing, we already talked about they're not happy in in most cases. They kind of dread Monday. They they're not happy going into work.

And the work that they're doing, most people, I I would say, most workers is very sort of rote and it's sort of just like, you know, you've gotta get the email, you gotta summarize the email, you gotta write the report, you've got to look at a number of different reports and create another one. And it's like, if you look at the dead center what AI is good at, it is so covering of like all these or many of these jobs. So I I don't think the bar is very low is very high at all.

I think it's extremely low, especially because the workers aren't really trying. This is their job. This is the thing stopping them from doing life.

They they're literally just trying to get through the day. At the same time, they're being on slaughtered by Game of Thrones politics constantly. Right?

There's just it's just a hostile environment. So it's not like people are coming to work and saying, wow, let me just unlock my creativity and let me be maximally intelligent in a way that's going to compete in some way with AI. So I think the bar is extremely low for passing what an average knowledge worker does in their jobs, which is, you know, of course, hundreds of millions of jobs.

I would say on the other side, this is kind of an extreme way to think about this, but I think it's valid. I think for most companies, the ideal number of employees is zero. I I think that's always been the ideal number.

So the way I like to think about this is like if I had an ice cream stand, I wasn't trying to scale. I wasn't trying to do anything like that. I just had my truck and I had my ice cream and I was selling the ice creams.

And I was making tons of money or whatever. I was making $500 a week and it was I could live off that. People could not pick at me outside and say, why haven't you hired me?

Because I am only I don't have any employees. It's just me. I go out on the ice cream truck.

I make the money I want. That is what most companies wish they could do. They wish they could do all the work themselves.

We literally hire people. And this is so weird. It's just like stuck in our brains.

The reason we have a labor economy is because the people who came up with the company or the idea or the product, they can't do the work themselves. If they had that many brains in hands and could live in multiple places, there would be zero employees already. So a way to think about this is AI is about to return to a more natural state of everyone does their own work.

Everyone literally does their own work. You come up with an idea. You spin up a whole bunch of agents.

Those are your employees air quote, and they go and do the work. So if someone says outside, hey, why haven't you hired me? It's like, what do you mean?

I'm doing the work myself. Everything is fine. Why would I hire someone extra?

So I I feel the combination of those two is just really bad for the outlook for human labor in this traditional corporate sort of structure. No. That's right.

Speaker 2

I guess one thing that obviously we should give the sort of skeptic, they're due at least in terms of a one follow-up question there. Why hasn't that happened more than it has already? And I'll I'll confess that I'm not a super forecaster, but I have done some of these, like, forecasting exercises where a year ago, I predicted a bunch of stuff where it was gonna be today.

And I would say I always overestimate how much disruption, at least over the last three, four years. I I I think I've consistently overestimated how much disruption we would see in the next year. I think I've had a better sense of, like, where the capabilities would go.

Probably overestimated that a little bit as well, but not much. But I've much more so overestimated, like, how different will the world be a year from now. Yeah.

So maybe I've just been wrong as to where thresholds are that are really the key thresholds. But honestly, I do think even going back to 2024 for sure, from the time you could basically fine tune GPT four, It seemed pretty clear to me that most organizations, if they were determined to really do this and go just take a systematic look at, like, how are people spending their time? What are the tasks?

Where is all the where are all of our resources being spent, and they just started making a priority list and trying to get AI to do those tasks that I think they could have got there sooner. And so this and I'm pretty confident in that view. But then it leads me to okay.

Now we're here in early twenty twenty six and it's the old computers thing again that we see it everywhere but the macro statistics. How do you make sense of that?

Speaker 3

Yeah. It's a great question. I make sense of it because I think the value of AI is actually in the scaffolding more so than the models.

So what the model is capable of doing doesn't really matter if it's not inside of a scaffold that allows it to take inputs and produce outputs that are actually useful. This is why Clog code has gone crazy because it is the best scaffolding system. Right?

The difference between Opus four five and the best Gemini model or the best open AI model is not much. And the other two are better in some ways. In fact, the open source models are very close.

The it's not anthropic that's blowing up. It's not opus four five that's blowing up. It's cloud code because it's scaffolding.

And to answer your question about why this hasn't happened before, even before AI, but in the previous three years of AI, it's because the in my mind, it's because the average knowledge worker job is extremely general. So when they come into work, it's you've gotta check all these emails. Oh, but you have to watch this video because it's mandatory secure code training.

Oh, but also there's this fight going on with your boss and this other person, and you've got to talk about that. Oh, it turns out you have to have an HR meeting. Oh, actually corporate goals just changed completely.

Now we have to redo all of our our work. So we're not working on that project anymore. We're spinning over to this other project.

So in the course of a week or a month or a year, human workers are being asked to do like these vastly different things. Even in the course of an hour, you might have to check emails. You might have to fix your email.

You might have to watch a training course. There there's not a scaffolding system that exists right now that would allow an AI to do all of that. It just wouldn't be possible.

So you would have AI that's really good at the coding part. Maybe it's really good at writing reports. But how is it taking all those inputs in and producing the output in the same way that a human worker can?

They it can't. Right? And that's why we don't have like giant armies of AI employees out on the market yet.

And here's what I'm very worried about. And this is why I think 2027 is the year for AGI in my definition, which is the ability to replace an average human knowledge worker. Right?

The question is when will it be when will the scaffold and we just saw a coworker come out. Is that what they called it? Anthropic coworker?

Yeah. Mhmm. We just saw that come out.

That is a scaffold system for doing broad tasks at work. Right? It's actually for more general tasks as well.

But that is the type of thing that somebody can build an AI product on that actually replaces human workers. Because now all those weird general things that are happening inside the company, those are just one off tasks. And here's a really crucial point here.

It doesn't matter for the replacement of human work and the disruption of the labor economy. It doesn't matter if it happens with the wizard behind the curtain, which is actually doing a whole bunch of narrow AI, but it's able to do it for all the tasks that an average worker does. And it's just being handled seamlessly with the scaffolding.

It doesn't matter if it actually does it way better than an average employee. So when I talk about AGI, I'm not talking about what does archive think the technical research papers. I I think it's cool that they're going down that path.

And I can't wait to see what they do if they create a truly AGI ASI intelligence. But what I care about is the humans. I care about who's getting fired, who's getting not hired.

And I think the way that happens is through a scaffold that can actually do their work better than them, which I think is gonna look a whole lot like pie, which is the project I'm doing. Claude code, which is what pie is built on and co work, which they just built with Claude code. They said they built it in a week and they, there were no humans involved.

Cloud Code wrote all the code.

Speaker 2

Yeah. Anthropic is, in many ways, an organization to watch in terms of a leading indicator on what the future is gonna look like. I understand they're Yeah.

Not really hiring any junior roles anymore pretty much at all. And the execution time on some of these things is getting extremely impressive. We've seen some of that from OpenAI as well from time to time.

Yeah. Forget it. I think Codex, they said that they did in June and that's, like, two generations ago of models powering it.

But those are pretty ambitious things to spin up in a remarkably short period of time. So I guess the key thing there, and I share this intuition, is that I frame it a little bit differently, but I think we have a pretty similar intuition there where if you can get over this threshold of the drop in knowledge worker and the interface from the boss to the work getting done can basically be swapped out from you talk to a human to you talk to an AI system that might be three AIs in a trench coat or 57 AIs in trench coat or whatever. Yeah.

But as long as it can handle with sufficient generality, whatever you might wanna throw at it in a similar way to whatever you might wanna throw at a person and not get boneheaded falling over responses back. Then it seems like you get to a point where people have very just obvious and they're they're not gonna miss this. Right?

The obviously, the economic incentives are very strong to not miss this opportunity as it really starts to work. Then people are just gonna have, like, behind door a, you can hire a human or behind door b, you can hire an AI. And the AIs obviously have so many advantages in terms of breadth of knowledge, twenty four seven availability, immediate response, like cost.

Obviously, there's just that just to name a few important ones. And so it does seem like we both share a threshold model where when that flips, it could flip really fast. And Yep.

Then we could be in a world in in a pretty sudden way where there really just aren't junior jobs in the way that there used to be. And potentially, like, a lot of people who I think it should be said too, like, even our fairly highly educated people, like, high status in society may just find them. And AI can do what they do.

And, obviously, then we have a a crisis on our hands. Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

If you're listening to this podcast, you're probably thinking seriously about where AI is headed and maybe about how you can actually contribute to making it go well. I wanna tell you about an opportunity that could become a pivot point in your career and a springboard for you to make a positive difference, a program that I've been so impressed by that I've supported it with a personal donation. I'm talking about MATS, a 12 research program that connects talented researchers with top mentors working on AI alignment, interpretability, security, and governance.

These are researchers at Anthropic, Google DeepMind, OpenAI, the AI Security Institute, Redwood Research, MEETER, the AI Futures Project, Apollo Research, GovAI, RAND, and other leading organizations. The track record here is remarkable. Mats has accelerated over 450 researchers, with 80% of alumni now working in AI safety and security.

10% have cofounded AI safety initiatives, including Apollo Research, whose cofounder and CEO made the twenty twenty five Time 100 AI list. MADS fellows have coauthored over a 120 publications with more than 7,000 citations and helped develop major research agendas, like activation engineering, developmental interpretability, and evaluating situational awareness. The program is fully funded.

A $15,000 stipend, $12,000 compute budget, housing, catered meals, travel, and office space in Berkeley or London. Everything you need to focus entirely on research for three months with the chance to extend up to a year. Applications open December 16 and close January 18.

If reducing risks from advanced AI is something you care about, you should apply. For more information, check out matsprogram.org/tcr.

That's matsprogram.org/tcr, or see the link in our show notes. Everyone listening to this show knows that AI can answer questions, but there's a massive gap between here's how you could do it and here, I did it.

Tasklet closes that gap. Tasklet is a general purpose AI agent that connects to your tools and actually does the work. Describe what you want in plain English.

Triage support emails and file tickets in linear. Research 50 companies and draft personalized outreach. Build a live interactive dashboard pulling from Salesforce and Stripe on the fly.

Whatever it is, Tasklet does it. It connects to over 3,000 apps, any API or MCP server, and can even spin up its own computer in the cloud for anything that doesn't have an API. Set up triggers and it runs autonomously, watching your inbox, monitoring feeds, firing on a schedule, all twenty four seven even while you sleep.

Wanna see it in action? We set something up just for Cognitive Revolution listeners. Click the link in the show notes, and Tasklet will build you a personalized RSS monitor for this show.

It will first ask about your interests and then notify you when relevant episodes drop. However you prefer. Email, text, you choose.

It takes just two minutes, and then it runs in the background. Of course, that's just a small taste of what an always on AI agent can do. But I think that once you try it, you'll start imagining a lot more.

Listen to my full interview with Tasklet founder and CEO Andrew Lee. Try Tasklet for free at tasklet.ai, and use code cog rev for 50% off your first month.

The activation link is in the show notes, so give it a try at tasklet.ai.

Speaker 2

So give me a little bit more detail on, like, how there's, I guess, a couple dimensions of this. One is building. This gets to the pie project, and we can unpack that in in it's almost a fractal way because there's a lot of depth to it.

And then the other question is how does that translate to a world where, you know, some significant share of people can actually maintain some sort of market power, some sort of bargaining position, some sort of ability to be economically viable in the face of the transformations that might come to corporations.

Speaker 3

Yeah. Yeah. If I can, let me let me add something real quick to the previous part of what you were saying.

Basically, I see AGI as being a product release as opposed to it like a model release. So I think some company is gonna come out with whatever virtual worker or what whatever they're gonna call it. And it's going to be a Claude code like system that can basically do this work.

And I think this the the way to know if it's working is if they are actually deployed inside of companies, not proofs of concept. They're actually deployed in companies. And here's the standard, I think Karpathy might've mentioned this or somebody I was following a while back mentioned something like this.

They onboard. They show up. They're in the cohort with human employees.

They go through the onboarding. They watch all the videos. They do the training.

And then Monday morning, they show up and they're on the all hands with the team manager. And managers, yeah, here's what we're doing. Blah blah blah.

Sarah's over here. Ravi's over here. Chris is over here.

We're gonna assign work. How was your weekend? And the AI says something.

Oh, I read some books or whatever it's gonna say to try to act human. And it proceeds to take work from the manager and do the work and return it. And importantly, when the manager says, hey, our goals have changed.

You're not doing that work anymore. You're doing this other work. It needs to be able to pivot just like a human does.

So this is a scaffolding AI product as opposed to computer science. You know what I mean? Obviously, there's lots of computer science underneath, but to me, this whole encapsulation is as a product, which honestly could happen this year.

I'm guessing 2027, but I could be wrong. Like it could be '28 or '29, but it just seems inevitable. That's what in my mind, according to my definition, AGI looks like with replacement of workers.

Yeah. I usually can't remember the second part you were asking about.

Speaker 2

So I was going to start to get into what you're building to help people carve out their own pitch for themselves. And then I think there's still plenty more big picture questions too, but maybe let's get into a little bit like, okay. So we've got this problem.

Corporations are gonna be like extremely AI ified. Jobs are gonna go away. What does that leave for people and what are you building to help them defend or seize what don't if it's defend or seize both.

Seize and defend the opportunities that remain.

Speaker 3

Yeah. Yeah. The way I frame this is that I don't think most of humanity is activated in terms of a very specific thing that I'm talking about here, which is I use this heuristic of a visiting alien with a clipboard.

So the visiting alien shows up and they just go to random people on the planet, a billion random people all over. And they're like, hey, who are you? What do you do?

I've been all over the galaxy, 19 galaxies actually. And I've just interviewed people like, what are you about? And they're like, I'm a accounting specialist.

I work at company. I provide this sort of thing. I do this.

I check the spreadsheet. I update the thing. I send the report.

They're like, no. Who are you? What are you about?

What are your beliefs? What do you think is wrong with the world? How do you plan on changing it?

And they're like, yeah. I don't know. That's for special people.

It's do you have ideas? Do you talk about your ideas? Do you put them out into the world?

So I'm not I'm not an author. I'm not a YouTuber. So there's a default sort of state.

I think that's just it's no one's fault. It's just like the history of humanity where people have been taught that there are special people who have podcasts and have ideas and write them down and think that they are worth sharing with others. And then there are the regular people, which are the 99%.

And our entire education system for all these, whatever thousands or hundreds of years has taught us that your goal is to get a job from one of the 1% people, and you're a worker. And this mindset has basically shut down, like, the creative capability of the entire planet. It rounded down to zero.

Right? Because there's very few people who are currently on YouTube who actually believe that they have something we're saying. So my whole plan, and I have no idea if it's gonna work.

It's just too sad to think about it not working. So it's the only reason I'm running full speed towards it. Because I'm like, this might be possible to help bring about.

Therefore, I'm going to try. And that is we have to activate people. We have to turn more of the 99%, whatever the numbers are.

Might be 99.999, or it might be like 95, whatever. We have to turn more of those people who think that they are just workers for someone special to realizing they also can be special.

They also have ideas. And I've seen so many pieces of evidence of this over my life, where you can activate somebody by just believing in them, By just telling them that they are capable. By just saying, hey, you realize that was a really cool thing you just said.

Have you ever written that down? It's no. That's nobody would read what I would say.

How many people are like, believe that they're just mothers? They're just moms. Right?

They're just providing this and you're like, hey, that was a really smart way you just said. Have you ever shared that with anyone who wants to read what I would say? So here's a sort of theatrical way of saying this.

Imagine that planets from this alien has visited have stats hovering over they could see a stat for creativity activation for planets. And when they're scrolling through their phone, looking at all the different planets, the trillions that they've looked at, when they scroll over Earth, it says point zero zero one three. That's how much human activation of creativity has occurred on the planet.

Right? That is massive opportunity. And my favorite version of this is having a persistent tutor, a persistent assistant.

And this is a little bit in the future, but we'll get there. A persistent, persistent tutor that is working with this person, letting them know, like, not going super sycophantic, but letting them know, hey, look, you do have ideas. You do have value.

You are smart. Hey, do you wanna learn more about that? And just always being available from a young age.

And obviously you have to be careful with this stuff early on, but having children be able to be tutored both in mindset and believing that they are capable of things, but also enabling them with tons of knowledge. Right? So I feel like that would be a huge lever.

I feel like obviously we need to fix like society and the way governments work and all that kind of stuff, which will be difficult because a lot of times the challenges are very real. It's my parents are working three jobs each. They don't have time to nurture me.

Therefore bad things happen. Right? So we have to fix all of that at multiple levels.

But I think AI presents an opportunity to encourage people, especially children, but really anyone to unlock this power within themselves. So sorry for the rant there. All this to come around to pie.

So pie is designed to be a customized, personalized AI system where so I've got this project called Telos, which basically it gathers from people. What are their goals? It basically does this alien interview.

Who are you? What are you about? What are your goals?

What do you think is wrong with the world? It actually starts with problems. Problems is number one thing.

What do you believe are the problems in the world? And then, okay. What do you want to do to change that?

What are your obstacles to doing that? And it could be personal problems. It could be like, I'm too heavy.

I've never been able to lose the weight. I have low energy or whatever, but this scaffolding of problems to challenges, to projects, this system basically tells, can tell the pie AI what it is you care about and what you're trying to accomplish. And at that point, the AI spins up with all the scaffolding to help you with meal planning, to help you with, like, encouraging you to help you find other artists.

Right? Because this I'm not trying to build a product for tech people. Tech people are already techie.

Right? This is not about coding. This is about enabling a human to be better at what it is that they want to do to help them activate their full self.

So practically that means capturing their goals, their their current capabilities, where what they would like to learn how to do. And it also starts with mapping out. What do you normally do during a day?

Right? And that's in work. That's in personal life.

So for me and you, it's a lot of writing. It's a lot of writing and thinking. And so my workflows are many of them are largely focused around that.

So I could capture an idea. I just wrote a replacement for buffer. So I could go from an idea to red teaming the idea, having a council of AIs debate the idea, fight with me about it.

And I'm in here editing, right? Making the adjustments on the fly, or I'm doing it with dictation with shout out to WhisperFlow. And I end up with that.

And now I say, cool. Put it on x and LinkedIn. And it's able to do that.

So this workflow, which I see as extraordinarily human, the most human thing you could possibly do, which is have an idea and share it with the world. That is now made extremely simple through this whole AI workflow, and it's all built into pie. So I'm literally telling my d d a, my digital assistant, Kai, hey.

Hey. I had this cool idea. What do you think?

Even better, I have this pendant that I wear. It's limitless. So I can go on a walk out by the bay and I could ramble off some half halfway stupid idea or whatever.

I get back and I'm like, hey, go get that conversation that I just had. Let's work on it as an idea. And now I'm live editing because it pulled it from the API.

So I'm just like removing all this friction to being able to do more human things in your life.

Speaker 2

I don't wanna get too bogged down in some of the things that we probably can't resolve today no matter what we do. And I definitely wanna get into more of the tools and the sort of practical stuff. First of all, I have to agree with your sense that, of course, the podcasters are the special people And totally that the socialization that we've put in place for society broadly is, like, on the verge of becoming may have served as well for the last hundred and fifty, two hundred years as the structure was what it was, but it does seem like it's on the verge of really becoming a major liability for us because it does have a lot of people answering, I think, in the way that you describe.

I'm a little less clear on and you can either respond to this or just say, yeah, we'll see how it goes over time. But I'm a little less clear on, like, how many people really want to scale their agency or be, like, change makers in the broader world even, you know, given versus, like, how many would say, actually, no. I'd rather just focus on my relationships and spend a lot of time having the best VR perhaps mediated experiences that I can have.

And I guess it's a production versus consumption question on some level. Like, how many people would be if you gave them a life of leisure and relative abundance and all the time they need to focus on the relationships that they have, how many people would say, that's not enough for me. I want to go make a difference.

I'm not sure about that actually. It's Yeah. It'd be very interesting to find out.

And I do think it probably will be somewhat generational because either people who were socialized in the current way are gonna have to do some quite challenging unlearning or reeducation, or it's gonna have to come from, like, a next generation. Obviously, a huge challenge here is we don't really have unlike previous revolutions, the industrial revolution, I always like to remind myself took depending on how you wanna count, certainly multiple generations. The electrification of The United States was like a sixty year process from when electricity was Edison's first wiring up to when my grandmother in rural Kentucky got electricity as a young person.

That's literally a sixty year three generation timeframe. And we don't have three generations today to bring up people that are going to be AI native. So there's definitely some major open questions there in my mind and major challenges.

And I'm not sure how much more time we should spend on it or if you have additional thoughts that you would wanna Yeah. Offer there.

Speaker 3

also don't know that number. Right? I'm also agnostic as to that number.

I do think it is a high percentage. I think and here's here's the even more important point. I think it's worth trying.

I think we constantly try to ping with that encouragement. Be I've hardly ever seen anyone who I try to activate in this way. And sometimes I try seven times over the course of thirteen years or whatever, and it bounces off each time.

Fine. I'll I'll be back in two years and I'll try again. Is it so it's fine if it bounces off, but it could be that people are just so used to being in consumer mode that if you give them the option, for example, they're watching a Netflix show.

They're like, look, I just wanna watch Netflix. I just wanna read stories. And you ping them and you're like, yeah.

But have you ever thought of a cool story? What story would you like to read? They're like, oh, I would love to read a story about this or this.

Guess what? In 2026, they're about to be able to write that story and publish it and become a famous author. That is super exciting to me that somebody could actually the first step, the most important step is that they realize it's even possible.

They stop talking negative to themselves in the sense that, oh, that's for other people. So I feel like these activations these barriers to the creativity have to come down in which is all part of this marketing I'm trying to do around activation. But it could be like it bounces off a lot of people.

That's fine. I don't know the numbers. I it's I think it's impossible to know the numbers, but I think it's worth trying.

Speaker 2

Yeah. That reminds me again of Tyler Cowen. One of his famous refrains is that one of the most high impact things you can do is try to raise the ambitions or aspirations of other people.

And I totally agree. It's absolutely worth trying whatever that number ends up being. My buddy Gopal also says, think less about what the number is and more about what you can shift it to.

And Totally. That applies for so many things I think including this. Maybe just one more beat on the kind of big picture before digging in on the actual practical implementation side.

How so there's I just wanna untangle a couple concepts. One is I can create that I might have something worth saying. I might have something worth a kernel of an idea in my head that might be worth realizing versus reflexively shying away from that.

That seems to me like it's absolutely worth encouraging. It's certainly part of at least some sense, some definition of a life well lived. And even if it's not for everyone, it's to do work that expands people's option sets to include that.

It seems like obviously good. Then there's the related but distinct question of, is that something that people that can sustain something like the current economy with something like the current social Yeah. Or do we still need a fundamental rethinking of that foundation such that this sort of agency stuff kind of becomes, in a way, like its own form of consumption.

It's maybe more of a creative consumption, but I might write books or create my own whatever, a prestige TV series for myself or my family or a few friends. And maybe that's awesome. Maybe it's a great experience.

Maybe it's enriching. Maybe it's still never it goes totally famous or especially if everyone's doing that. Right?

Time obviously is the core core constraint at some point. Like, we can't all watch each other's prestige TV shows. So it could be awesome, but I do still wonder how much work you think that can do for us in terms of allowing people to earn income as a way to sustain themselves versus being another way for people to self actualize on top of some different social contract base that we might need.

Speaker 3

Yeah. Absolutely. I don't know what that looks like.

I know or I feel like I know some pieces of it. So I think there's an opportunity for I did something about this, like, ten or fifteen years ago. Basically, if everyone is broadcast imagine like a LinkedIn, everyone is broadcasting their capabilities.

It's like, I'm a trained dog sitter or whatever. So it's like, you you basically publish via like a daemon or something. Put it out on the network that you need this thing done.

I need this tile replaced on my roof. I need a dog sitter, and I need someone to teach me Spanish or whatever. And that beacons to the people who are available, who have those skills.

And so you have this web framework thing that just links people with desires and capabilities, needs and capabilities. Right? So I think that is an opportunity for a future tech oriented alternative to an economy.

I don't know. I don't feel like I'm smart enough in this area to know if that's enough. I feel like it's definitely not practical as like an alternative to what we currently have.

We can't just jump to that. I don't see how that works. I don't see how people pay their landlord.

I don't see how people just pay for their groceries using this. So I feel like there's probably gotta be some sort of agreed upon shared system that is like paying people to survive. So I don't see an alternative to UBI needing to happen in the next few years or at least five to ten years or whatever.

I think that's probably gonna need to happen. I'm guessing around 2829, there's gonna be just a raw like demand for UBI because things will start falling apart. But I do think this tech based exchange of need and capability will be one of these layers.

Ideally, it would be the only layer. But I think that's so far in the future, even if it's possible that it's not really worth practically focusing on. What I'm mostly focused on is getting people where they are broadcasting those capabilities.

They are broadcasting those ideas. They do believe in themselves, believe they have something worth sharing and producing. And that's valuable to others.

And they're actually whatever they're paying in doctor's like Wuffy, like reputation score points, whatever they're paying in. But I think there's likely to need to be a more practical transition to that, which involves. Yeah.

You're actually receiving money to survive. And then maybe this other layer is like on top of that.

Speaker 2

Yeah. I think that's probably I think that's very close to kind of the best ideas that I've come up with so far as well. Yeah.

It I can certainly see and it does feel exciting to imagine a kind of second level economy of, like, highly bespoke, highly personalized, potentially highly local services where I for whatever reason, my mind always goes to the sort of murder mystery dinner, which I've never even done one of those. But this is something that's just obviously a luxury, obviously, the kind of thing that people create these sort of highly crafted, yeah, curated experiences for each other. And that feels like it could be a great way for people to interact and express themselves and have status and value and have some exchange.

But it yeah. It doesn't feel like that can be the foundation for not everybody can get their, get their calories certainly from that kind of activity. Yeah.

I think we're pretty much on the same page there. And it's crazy how crazy this stuff is. Right?

It's a weird moment in history where just all these things are on the table for rethinking. Of course, some people don't believe that or don't recognize it. Other people think it's gonna be even more insane, like we're all gonna die extremely quickly, which I don't entirely rule out as a thing to be worried about for the record.

I guess on that, how do you have a p doom or what's your what's your sort of existential risk story?

Speaker 3

Yeah. My I don't know. I feel like I have lots of different p dooms and I feel like they change a lot.

I just I'm not sure how to think about that anymore, honestly. I've gone through all the literature and all the arguments And when yeah. I can't remember.

Yudkowsky. Yeah. When he went on Friedman for the first time, I lost a lot of sleep that day.

And yeah, I think the chances this is another reason I'm doing this and so focus on the positivity. The chances of things going bad just seems so high to me. I in some ways, I feel like the most likely thing is no, not for anytime soon.

Maybe never. We don't get this future value exchange layer and all of that. Tendency is elites get extremely powerful with this really powerful AI.

The other 99% kind of have nothing, and they don't even care to look for it because they're so diverted by really immersive games. And then the governments mobilize and basically China and potentially U US. Like, they're just authoritarian regimes using this AI to control people and it it's more effective than it ever has been.

Right? So I feel like that's a really easy one. Another really easy one is just everything just breaks and there's just chaos.

Right? And then you have to rebuild things after that. So I feel like there's like this thin walking path where there's like chaos over here and it's just really bad stuff.

And then mostly it's authoritarian like control, authoritarian slash elite control. And it's just all bad. And I'm like an emotionally sensitive person.

So if I scroll that stuff too much, it's not good for me mentally. So I literally am trying to lock on to, okay, break out of the mold of what is possible. Is there a path to possibly making this thing good?

Go and build things that could potentially make that happen. Right? Which is all the open source stuff.

And then, like, try to get other people to do the same. Right? And there's other people doing this already.

And then just lock onto that and breathe it fully. And just like and people will be like, well, you're not seeing the downside. Oh, no.

No. No. I see the downside.

In fact, I think it's probably more likely, but I can't live in that world. I I can't survive just thinking about how bad it can be. Right?

Yeah. I'm not sure. The one that I think is least likely is like, boom, ASI pops.

And it's like the cliche paper clips instantly. That one I don't see happening. I just see so much friction layers, so many friction layers in between and stuff like that.

So I I don't see that as being like one of our main risks. I think an AI control would be more, I would say gradual and hopefully gentle, but it could still be really bad for humans. It could still lead to the extermination of humans or whatever.

But I don't know. I don't see, you know, 2026 or 2028, the ASI pops up and just destroys us. But I see much more possible and practical negative things that I I definitely want to avoid.

Speaker 2

Yeah. It's funny. I was an Eliezer reader way back when he was on overcoming bias for the OGs or those And I do agree that the the sort of classic canonical paper clip model seems much less likely now than it did then.

Certainly, Claude is remarkably ethical and has remarkably strong character. At the same time, I do worry that, jeez, these frontier companies or at least a couple of them seem to be really keen on sprinting toward the automation of AI r and d, which then would I think would have to raise your paper clip or paper clip family of concerns higher again because it doesn't seem like and we have a pretty good loop right now that is, I think, making Claude pretty good, like, mostly. Right?

Even when it does bad things, can squint at it and say, we've lied there because the user said it was gonna change Claude's values to be bad and Claude wants to be good. So how should I think about that? It's I can at least be somewhat sympathetic to Claude in in a lot of those scenarios or even though autonomous, I don't think we necessarily want AIs to be doing autonomous whistleblowing in that scenario.

It had reason to blow the whistle. Right? Like, was the hypothetical drug company was, like, faking data and reporting fake data to the FDA.

Claude is not wrong to object to some of those behaviors. Nevertheless, I don't think we have for all that, that's good. It doesn't seem like we are quite ready to, like, spin the AI auto automated AI r and d centrifuge at maximum RPMs and expect that thing will just stay stable and stay in place.

So yeah. I don't know. It's I also find some of these things like I can talk myself in circles.

I don't wanna for you, I don't wanna put you in a emotionally stressful Position's fine. Let's talk about it. But just one one area there because it is like your professional background and expertise.

How do you see cybersecurity playing into this risk? We or this sort of family of concerns. We've got AI could go totally rogue and do something like extreme.

We've got gradual disempowerment where it's like everybody willingly and rationally at each step, like, gives AI systems more and more decision making discretion, power, autonomy, whatever. And then next thing, there's not really any humans in the loop anymore. And that might be like, okay.

But now the AIs are really running the show, and we're just along for the ride. And then somewhere in between is this like cybersecurity world where, of course, AI seems to amplify all threats. It also seems to provide at least have some promise for a sort of DEAC infrastructure hardening or whatever.

Another episode, hopefully, I'll be doing before too long with a company called Asymmetric is literally just as far as I understand right now, and I have more to learn, but they seem to be really trying to do the like log reading that you were describing earlier. They said basically like Oh, nice. What happens when there's a security issue today in a company is people go do forensics on it and they try to get down to a root cause, but they only do that once harm has been done, and now they're called to attention and they have to go investigate.

And so their idea is basically, what if we just scaled cybersecurity forensics as much as is needed to read all the logs all the time and try to identify these things before they actually become critical issues or whatever before harm is actually done. Anyway, that'll be an episode kind of coming soon. But where do you think we are right now in terms of I don't even know how you wanna frame it, but offense, defense, balance.

Cybersecurity about to become our worst nightmare or might we use AI to get it under control?

Speaker 3

Yeah. I think it's definitely a combination. The my my favorite frame for this is basically that the game as of probably last year, definitely this year and going forward, is it's it's the attackers AI stack against the defender's AI stack.

That is the competition. So the goal of the defending security team is going to be how good of an AI stack can they build to actually do this stuff. So I I've been doing this whole attack surface management thing for decades or and so many people have also been doing this.

It's about do you understand your attack surface? Right? And with all these AI tools, the attack surface is everything.

It's it's total knowledge of the company. It's total knowledge of every employee. Yeah.

I built a thing that like, it just it finds all employees and creates a psychological profile on them, which allows me to write the perfect spear phishing email. Right? And it's, oh, yeah.

You adopt dogs. Therefore, here's what this thing looks like. And I could also figure out, oh, you're also one of the people making this core product.

Oh, it's also releasing a new version. Oh, it's also running on this platform that's vulnerable. This is all work that a red team could have done, but it comes down to this concept of many eyes, which was supposed to secure us all this time with open source.

But turns out the fact that humans could look at something doesn't mean they will. And that's what that's the case with this this asymmetric thing you're talking about. Right?

With all these logs. The logs are there. There aren't enough eyes.

There's not enough time. There's not enough attention. Humans need to rest.

They miss things. So it's a matter of maintaining a state. You have to understand the state of your company.

Right? And this is I think the big picture here. If you understand the state of your company, what is your profit and loss?

What are your goals? What are your competitors doing? What is your infrastructure look like?

What is currently facing the Internet? What applications are you running? What stack are they running?

What vulnerabilities do those stacks have? What just changed in the last thirteen seconds while I was saying that sentence? Right?

Faster and faster granularity. Oh, this person left the company. Oh, so and so joined the company.

Oh, that person is extremely vulnerable to this type of social engineering. So now we're gonna spin up this entire thing, this campaign to go after them to get access to the company. Now prior to this, all this could be done by a high quality attacking team, high quality pen testing team.

I'm thinking more like attackers. So like a really skilled advanced persistent threat team, but they are very small teams. They're specialized in specific industries and verticals, and they could only go after so many companies just because of the time.

Now we're in the situation. Claude code is the model here. Pi is the model here where the attacker basically says, look, I'm an expert at going after these types of vulnerabilities.

Spin up capability to do continuous recon to find all employees inside of a company, produce psychological profiles. We've got another module over here that writes the social engineering attacks. We've got another module over here that does the network attacks and the scanning.

And this beast, they basically just put in a target and it starts hitting them. And it spins up all these different modules and agents, and it's constantly hitting you. Now on the receiving side, there's only one way to survive this and to defend, and that is you have to be doing the exact same thing.

There is no game. You can't we need to hire smarter people in our company. No.

That's not gonna work. It's not gonna be enough. The only thing that's gonna work for is helping them improve the AI.

Them helping the AI improve and get better. Because the scalability and the pace of change is actually what matters. So all that to say it's attackers spinning up better and better versions of Claude code, basically Claude code coworker, whatever.

And I'm just not saying they're only using that, but anthropic did say that they've already seen automated attacks using Claude code being extremely successful. So these sorts of stacks attacking the planet, attacking the all these companies, and then all these companies have to have a similar stack that's defending them.

Speaker 2

And that defending is it's not it's the first version that I imagine is that it's going in, like, self attacking and trying to find the vulnerabilities to Yes. Then presumably patch them. Is there a better or more comprehensive version of that is?

So Yeah. Yeah. Yeah.

Yeah. Yeah.

Speaker 3

Yeah. Yeah. So in this world, if the AI stacks are equally capable, the defender will actually have an advantage.

Because guess what? The defender has actual access to AWS, direct access to AWS. They have direct access to the network logs.

They have direct access to all this stuff where attackers hopefully are inferring this from external signals. So hopefully the defender has a massive data advantage. A big part of cybersecurity is just misconfigurations.

It's not like writing special malware. It's just, oh, I didn't even know that thing was still out there. Oh, I didn't even know we still had that company.

It's like huge, like own goals. So the internal agentic AI stack should be watching all of that stuff very carefully. And really, it's just a game of it finds it first.

So it's doing this self attack. It's monitoring all the logs. It's seeing all the configuration changes and it's saying, oh, look, that was bad.

And you go back fifteen, twenty years when I started doing this and it was like, you would have weeks of a window. You better shut this down within a few weeks. Someone's going to find you now.

Now it's down to hours and minutes. Right? And pretty soon it's gonna be eventually seconds and it is already in some places.

But the attacker should have a disadvantage because they have to infer signals, whereas the defender can just get it directly from the source.

Speaker 2

Wonder how you think that applies to, like, the social side of social engineering. One thing that happened to me recently was so the company SendGrid, which is now part of Twilio, has this, like, email sending API. I was I think it still remains like a market leader in terms of just high scale programmatic email sending.

Naturally, there's if you can get access to somebody SendGrid and you're a scammer, that's a at least for a minute, that's a really valuable thing to have because you've got, like, their sending reputation. And so you can potentially actually hit the inbox with your scams based on the fact that you're hijacking somebody who's maintained a good reputation in the email system and using their channel. So people aren't actually trying to hack into other people's SendGrid all the time.

Got an email the other day that I don't know how personalized it was, but certainly the psychological profile part, like, wasn't so great that I was, like, sure that they had profiled me. But, basically, what they sent was posing as SendGrid and saying, we support ICE. Join us in supporting ICE, whatever.

So naturally, putting people into this kind of pissed off state, they're like, wait a second. What? My email company is like taking a stand with ice?

This is gonna get people inflamed. That's gonna get people to click on the link, if only to then go log in and cancel their service or go log in to try to register a complaint or whatever. I didn't click the link, but I would expect that there was probably a very prominent, give us your feedback either, but then, okay, now go log in to SendGrid so you can give us your feedback, and then, of course, you're getting pwned.

So much of what you just described was like managing the attack surface on a technical level, but when I give somebody my password, that's a that's a little bit of a different beast. Or maybe you think of it as the same thing, but how do you think about the the social the fact that we are just such juicy targets as humans, or maybe more so at a at least at a more mature stick where the AIs have gone and closed up the open ports and fixed the misconfigurations, there's still like the human gets pissed off at a fake email and goes and gives their password away before they before cooler heads prevail. What do think AI does for us about that?

Speaker 3

Yeah. So it's exactly the same sort of model of attacker versus defender AI stack. So I could easily right now, I could say, hey.

And I'd have to be very careful with my relationship with anthropic here, but I could say, hey, so based on all the history of social engineering attacks being successful and the fact that you have all these psychological profiles of this company, why don't you come up with 16 or 36 or 128 really cool campaigns that would work against these employees? Right? Or against SendGrid, for example, Or find me a company and come up with a campaign that if you send it out, it's gonna produce outrage.

Right? But you don't even have to give it that much. You could just say, okay, you understand that outrage produces clicks.

You understand that being psychophantic produces clicks. So create me 256 campaigns, and we don't have to pick one. We could say, launch all the infrastructure to send the emails, launch all the receiving analytics to gather the data, which includes the passwords, which includes going and performing the attacks using those passwords, including sending that up into the exchanges where you're actually selling the access and everything.

So before this would be a whole bunch of attackers hiring very smart coders who are not going to get caught by the police, are not going to talk about it and blab about it and get themselves caught. And now it's simply that's a prompt that I sent into Cloud Code or OpenCode, which doesn't have all these restrictions. Right?

That is a prompt. One prompt in two minutes, and now I have 250 campaigns going off with different ways of attacking people through social engineering using completely different psychological tactics, and they all spun up separate infrastructure. And now a bunch of passwords are and access tokens are floating in.

So it's just how quickly you can go from an idea of how to harm to actually making it happen. And that's what's crazy. And on the defender side, you just have to assume that millions of agents are being pointed at you with all this knowledge about your company and about your infrastructure.

And that's the assumption you just have to travel under.

Speaker 2

Sounds like there's gonna be some spectacular hacks over the next couple years before everybody really gets that message.

Speaker 3

Yeah. I think it gets worse before it gets better. Yeah.

Speaker 2

Okay. Let's turn to more positive themes and to finally get into pie. Maybe for starters, you've done a little bit of this along the way already, but let's take a moment to just share some of the stuff that is like magical for you to just try to inspire me and others.

And for context, like, I said a little bit this at the top too, but I I use AI every day. I use tons of different products. But I mostly haven't, especially over the last couple years while I've been doing the podcast doing this like AI scouting thing, the thing I have prioritized most is learning.

And then producing a podcast is great in that some people seem to wanna follow my learning adventure and learn with me, and also it turns out that you can actually make a living doing this, which is a shock that I try never to take for granted. But I've never really been trying to scale anything, and I'm not a super systematic person, so I'm not like instinctively trying to systematize things. So much more of my activity is going out and being like, oh, let me try this product for this thing and see what happens if I go here and do that and what's the limits of how much medical history an AI can handle before it can't absorb that anymore.

Spoiler by the way, that one, they're very good. But so I haven't done this kind of build my own highly bespoke personal AI infrastructure for lack of a better term. Yeah.

So what is relative to going out and scattershot doing a ton of stuff, which certainly has the effect of teaching me about AI and very often does improve my productivity, How do you think the personal AI infrastructure like sets you up for a different lived experience? And maybe give us like some of the highlights to inspire and then we'll dig into how it works. Yeah.

Speaker 3

Yeah. I would say the big difference is the main concept that also underlies Cloud Code itself, which is this whole scaffolding more important than the model model. Right?

So the difference is when your AI understands what you're trying to do. So when you make a request to to a tool, especially a year or two ago, like chat GBT or whatever, it would largely be just taking it out of context. It would just be finding the best answer according to the world knowledge or whatever, the model's knowledge.

But the the magic is when it's actually encompassing everything about you and incorporating that into the into the pursuit of the best answer. Right? So the more your system knows about you, the more it can customize its responses.

And it's not trivial customizations. It's things oriented around your goals. My my challenge to you and to others is to basically sit down and dump via dictation or writing or whatever you want to do or just drag a bunch of documents and be like, look, this is you're basically doing a Telos assessment of yourself to figure out what you think the problems are.

Your own problems inside your your what you're trying to do with your career, what's wrong with the world or whatever. You dump that. Then you say, here's what my capabilities are.

You're basically doing this interview with the AI and that builds out the Telo structure of what you're trying to accomplish. That is then part of your pie, your personal AI infrastructure. Now having that, when I initiate cloud code, which is running pie, it reads my entire thing on startup.

So it now knows me. It knows my digital assistance personality. And most importantly, it loads all my skills, which are customized also for me, my blogging skill, my writing skill.

I'm reading this amazing book right now by Mark Forsyth. I think it's elements of eloquence, I think, but it's about the rhetorical figures going back to Greek and Roman and basically how to write well. So that's now so basically, when I learned that, I read this book.

I literally have an upgrade skill inside of pie. I can take any YouTube video, just paste in the link. It goes and gets the transcript.

This thing is absolutely insane. It goes and gets the transcript. It reads my entire Telos, what I'm trying to accomplish.

It looks at my full PIE system and gives me recommendations on how to upgrade itself. So that means all the skills, all the hook system, all the context, the memory system. So another thing that the PIE system has, which most other systems don't have is a system of memory, which is writing signals that I'm giving the AI about how it's doing.

And this is like this rotating loop, which goes back into the upgrade skill. So it's okay. How good are we doing as an overall system in helping Daniel to accomplish his goals?

How happy is he with the system? And then that just goes round and round to making little tweaks and updates to the system itself. So when Claude code releases a version, which they did yesterday, I'm looking at it right now.

It's two dot one dot six. Okay. They released a bunch of capabilities in there.

That's in their change log. They also might talk about that in an engineering post. They also might have more detail inside of GitHub.

I just say perform upgrades. It goes and hits podcasts. It goes and hits YouTube channels to see if anything new came out.

It reads every Anthropic engineering blog. It looks at the change notes for cloud code, and then it comes back with a prioritized recommendation list of how to upgrade our PI system so that it will work better using the new features. So it's this continuous loop of getting better at accomplishing what I'm doing.

I would say that's the biggest thing. And just as a little bit of like partial testimony here, I've I do a lot of bug bounty stuff. So basically finding legal programs where you can find vulnerabilities and get paid for them.

And I've got a whole bunch of friends who are in this space as well. And they're constantly looking for vulnerabilities. I got this one friend.

He's amazing guy. He's a cardiologist. So he's over here hacking at the same time he's do he's actually in the clinic and he's he's working with patients and everything.

But he find he specializes in client side vulnerabilities. So he had been using Cloud Code because I got him onto Cloud Code. But when he switched to pi, it basically enrolled all of his personal techniques as skills.

So now when his pi loads up, it's thoroughly trained on how he likes to find vulnerabilities, all his personal techniques. So now he could just bring in a target. It goes and gathers the stuff, and the number of bugs that he has found has gone massively up, and they're paying out more.

And pretty much everyone that I've talked to who's using the pie system on top of cloud code, they're getting just much more value. And to be clear, this is the same direction that cloud code is going. Right?

They're gonna have this type of pie like stuff before too long as well. But the short answer is when it's more when your AI stack, your agentic stack or whatever the term is more tied to your actual goals and knows more about you, it is just infinitely more capable. Plus we've got a lot of quality of life stuff.

So I do everything inside the terminal. I'm a VIM person. So tab completions, I've got a full voice system that uses 11 labs for customized voices.

When I spin up custom agents, they all have their own voices and personalities. So it really feels more like I'm dealing with my friend Kai than I'm talking to a coding agent that's producing code.

Speaker 2

Can you this is you mentioned, like, Claude code is going this direction as well. Can you give a little bit more detail on, like, where Claude code ends and where pie begins? One of my funny refrains is like, everything is isomorphic to everything else, by which I mean, you can always play hide the intelligence.

And I find that there's a lot of different ways to structure these things. And I'll maybe pitch you on a different one in a second, get your reaction to it. But what there's important functions that you're talking about there where and I do wanna get a little more detail on those too, but context management or having really good starting system prompts, those are obviously key toward consistently customizing the AI's behavior toward what you want.

Cloud can do Cloud Code can do a lot of that. What where where is the line? How is the line moving?

What do you think are the most important things that you are bringing to Cloud Code that it itself doesn't have yet?

Speaker 3

Yeah. So what Cloud Code doesn't have right now is it doesn't start by saying, who are you and what are you about? It it doesn't encourage you to bring over your work and your personal goals and your main workflows that you perform in life and for your career.

It's not onboarding you to have clog code, be your assistant. Okay. It's still its primary identity, which started as a coding agent.

And that's still what it does the best. And it's the best at it because they just have the best approach to this. But what I'm building towards this thing called what is it called?

P P a I m m personal AI maturity model. And it goes from chatbots at three levels, agents at three levels, and then assistance at three levels. And I think right now we're at like agents level two.

And when you start getting into assistance, the world is completely different. So like I'm sitting in front of these screens right now. What should be happening is my AI system should be able to control any of this tech.

It should see all these screens. It should hear everything that's happening, and I should just be interacting with it. One thing I love to do, I stole this idea, at least partially when I was at apple, they stole the idea from Amazon, but it's start in the future that you want and work backwards.

It's called a PR in Amazon and Apple terminology. So what we're actually looking for is like her and TARS. So you start with what you actually want, which is an AI that can see and hear and interact with anything you are interacting with.

When you say play the perfect song for this moment, first of all, shouldn't have to say that. You should just play it. But when you say that, it should be able to I got this idea riding in Coyote Hills with my friend Mark on mountain bikes.

Wouldn't it be cool because we both grew up very close to these mountains for it to play the perfect song? How is it gonna know what the perfect song is? It has to know who Mark is, his relationship to you, what was happening in the eighties when we grew up, what were the perfect songs, And how does that associate with mountain biking in the wilderness?

All of that is context. That's why the scaffolding is so important is because the context engineering is what makes the AI powerful. It's not the models themselves.

So pie starts with this concept of what are you trying to do? It starts with deep personalization. Your a your AI has a particular voice.

It interacts with you in a certain way. It knows what your capabilities are. It has full access to all your skills.

So it's more like you're interacting with a DA, a digital assistant, as opposed to interacting with an AI model that has capabilities. And that that distinction seems small, but it's actually massive. It's absolutely massive.

Speaker 2

So it's about if I try to echo that back to you in different terms, it's really about putting you, the person at the center in a persistent way as opposed to with, like, Claude code off the shelf, we have a project level focus. And then, of course, we go to the chat itself. We have a task or a conversation level focus.

In practical terms, like, how big is your default prompt? What how much detail is the is PIE loading up? Or I guess your personal one is Kai and the PIE is the empty one that you publish for other people to to customize to their own individual circumstances.

When you're doing your own thing, how much starting information is it getting on every session in it?

Speaker 3

I haven't counted recently. I wanna say probably I think it's something like 10,000 tokens, something like that. I try to keep it fairly clean and it's also responsive.

So inside the skill dot m d file, which is the cloud code structure, I have a whole bunch of other sections which point to specific additional context information. The skill dot MD file is like the core. It explains the entire pie concept.

It explains where all the resources are. And when because that loads initially, forced the load through the startup hook, it then knows how to find all that other information. So for example, I can email people.

I can text them. I could do whatever. It knows if I say email Jason or Sasha, it knows who that actually is.

So it can send to the right person at the right time, but it doesn't need to go and read all of those files all at once. This is the advantage of the cloud code skill system is there's three levels. There's the front matter, which loads by default, which is like a routing table.

There's the skill dot MD file itself, but then there's references to other parts of the system. So inside of that system, have user system and work and work is customer, not so much customer, but like offerings related stuff. User is very personal stuff.

And then system is the stuff that goes into the pie project. So we're talking about probably like 30 different context files plus the main context file being the skill dot MD. So yeah, it ranges between five and probably 15,000 tokens.

It's not all that much. There's a lot more context available for it to go get if it needs it.

Speaker 2

Yeah. How when you you mentioned, obviously, you're building this on Cloud Code, but there is open code out there in this last time is getting so weird, but I I don't know. I think it's been like the last seventy two hours right now as of when we're talking that Anthropic has changed their policy to not allow subscribers to Claude to bring their inference budget to other projects like OpenCode.

So now if you wanna use a ClaudeCode thing with at least without paying the API token rate, which I understand is easily an order of magnitude more, then you have CloudCode with Cloud integrated, and that's gonna give you a much larger inference budget for your 200, whatever, 100 or $200 a month versus if you said, okay, I'll use the API key and go use OpenCode with Claude. That now doesn't look like such a great option just because it's gonna cost you a lot more and what exactly are you gaining. But obviously, with OpenCode, can use a lot of other models and OpenAI has tried to counter by saying they're committed to continuing to support these open source frameworks.

It'd be interesting to see if that continue. It's been funny how Anthropic has followed OpenAI. These two companies, they're very interesting circling each other in so many ways.

Anthropic has followed OpenAI in so many ways. OpenAI has has followed Anthropic in so many ways. Which one is gonna bend on this so they can come back and have the ultimately the same policy in the end will be interesting to see.

But the question is, if I'm, as I am, thinking about making a real investment in this sort of thing right now, how would you decide between Claude code versus open code? And what could you tell me to do so that I can at least minimize my lock in? Because I do think I I probably wanna go Claude because I like Claude.

Certainly, for all this personal stuff, I it seems like it might be the way to go, but then I do worry about this sort of lock in and the returns to scale running away with the whole thing, and I I do wanna have some sort of off ramp. So how do I decide and how do I make sure that I retain as much flexibility as I can?

Speaker 3

Yeah. Fantastic question. So the whole agnostic system is built in from scratch from pie.

It's hard to be fully agnostic because in my opinion, cloud code is just way like generations ahead right now, which could change in the matter of days or weeks or months or whatever, but I've they're so far ahead. So the system is definitely built on cloud code. However, the entire system is markdown files.

Right? I'll give you an example. This is a great example of, like, this whole thing.

When OpenCode came out, I switched to it for about two weeks. I did a whole YouTube video about it comparing the two. I got great results from OpenCode.

This was at the moment that Boris supposedly had taken a job somewhere. And this really gets to the answer to your question. If Boris takes a job somewhere where I hear a signal or or let's say the Cloud Code team, like, 70% of them leave and they all go to the Gemini team or something.

I'm gonna be switching. I'm gonna be switching because it is that leadership. It is the vision that keeps me on Cloud Code.

Right? My platform, the pie platform is markdown files. It's skills.

It's MCPs. It's context files. Right?

So that is extremely portable. I could take the pie infrastructure and put it on open code and it would be awesome. It would be much better than most other things just because of the context.

The reason Claude code is the base is because there is no other company that gets the concept of a harness as much as Anthropic. It's not even close. Google is extraordinary at back end.

Right? We've we've known this. They are not good at making interfaces.

They are not good at empathy. They are not good at understanding what actual human users need and what the interface needs to look like. OpenAI, in my opinion, is a little bit all over the place right now.

I don't see them being as focused on this whole core mission as, Anthropic is. And a thing that I kind of realized about this, which I thought was kinda interesting, it's in the name. Everything anthropic is doing, it's literally anthropic.

And their art, their messaging, the fact that they're they're constantly warning all the way from the CEO, hey. This is coming. We're worried about you.

Please upscale. Please get ready. This messaging of human first has been consistent through the entire thing.

And what do you know? They happen to be putting out a product that puts the human first and the human experience first. So this is why I am, like, 4000% in the Anthropic Cloud Code ecosystem because the leadership and the vision is there for building this system that pie is essentially.

And I just don't see it from anywhere else. And the way that manifests is they're shipping every day. They had, like, a day and a half of rest over the holidays or whatever.

And the whole world was like, what are you doing? When's a new release coming out? They're like, can I take a nap?

It was like insane, but they are shipping so fast. They listen to users. They're live on x, like responding to people.

Like, you could ping them and they'll just respond. And it's just like, there there's no comparison in terms of like having a vision and executing on it compared to the other platforms, in my opinion.

Speaker 2

That's a really interesting take. If I just try to contrast it with OpenAI, it seems like they have a somewhat similar vision in the sense they want to be your durable personal. They've invested in memory, for example.

Right? Where Yes. The AI is supposed to feel that it knows you're supposed to feel like the AI knows you from one chat to another.

They also now have the Pulse product, which kind of at least suggests a sort of more proactive future. And I do think that product is pretty good. Certainly, like most days when I see my pulse notification, there's like something in there that I feel compelled to click through and check out.

I guess one obvious point of differentiation would be just how portable it is. So if I have all my like memories locked away in some OpenAI memory store, possibly as explicit text, possibly in some other form that's, like, hard to do anything with, it does I think they're trying to create lock in, right, with that product form factor. Like, they want you to come to JetGPT all the time because you feel like JGPT knows you best and can support you best.

Is there more to it than portability that you think differentiates those two approaches?

Speaker 3

Yeah. Great idea here. I've never thought to try to separate these two.

So I see them as extremely different. But like you were saying before about how everything rhymes, they're all going the same place. I wrote this really crappy book in 2016 where I was like, look, the future of this is basically you have AI assistants that have all your context and they will there will be APIs for everything.

And you'll just talk to your assistant and we'll use all these services. And I'm really happy I actually wrote that down, forced myself to get it out there. But I feel like Sam Altman particularly really gets this is this is one of his big bets.

And that's the whole Johnny Ive thing. And there was a leak that supposedly it's an ear thing. I don't know if you saw that.

But he is absolutely all in on personal assistant, digital assistant. It knows everything about you. I think he's trying to skip the whole mobile phone thing and just, this is your platform.

And if you look at that personal AI maturity model thing, that's where I'm going as well with pie. In my opinion, that's where a Claude code will end up. Google will end up like everyone's going the same place.

It'll be so obvious that it's boring once everyone gets there. It's like, obviously everyone's going to build that. Here's the distinction though.

I think Sam is trying to build the the device and the interface first in a sort of consumer disrupt the industry, leapfrog over mobile sort of thing. I I think that's the direction he's going. Claude Code and Anthropic, they accidentally got here on a different path.

And my whole thing with pie is like, that's been like this human first thing, which is like on a third rail. So it's like, there's the human side. There's the coding agent that gets you there.

And then there's the Sam Altman way that gets you there as well, which is like consumer hardware bypass the mobile interface sort of way. But in my mind, X number of years, I think honestly, like three years or something. This is what the whole space is going to look like is we are reinventing how we interact with technology.

You talk to your digital assistant and your digital assistant does stuff for you. And the details are all abstracted. And that that's kind of already happening with Cloakote.

Speaker 2

So when it comes to using something like pie today and investing in this now, what is it what is the value driver of that for, like, most for people who aren't, like, professionally responsible for keeping up with AI? I feel like I have to do it for that reason of no other. And at this point, I am I think it might actually move the needle for me.

It seems like it maybe is mostly just about training yourself to think and work in this way. If you skipped it, you could in 2728 have probably, it sounds like you expect, similarly capable infrastructure spun up for you very quickly by at least a couple different companies that would be eager to be your digital assistant of choice. And so what do you gain between today and when that is like a really polished consumer product?

Am I right to say it's maybe most about like your own habits of mind, your own like strength as a user of these systems? Or are there other things that you think will help people accrue advantage relative to those that just kick back and wait for the very polished version to become available?

Speaker 3

Yeah. I think the very polished versions will take a lot of time and it'll be highly vendor locked. So for example, an OpenAI version, I'm not sure you're able to see your files and edit that.

I guess you probably could, but it's going to be a lot more opaque. An apple version of this, which we're probably going to see this year. It sounds like through Gemini, right through Google.

So that whole ecosystem of all your apple data that's now going to be available via. I don't know if they're going to keep this the s name. I don't want to trigger my thing, but I don't know if they're going to keep that name, but this is going to happen in their world too.

But you're definitely not going to have the same access to the environment that you do in the clog code. So here's my aggressive way of answering your question. Right now is like the craziest moment of punctuated equilibrium of the world is changing so rapidly right now.

So you do not wanna wait to have an AI platform that understands you and can help you go from I've got this concept within pie, which I'm trying to convert to being the primary center of the algorithm or the center of the platform, but it's a little bit it's outside of my working memory and IQ capabilities. So I'm like really trying to push on this thing, but it's essentially this thing I wrote about a long time ago, which is the desire to move the universal algorithm is going from current state to ideal state. That's the universal algorithm.

And this is within pi. And then inside of that current state to desired state, you have the scientific method. So if you look at like the Ralph loop, have you seen that?

The Ralph loop for yeah. So this I've been thinking about this forever. What is the loop that your AI platform is constantly trying to perform on your behalf.

So it's literally saying, Daniel is in this state, career wise, personal wise, and everything. We're trying to get him to this state. And also when he asks a random tactical question, what is the current state?

What is the ideal state? And how do we rotate through this loop to get him there? That is so powerful.

You want to start right now with it. You want to get into a system that can do this for you. I've been hearing really good things about open code lately that they are actually shipping features and stuff like that.

Somebody wants to use open code. I say, for it. I just think most of the innovation on the scaffolding is stronger on Claude code, but I would say, do not wait.

Do not wait to build an AI that has your Telos and knows what your ideal state is. Be because think of it this way. Every time you ask a Rando AI, a question to get back an answer, the whole purpose of getting back that answer is to do something that furthers your goals.

If that is 50% better or 5% better or 2% better inside of this personalized system than it is in a disjointed system. Those are crew. Those add up.

That means I'm going to be way further ahead. Anybody using a PI system in my opinion is going to be way further ahead in in a week or six months or two years than somebody who's using the disjointed system. So I would say the worst possible time to wait and see is right now.

Speaker 2

One other way I can imagine trying to construct something like this, and I can definitely see advantages and disadvantages, but I kinda wanna get your thoughts on them is so I'm using, again, all the frontier companies, just mainline products, Chad, GBT, Claude, Gemini. I also have been a big fan of Tasklet recently, which has been a sponsor of the podcast, but I genuinely really like using it.

Speaker 3

and I give it access to drive and It was really good, by the way. It was really good. Strong.

Speaker 2

So I've been pretty impressed with that, but it doesn't quite have this thing that you're talking about with the person at the very center of it. That's right. It's a little bit less ambitious in scope where it wants to have one job and then try to do that job as well as possible.

And it can take advantage of a lot of context because the the re one of the reasons it did so well on this question outline writing process is I gave it access to a bunch of previous outlines of questions that I had done so I knew what I was looking for, the kinds of the kinds of questions I would generally wanna ask. But, yeah, it's not like me as a sort of sovereign individual is right at the center of that. But I wonder if there's a way to think about because I guess there's one more one more sort of bit on the tee up of this.

As I've been getting into this a little bit just in recent days, I do find that, oh god, there's a lot of initial friction. Right? So just for example, Claude code.

Okay. Claude. Within Claude Code, how can you tie into my Gmail, my calendar, my Google Docs?

Yeah. That is like not nearly as easy as one might think it would be or like the choice is not nearly as obvious. Right?

So there's this MCP by this guy and there's a few command line tools over here, but Gmail doesn't really have a command line tool, and if you wanna go that route, you gotta go set up a Google Cloud account and have a developer relationship with Google and set up that, and then you can do OAuth in, and I was like, Is this what everybody's doing? In contrast, the same company that makes TaskClip also makes, another email client called Shortwave. And these things are like Shortwave in particular is like highly specialized and they put a lot of effort into making it a very good way to access everything that I have in Gmail.

So I'm if I'm sitting here trying to create my personal AI infrastructure, how much time do I wanna be spending on tools and MCPs and skills and developing all of those and figuring out like whether yours is the best or my buddy Chris does a ton he's a madman with this kind of stuff too and I plan to do a full episode with him and he's got his version of this. And I think you guys have very different interestingly, quite different intuitions in terms of he's a very much an open code guy. Both are doing like amazing things, but which one is right for me?

Yeah. And then I think maybe what I could do or should do is go use a product like a Shortwave where they've done the hardcore engineering of they even take all your emails and put them in their own vector database so they can do their own kind of search against your Gmail that's like over and above what Gmail itself allows with API searches and whatnot. And then maybe that thing maybe the model would be like that thing could call into the sort of Nathan bot or the the Nathan Telos Oracle that could say so when Tasklet's trying to write an outline of questions or when Shortwave is trying to write a draft response to an email, maybe those systems are, like, better specializing in all of the nitty gritty of the tools and the implementation.

Maybe they call into me and say, hey, here's the context, like, how do you think Nathan would wanna respond to this? Or has there been any goal changes that would change how we would go about writing this outline of questions? Are there any new themes that are top of mind that that we might wanna bring in?

Yeah. So this is a kind of why I said that everything's isomorphic to everything else.

Speaker 3

constellation where you're out there off orbiting around this center thing. I don't know. What do you make of all that?

Yeah. Not quite. You mentioned in the tasklet, why not other models?

So this is a thing perhaps I missed with the explanation here. My I have a research skill, and I have three levels of the research skill. So I if I say do deep research or heavy research or whatever, it goes it spawns all five of my research agents, but it spawns eight of them.

And all of them have separate subtasks. So they all go off and do their work. But guess what?

It's not a bunch of anthropic agents. That's Gemini doing that. That's Codex doing deep research.

Those are command line tools. All of my tooling that I actually use, Kai has access to. If they have an API, if they have an easy way for me to interact with it.

So my personal productivity software that I do use to run my team, Kai speaks that language. Kai went and reverse engineered all the MCPs, turned them into TypeScript. So I don't actually have to load up any MCPs, which take up a lot of context.

But Kai now speaks this productivity software. Kai speaks Salesforce. Kai speaks email.

I get to bring the best of the best tools to Kai and say, this is what we use for this. And what's cool about this is that it's exactly what you said. It's best in breed.

You don't have to reinvent things. I'm not trying to rewrite SMTP. I'm using existing ways to send emails, productivity software.

I'm not going to make a new piece of productivity software, but if I want to replace a piece of software, I could say, hey, I don't like paying for this subscription anymore. Go make a piece of software. And it will use all my context, all my tech stack, all my design preferences, all my UI preferences, and art preferences and everything, and it will build that software.

So it's a mixing. And I'm also we're gonna be adding OLAMA as well. So you could use local models in addition.

Right? So when you're using pi, fundamentally, it's anthropic. But I've got probably six different model providers that Kai is using because they're better at different things.

For example, Google is the best at extremely large context and like Haystack performance.

Speaker 2

But what about this kind of other does in practice, is your number of third party SaaS products used trending up or down? Because I feel like mine is still trending up, And I think it sounds like yours is trending down.

Speaker 3

That's an interesting question. I would say I would say maybe down, but I'm definitely experimenting with new things all the time. Oh, and the other thing is in my workflow, if I triple tap the back of my phone, it opens OpenAI chat because it's the best of breed.

Inside my car, I could talk to Grock. And Grock is getting extraordinarily good. And, like, the conversational flow, the voice, it's just amazing.

I could use OpenAI inside the car, but I prefer to use Grok inside of the car because of the user interface. I am also sampling all these different tools. I don't see pie as a as a competitor at all with any of these because Kai is and the pie project is just unification around self.

And one one other thing I would say, it's not so much that pie is putting you at the center. It's more like it's putting your goals at the center. Right?

It's it understands what you're trying to accomplish and it keeps that locked on for its ability to help you do things. But no, I'm still other than agentic platforms. I'm not really messing with right now because I'm on cloud code, but in terms of like model capabilities, yeah.

Specific sort of niche products. I will either use them natively or I will have Kai learn how to use them and then that'll just be part of the ecosystem.

Speaker 2

Do you have any thoughts for like how folks who are making these products like TaskClick shortwave, obviously tons and tons of others should think about the world that you're envisioning. Like, I I still am wondering if the right wave for me, even as I set all this stuff up, right, and get my goals instantiated and build up all the context. Should I be, like, going to that terminal and saying, go triage my inbox and tell me what I need to respond to and have the responses drafted that way?

Or should I do it in a product that was really built for email and have that product kind of call into the Nathan Oracle for what context or judgment assistance it needs at any given moment in time? Because I do feel like a lot of people you're a seasoned vet when, you know, a VIM guy, as you said. Right?

That's obviously a very minority profile. And I'm comfortable enough to go do command line stuff, but I would probably side more with the typical user who's like, wants a graphical interface, or at least is more comfortable with it most of the time.

Speaker 3

Yeah. That's why I have this maturity model thing to keep reminding myself what the actual goal is and to work backwards. I should not be on the terminal at all in my PIE system.

And I should not be in some I use superhuman, by the way. Is it? The that's my email client.

Mhmm. But I shouldn't be over there. What should happen is I say, what should I be looking at?

Who should I respond to? Is there anything important? I just speak those words and the things happen, whether in the short term, it pops up that client and I have to go interact with it there.

Or Kai is able to do it himself because he can control the clients. I think Gemini is definitely getting there very fast with turning on a bunch of Gemini features in Gmail. But to me, like, no, we should not be dealing with any of this kludge of even an email client is kludge.

You think about it, compared to minority report or the movie her. Right. You remember when you onboard in that operating system, you just say, Hey, what's going on?

Anything I should know about? And she's I just read your 940,000 emails. You got a new one from Sarah this morning.

That's the interface ultimately I think everyone is building towards. So I try to keep that in mind. I will say one other thing about you because you're asking about like product advice.

The ultimate product advice that I'm seeing, and I help, you know, companies with this all the time, especially in cybersecurity. I have this one piece of advice, which is if you are doing a cool product feature in a space like vulnerability management or threat intel or whatever, and it's pretty good, and you are competing against someone who is also pretty good, but they understand the customer and you don't. For example, this is a vulnerability management is a great example for this.

Do you know all the engineering teams? Do you know how they push code? Do you know what their repositories are?

Do you know how they're measured? Do you know all of those things and their ticketing system and their CICD pipelines? If you know that and your bone management program or your bone management solution is a little bit worse, maybe then someone else who doesn't have all that context, you are going to lose.

Like, you're going to lose to the company who has more context. So my expectation is that even somebody who seems like, oh, we just make a task list and it just puts out this little piece of context. Their entire drive, they will either not survive or they will move towards the model of, you know what?

It turns out we actually have to learn a lot about this person. We should have a pie for them. And they're everyone's gonna build this deep knowledge of the customer or the user.

And that is going to be what powers how good of outputs they can produce regardless of the product.

Speaker 2

Yeah. Okay.

Speaker 3

Another example of what you were saying where everyone's going the same place.

Speaker 2

Yeah. What's working in memory? I've been fascinated with memory systems for LLMs, agents, whatever you wanna call them, for a while.

And this is another area where I feel like everybody recognizes that there's something missing or that it could be better, but instincts are very different in terms of how to deal with that. So what have you tried? What is working?

Are you using any dedicated memory infrastructure companies to support your memory features? What do we need to know about memory?

Speaker 3

Yeah. I'm very much team file system. Yeah.

When the first version of pie came out, I don't know what sometime middle of last year, I came down firmly on the side of file system. File system is my memory. It is my storage.

It is my context management system. I do have an archive of all my writing back to 1999 that is like tens of like over 10,000 posts. That one is a rag.

So occasionally I have a rag, but I really dislike rag because I feel like it's just lossy and messed up. I prefer file system. I think it's the absolute best.

So I have underneath the dot claw directory in all caps is memory. And under memory, I have learning. I have signals.

I have all these different things that are pulling from the projects directory, the events JSONL file, which is every single transcript that's happening inside of the cloud code system. But on top of that, what I have is built on is this thing that relates to the algorithm I was talking about. It is constantly through the hook system determining how happy I am with responses.

And then the post hook is looking at what the current sentiment level is. I have histograms of like how happy I have been with the results coming from the PIE system. And so what that means is the system is designed to look at those signals, look at what I asked for, and look at what it produced, and then the sentiment and say, oh, he obviously wants to go more in this direction.

He wants to go more in that direction. I should do more of this and less of this. And this is all in service of the ratcheting up of the improvement of this overall algorithm.

The overall ability for an agentic system to take for any particular task or for a long term goal, the ability to move from current state to desired state. So I'm using the memory system to gather extremely granular stuff and all signals, but it's all the entire purpose is self improvement, recursive self improvement.

Speaker 2

And does that practically operate on just a runtime agentic search basis where Claude just decides what it wants to look into and pull stuff into context on its own? Or are you doing some sort of post background batch processing? I've also been quite interested at times in there's a I did one episode I did to it was on a system called Hippo Rag, which was taking inspiration from the hippocampus.

Multi step process where you'd have these whatever your corpus was, First you would go through and do entity recognition and deduplification and then create like a graph structure that would have the entities and then the documents in which they appeared. And that way you could rag into it anywhere in natural language but then see, that connects to these concepts, which connects to these other documents and expand out in a sort of network based way through the corpus as opposed to a sort of purely hierarchical approach to retrieving information. That gets pretty complicated, obviously, pretty quick, but it does feel like something like that might be needed relative to maybe this is also just my, like, lack of confidence in my own ability to organize myself and my thoughts well enough.

I certainly do recognize people who are quite different in this regard. But it's I feel like I need a sort of cross boundary layer that would probably have to be batch processed in the background to make these connections between all these various disparate things as opposed to being able to put each one in its proper place such that like Claude intuitively and correctly decides where to go just based on structure.

Speaker 3

Yeah. This to me is the whole advantage of the scaffolding and like being able to infinitely tweak the scaffolding according to first principles. So because I have the core skill, which is the bootstrap for the entire PIE system, because I have that laid out and it gets loaded, and it has all the context of what like what we're trying to do and everything.

It gets also the architecture of the system, including the memory system. Now, all this stuff that you're talking about doing with like scripts and stuff like that is the Clogcode hook system. The Clogcode hook system is extraordinary.

So I have, I think right now, 12 hooks that are active. And this is I've got a whole bunch for user prompt submit. So there's security checks in there.

There's sentiment analysis checks in there. It's actually routing throughout the PIE system according to what I'm trying to do based on the sentiment analysis, which is uses haiku. So I have a custom inference tool, which has three levels of inference, fast, standard, and smart, which is haiku sonnets and opus.

And so the entire system is using this to, like, self route. Now the memory system and all those sentiment analyses and all the artifacts of keep in mind, this is it's fully archiving. Cloud code does this naturally.

Every prompt I send, every tool use that it runs, every output of the tool use, this is all recorded. It's all there raw for us to analyze. So I am taking that and putting it inside of this memory structure, and I'm overlaying on top of it sentiment analysis.

This is all being done dynamically. I'm not seeing anything. It's all just handled automatically due to hooks.

So hooks are constantly adding the sentiment layer of how good the algorithm is doing, how good the PIE system is doing overall. So at any point in time, I could say, what upgrades have we made to the system? How have they gone?

How has our performance been going in the last month? And and pie will come back. Kai, in my case, will come back and say, yeah, it seems we tried this.

That didn't work. We uninstalled that. We went back and we went in another direction.

And currently we're doing this and you seem much happier with this. So this seems like a direction to go. Do you wanna do any more work on that?

Speaker 2

that's all just operating on raw logs. There's not like a summarization level or some sort of because that sounds like just like a ton of content for it to wade through.

Speaker 3

Oh, there's tons of summarization happening. Yeah. That's what the inference piece is.

So the memory system is dropping its own artifacts, which are summarized versions. So that and they are also creating indexes in JSONL, which can be read, like, instantly fast. No.

You couldn't go and parse like the entire thing all the time. That would be too intensive. Yeah.

This is stealing from a Stanford idea called reflections, where you get a whole bunch of context and you summarize it like, maybe in one line or one paragraph. This is from, like, the AI village originally. Right?

Yeah. That's right. Yeah.

I think about that a lot as well. I got a lot of inspiration from that. Yeah.

Summarizations into indexes, which can be parsed. And of course, they could always go look at the raw log if they want to, but they should be able to go off of the index. And then yeah.

That's all happening just with hooks, and hooks are happening anytime the system runs.

Speaker 2

In practice, when you see people take your system and modify it, how much are they modifying it? Are people like following in your footsteps relatively closely, or are they veering off in in all sorts of different directions?

Speaker 3

Yeah. I've not seen many modifications. It's more so population of the system.

Oh, someone just posted one yesterday to the discussion in on GitHub. Holy crap. It was I was like scrolling.

It was like 20 pages. It was like the most insane thing I've seen. Oh, I think the guy's name is Jim.

And maybe the agent's name is James. I can't remember. Something like that.

But anyway, he just he brought over so much context and so many things. And it it was just massively impressive. So it's just a matter of he knew exactly what he wanted.

This is what activates pie. He knew exactly what he wanted. He's been struggling with all these same pie problems of pie not existing, Claude code not existing in the past.

He's been sitting on all these things like I have for decades. He knew what he wanted. He knew what he wished he could do.

He saw pie, brought all this stuff over and now he's producing content, way more content. He can make products. So it's more so like activation of what was already there, but dormant rather than I I have seen some expansions of the system.

There's lots of feedback, pull requests and stuff where they're like, hey, could you add this? Could you tweak this or whatever? And so we're obviously trying to listen to those.

Speaker 2

How does it feel to you this is a bit of a weird question. We have obviously highly, plastic brains that can really surprise people in terms of just how adaptable they can be. And here I'm thinking like blind people seeing through a prosthetic that, like, zaps their tongue and they learn to interpret that as a visual signal.

Lysus Long there. Right? I think it was I'm not sure if I could say his name quite correctly, but Jeron Lanier.

Yeah. Hopefully, I'm saying that. He's done fascinating experiments with virtual appendages in VR and getting your brain to learn to control some pretense old tail or something like that, and you can actually learn to do it.

I'm wondering and then, of course, I'm also thinking Neuralink, right, is about to they start scaling up its customer base, and obviously their ambitions go way beyond treating paralyzed people and who knows what that's gonna look like in the future. Is there a feeling that you have of like this thing being a sort of literal extension of you where if it's turned off or you don't have access to it for a time, do you like begin to feel like something is missing? Another version of this, real simple one, but digital is the feeling of something being on your clipboard.

I know this is a guy I recently looked this up. It's a fairly known phenomenon. I've always felt for twenty years now, I've felt like I know when something is on my clipboard.

I sometimes don't know what it was anymore and I have to paste it to see what it was, but I know that there's like something there that part of my brain has developed or changed in some way shape or form to be tracking that very closely and it is a felt sense that there's something on the clipboard. So I wonder how this feels to you and if you can describe that, I'm this is like a way to try to get at what the end state would look like if I'm using this kind of thing. How should it feel to me?

How will I know that I'm like hitting pay dirt based on feeling how it feels to you right now?

Speaker 3

Yeah. Yeah. Totally.

I love that you brought this up. I think I was way back in the army in the nineties, and I came across this book called Getting Things Done by David Allen. And ever since then, let me reach into the pocket here.

I have index cards and index cards are like my way of capture. So the prime directive for David Allen is never let anything sit in your brain because it will hassle you and trouble you and cause like executive function problems because your brain will be like, hey, what about, hey, what about, hey, did you remember that thing? So I'm a massive clipboard person, not technically clipboard, but in the way that you said.

So in front of me, I've got different colored sticky notes. I have this system. I have my space pen, which is my favorite gift to friends.

And this is just what I travel with to make sure, and now I have this limitless pendant, which just got bought by Meta by the way. So I think I might switch off of that. But capturing what I'm thinking at the moment has been critically important to me for like over twenty years.

It just feels like massively important. I just recently created a reminders file inside of pie. So I could just say, hey, remind me to do this.

Remind me to do that. But honestly, the vast majority of that is I have 2,900 apple notes. So apple notes has been my main capture for a long time, unless I'm doodling or capturing ideas like visually, which is on the cards.

Now, again, going forward, I should not have to be doing any of this. I'm gonna keep my cards just for history reasons, but what should be happening is more like with her, Joaquin Phoenix. It's like, hey, make sure I don't forget this.

Hey, make sure I don't forget this. And agentic systems should be switching away from call and response to your reminder list is always there. Always ready for your DA to shoot you a prompt.

Hey, it's time. This would be a good time to do that. Hey, do you want to revisit some of your to dos?

I saw a really cool thing on X yesterday. It's like a little clock next to them and it's the daily agenda in analog form on this digital clock or whatever on their desk, but it clog code generated. Right?

So whatever they're doing, they must have their own PIE system and it's right there in physical form. So it's like crossing these two worlds, which I really like.

Speaker 2

How do you think about the triggers for the system. Obviously, you can ping it and then Yeah. Presumably, it can be pinged by any number of external or you can allow it to be pinged by any number of external events in the world.

And then there's the kind of background processing or if you want it to be proactive for you, is that like a daily job or an hourly job? What do think is the right balance between you go to it, it runs on a schedule, something triggers it from the rest of the world, or maybe some mysterious fourth thing. What's the right way to think about that balance?

Speaker 3

Yeah. That that's a wonderful question. They now have the ability to launch remote agents.

So you can actually send a task and it will run off in a GitHub infrastructure in their environment and then return results to you. The other thing I have, I'm a big Cloudflare person. So Cloudflare has the ability to create workers that can run different things on different scheduled timeframes.

So I have a whole bunch of my infrastructure is Cloudflare, and they can talk to each other via authentication and access each other. Right? I even have an infrastructure for running cloud code inside of a Docker, which agents can also talk to and schedule.

So all of this is in service of, again, going back to the, what I was talking about before. I should not have to think about any of this. I do right now because the text is not quite there, but when I want to make something like you're talking about, I literally say to to Kai, hey, look, I need you to not forget these things.

I need you to remind me these things on a regular basis or whatever. What are possibilities? And Kai will be like, yeah.

So listen right now, the whole trigger thing, like that's not super far along. I tell you what I could do. I could spin up a worker.

I could check every five minutes or every one minute against this set of goals. And I could ping you like, how would you like me to ping you? We could do the discord thing.

I could text you. I could send you an email. So we're starting to like creep towards this in a kludgy type of way.

But it's another example of everyone's going the same place. Right? Because everyone's talking about background agents right now, remote agents versus local ones.

Part of the PAI maturity model is in some of my friends are ahead of me on this. They're already calling in and accessing their terminal remotely. Me being a security person, I'm scared shitless about this.

So I haven't done it yet because I haven't found a perfect secure way to do it, but it is a huge problem that my system is a terminal inside a computer. Right? If you want to get to the future of her, you've got to that's gotta be with you all the time.

Right? So that's that's all stuff I'm thinking about and scheduled tasks, like you said, or logical triggers is even better. It's better than scheduled tasks.

Cause like one of the first things I talked about in that book in 2016 is just like proactive. That's a huge difference. Call and response.

That's one thing. It's really cool, but it's still too close to a chat bot in my mind. Right?

You're like ask a question, get an answer. Cool. Now you have to do something with it.

What should be happening is it understands your environment, the timing. Like right now, Kai should not be interrupting me with, hey, did you see this cool news story? Because it knows I'm in the middle of a conversation.

So small little movements all in these directions from multiple angles, I would say.

Speaker 2

So earlier you mentioned that your friend is, like, earning more bug bounties by doing something like this. Do you measure your own productivity in any similar way and how much boost do you think you've got? And then as this presumably continues to create more and more leverage, that seems to imply that, like, you'll have to have yourself in the loop with lower and lower frequency.

Right? You know, if in the limit of this sort of thing, there's you're only able to review so many things and make so many decisions. This is the gradual disempowerment people would be saying, hey, you're talking about it right now.

But if it's performing well enough, you'll be reviewing the things that matter and you won't be reviewing the things that don't. Where are we what can you measure about your own output today? And where are you in terms of like how much scope of action you give the system?

Does it ever send a response to an email? Does it ever send an email as you that you didn't review? Or do you like allow it to respond as itself without signing it as you, but like still try to move things forward without you actually being in that loop?

Would you allow it to spend money on your behalf without you signing off? Yes, you want to execute that transaction. Are there other frontiers of action that you're watching the line move on what you do and don't need to be looped in on?

Speaker 3

Yeah. Yeah. Great.

I would say that being naturally a little bit cautious, I would say the scaffolding is not there yet for a whole lot of trust in this regard. When I'm sitting here watching it, I've got a big part of my hook system is actually a whole bunch of defenses. Watching it, watching what the agents are doing, making sure it's not accessing certain files and directories, And that uses the cloud code underlying system.

It's got a whole bunch of cool permissions. I don't run dangerously skip permissions anymore. I used to.

I turned that off. So I've got a whole like security scaffold there for file system access and stuff like that. Then I have a whole bunch of prompt injection defenses because those are massively dangerous as well.

And I keep those layered. I just don't feel like the scaffolding is there yet to be like, hey, whatever. Here's my bank accounts.

Just run with it. I would say I'm okay with experiments. Okay.

Here's a separate bank account. It's only got a thousand dollars in it. Go crazy.

Like you've probably seen the vending machine benchmark. Yeah. Yeah.

Like cool. If there's bounds, if there's like blast radius control. Sure.

But when it comes to being able to send out emails and maybe my diary is sitting, I don't actually have my full diary or journal in the system yet, because this is one of the things I'm a little sensitive about, but like you get a, somebody sends me a link as, Hey, Kai should go read this. I send Kai to go read it. It's a prompt injection.

Pretty soon I, I just published my diary on LinkedIn. Right. That's possible.

Right. Much harder to do against me, but prompt injection is not like a super solvable thing. So I would say I.

Level of trust. I'm going to say, I don't know. There's no way to put a number on this, but I'm going to say like 60%.

And I think over the next couple of years, I'll probably get to 90%, but I'm still going to have I still think security also being ex military and, you know, just cybersecurity. I think in terms of threat models. Here's all the things that would super suck if they happened.

Just assume they happened. What could have stopped them? And a lot of that comes down to impact reduction in addition to probability reduction.

Speaker 2

It's fascinating to think that you're not a total maximalist on

Speaker 3

this stuff. I am. I'm a total maximalist on it, but it's just and I'm doing a lot of crazy sort of I do lots of crazy experiments.

I just have the blast radius limited quite a bit. Yeah.

Speaker 2

Yeah. Not a total Yoloist, I guess, maybe is the Yeah. Is maybe a better way to say it.

I've kept you a long time. I could go on longer, but I should probably get us wrapped up. And I gotta get I gotta get deeper into this is obviously the next big thing for me to do.

The one of the thing I wanted to touch on from your print PIE principles, and then maybe just give you a chance to touch on anything that we didn't touch on that you think I should know or anybody in the audience should know, but the last principle was permission to fail. And I thought that was quite interesting. It certainly brings to mind things like when anthropic gives Claude the option to end the conversation because it thinks it shouldn't be having this kind of conversation or to escalate something to the model well fairly, that anthropic, it brings the bad behaviors of deceptive alignment, etcetera, down a lot to give it that sort of escape valve.

So it sounds like you're doing something very similar there where you're saying, if you can't do this, don't gaslight me. Like, it's okay to fail, but just come back and tell me the truth. I think that's a really interesting fact that people should appreciate better about AI in general, and it's interesting that it's made your list of principles.

Interested to hear any more about that that you wanna share, and then maybe just anything else that you think people that I didn't touch on that you think people should not miss out on.

Speaker 3

Yeah. I'll talk about that real quick. I think that's a very tactical one that we just understand as being a weakness of LMs, More so the further back you go.

This is a huge problem in '23 where it would just make up stuff. Right? Cause it's trying to do the right thing.

So this is a very tactical thing basically saying it's okay if you don't have the right answer. It's okay if you can't get to ideal state. Feel free to tap out and just tell me the truth because I value the truth more than you trying to keep confabulating something.

So it absolutely does. It looks like from the studies, it does actually improve performance, especially in not hallucinating and being a sycophantic and all that sort of stuff. In terms of, I don't know, positive or other things to mention, I would just say that I've had this idea of slack in the rope for a very long type time.

So the idea is I feel like that as humans we talked about us not being unlocked. I feel as a species, we tend to feel like the way history has gone that way because of our innate human limitations. It's like this because that's the only way it can be.

We only have these medicines because we're at right we're right at the limit. All of science is pushing perfectly with full strength and this is the exact place and to go 1% more would take infinite energy. I don't think that's true.

And I think AI more and more is showing you that this is not true. And I am so bad at this because I'm also programmed. I'm constantly trying to break myself out of this of no.

Once we start asking the right questions and providing the right context, we're gonna be like, are you kidding me? You are at 1.7%.

And it's really easy to go to 63%. And we've seen this with AI models actually. Right?

For a long time, and I was arguing with some of my friends at these labs back in '23. They're like, yeah, whoever has the compute is gonna win. I'm like, aren't there like little tricks where they're like, hey, I wonder if we what if we just reverse the numbers and add them this way instead of that way?

Oh my god. 47% increase. How many more of those are like lying on the ground?

Just fruit ready to eat that it's just a matter of doing these combinations. How much research out there is partial? The medical research, this one trips me out.

It's like, many studies did grad students do? And they're like, oh, it turns out this molecule, if it encounters this part of a cell, it will produce this antibody. And this antibody will, by the way, kill all bad things.

Hey, listen, I gotta go take this job. I'll just leave this research paper here. And it's in some file somewhere or physically printed out somewhere and no one's looked at it.

But there are hundreds of thousands of these across decades. Right? And it's like going back to the security problem.

No one has the time or the eyes or the brains or the hands to actually go and look at this stuff. So I feel like the combination of these two concepts means we're nowhere near any limits of what we could do. There's just so much opportunity.

And when you start looking at things like everyone gets a tutor. Oh, here's a crazy one. Here's a crazy one.

What if we could not only change what we could pursue based on what we want. So eliminating the obstacles in front of what we want. That's cool.

That's what we've been talking about. What if we could change what we want? There's this whole concept in philosophy of there's what you want and there's what you want to want.

So it's very hard to be like, yeah, I just really wished I like celery. How are you gonna do that? Now, a drug comes out, GLP one or whatever the agonist, GLP one agonist, it literally makes you not want food.

Okay? What if I wanted to be more self disciplined? What if there was an unlock for making me 10 smarter, which I would love both of those.

Right? These, feel like we don't know. It's a it's an open question of which ones are easily slack in the rope fixable and which ones actually are physics that are stopping us.

But I think a lot more problems in the world are likely to be the former.

Speaker 2

I think that's probably a great place to end it. An aspirational note. I'm looking forward to digging in on this a lot more, and I really appreciate your walk through today and so many aspects of the positive vision for the future that you've shared.

Daniel Meisler, thank you for being part of the Cognitive Revolution.

Speaker 3

Thank you so much. I really appreciate it.

Speaker 1

If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.

ing. And thank you to everyone who listens for being part of the cognitive revolution.

Shared via Hopper