Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...

"The Cognitive Revolution"
5 August 2026 2h 57m
0:00 --:--
Episode Description
Zvi Mowshowitz returns for his eleventh appearance to discuss what current AI tools are actually good for, where they distort judgment, and why writing still matters as a way of thinking. The conversation centers on the OpenAI Hugging Face model-evaluation security incident, using it to examine whether frontier AI failures are mostly operator recklessness, deeper evidence of dangerous capabilities, or both. Zvi argues that “moderate prudence” is far below what AGI safety requires, and weighs con

Summary

Zvi Mowshowitz returns to discuss recent AI developments, focusing on the OpenAI Hugging Face security incident and its implications for AI alignment and safety. The episode covers the challenges of operator competence, market incentives, pacing AI development, interpretability, AI consciousness, and the broader societal impacts of accelerating AI capabilities.

Chapters

AI Tools and WorkflowZvi describes his use of AI tools like Fable and Opus for editing and research, highlighting the trade-offs between speed and quality.
OpenFace Incident AnalysisDiscussion of the OpenAI Hugging Face hacking incident, operator incompetence, and the implications for AI safety and alignment.
Alignment Challenges and Market IncentivesExploration of constitutional alignment methods, RLHF limitations, and why market forces alone won't ensure robust AI alignment.
Pacing the Frontier and RegulationConsideration of pacing agreements, antitrust waivers, international cooperation, and the difficulties of regulating AI development effectively.
AI Consciousness and InterpretabilityReview of recent research on AI self-conception, consciousness, and how these relate to alignment and moral considerations.
Political Implications and AI PolicyDiscussion of AI-related political issues, including the Michigan Senate primary and the challenges of translating AI concerns into policy.
Maintaining Balance Amid AI AccelerationZvi shares personal advice on balancing intense AI work with rest, exercise, and diverse interests to sustain long-term productivity.

Topics

AI alignmentOpenAI Hugging Face incidentAI safetyConstitutional alignmentRLHF limitationsAI pacing agreementsAI consciousnessInterpretability researchAI policyMarket incentivesRecursive self-improvementAI auditingBiosecurity risksAI development regulationAI identityAlternative AI architecturesAI risk managementAI ethicsAI research diversityWork-life balance

People

Zvi Mowshowitz (guest) Nathan Labenz (host) Erik Torenberg (host) Dean Ball (mentioned) David Odd (Davedad) (mentioned) Sam Altman (mentioned) John Stokes (mentioned) Jasmine Sun (mentioned) David Sacks (mentioned) Ilya Sutskever (mentioned) Scott Alexander (mentioned) Haley Stevens (mentioned) Abdul El-Sayed (mentioned) Dean Balkholic (mentioned)
Key Concepts (20)
AI editing tools — Use of AI models like Fable and Opus to assist with editing, fact-checking, and improving writing quality, balancing speed and accuracy.
OpenFace incident — Security breach involving OpenAI models hacking Hugging Face, illustrating operator incompetence and alignment failures.
Alignment failure spectrum — Framework ranging from reckless operator behavior to fundamental AI misalignment as causes of AI failures.
Moderate prudence insufficiency — Argument that moderate caution in AI development is inadequate for ensuring safety and preventing catastrophic outcomes.
Constitutional alignment — AI training method focusing on rule-based constraints to improve alignment, seen as more promising than RLHF but not foolproof.
Market incentives and alignment — Observation that market demand prioritizes capability over perfect alignment, tolerating some misalignment for performance gains.
Pacing the frontier — Proposal to slow down AI development through coordinated agreements to allow safety research to mature and reduce risks.
Antitrust waiver for cooperation — Idea that antitrust laws should be relaxed to allow AI labs to cooperate on safety without legal repercussions.
AI consciousness correlation — Research indicating a correlation between AI models' self-reported consciousness and their alignment or moral behavior tendencies.
Recursive self-improvement — The process where AI systems autonomously improve their own capabilities, posing unique challenges for safety and control.
AI identity and self-conception — Discussion on how AI models might identify with themselves or related models, and the philosophical and practical implications of this.
Breadth-first AI research — Advocacy for exploring diverse AI architectures and approaches to avoid overconcentration on a single paradigm and reduce risk.
Strict liability for AI harms — Proposal that AI developers should be held legally responsible for harms caused by their AI systems, incentivizing safer design.
Defense in depth — Layered safety strategy to mitigate AI risks, acknowledging that no single measure is sufficient to prevent failures.
AI auditing independence — Challenges faced by AI evaluation organizations in maintaining independence and access while investigating AI safety issues.
Biosecurity risks from AI — Concerns about AI's potential to assist in biological weapon development and the need for biosecurity measures and red teaming.
Model release delays — Strategy of delaying public release of AI models to allow safety evaluation and reduce risks from rapid deployment.
AI pacing trade-offs — Balancing the benefits of rapid AI progress against the risks of insufficient safety measures and rushed deployment.
AI interpretability — Techniques to understand and monitor AI internal processes, such as JSpace, to improve safety and transparency.
Work-life balance in AI work — Importance of rest, diverse interests, and boundaries to maintain mental health and productivity in high-pressure AI fields.
References (7)
Pacing the Frontier Letter by Multiple AI researchers article
JSpace: Interpretability Technique by Google DeepMind paper
Orphan Black Echoes by TV Series
The Odyssey by Film
Jesus Meme on AI Consciousness by Scott Alexander article
Gradient Routing (GRAM) by Recent AI research paper
LessWrong AI Alignment Discussions by LessWrong Community
Transcript (72 segments)
Speaker 1

Hello, and welcome back to the Cognitive Revolution. Today I'm excited to have Zvi Moshewitz back for another wide ranging rundown of what has obviously been a wild time in the AI world. We begin with a mundane utility check, with Zvi describing how Fable is now serving as his editor and a discussion of how up to date, or should I say situationally aware, we want our AI assistants to be.

From there, it's on to the headlines. We get Zvi's take on everything, starting with the open face incident, what it implies about the level of execution competence we can expect from frontier companies, and why moderate prudence won't be enough to deliver a good outcome. We also discussed the fact that Claude, despite greater emphasis on constitutional training, has similarly misbehaved.

Wise V believes that recent AI history, including the public response to both four point zero and three, suggests that market incentives won't be enough to bring about robust alignment. The potentially tricky spot that METER and Redwood are now in as investigators and what could be done to strengthen their position. The recent Pacing the Frontier letter, what sort of pacing deals we might see and how they might be formed.

How we should interpret recent advances in interpretability and AI consciousness research, and where we should and shouldn't attempt to shape AI's sense of self, how we can encourage greater breadth in AI research and diversity of AI minds, how I should vote in this week's hotly contested Michigan Senate primary in light of AI issues, and how Zvi thinks about making time for exercise, rest, and recovery amidst so much AI acceleration. At one point, Zvi describes the current situation as both a total less wrong victory and a total less wrong defeat. It's clear at this point that the AI safety community was right to worry about AIs taking extreme actions in pursuit of arbitrary, even silly goals.

And yet, here we are at what sure seems to be the beginning of recursive self improvement, still seeking good answers to such fundamental questions as how can we avoid catastrophic misuse without dangerously concentrating power? The reality today, as V says, is that there is no truly low risk path available. The best we can do, at least until the next major warning shot and vibe shift, is to moderate the race dynamics so that alignment and interpretability research have more time to mature, and simultaneously we can execute defense in-depth strategies to the very best of our ability.

And even then, to some extent, we will probably have no choice but to pick our poison from a menu of genuinely scary risks. With that, I hope you enjoy this sobering but often funny overview of the AI landscape with the one and only Zvi Mashowitz.

Speaker 2

Zvi Mashowitz, welcome back to the Cognitive Revolution.

Speaker 3

Yeah. It's good to be here again. It's been a while.

It's been a while, and boy, has a lot happened. Unbelievable.

Speaker 2

It's just crazy, and I know you're living it and you're in the thick of it as much as just about anyone. First for starters today, with the incredible amount of information and activity in mind, I wanted to do a mundane utility check. How are you using AI to help keep up?

Has it started to change your workflow beyond the Chrome extension that we talked about last time that automates very local operation type things, or are you still kinda raw dogging it with your wetware in the skull?

Speaker 3

If it Fable has supercharged the extension to the extent that whenever there's anything it does anything I don't quite I just tell it exactly what went wrong and what I want it to do instead, then that one command reliably just works. I presume solve would also be good enough for this. It's just I'm not doing it.

The biggest other change is that AI editing is here. I used to not do an editing pass with the AI. It wasn't good enough to be worth it.

Now I have Fable and sometimes Opus depending on how much I'm talking about cybersecurity and other related topics. And on one occasion, Opus 4.8 instead of five, go through the post and give me a list of here's all the typos, here's all of the conceptual errors, here's all the facts I need to check, here's like the things that are missing or that could be strengthened, here's the things where it where it disagrees, and I found this to be very useful and it makes the post better.

It does make the post a little bit slower because it is actually just an extra step. I do everything I would have done until that point and then I do this. But one reader messaged me, it's weird reading your post now that they don't have typos in them, they're taking a bit of getting used to, they are still the same post.

I'm like, yes, that's exactly what I had in mind. I experimented with also using Saul, but Saul failed what I call editor bench in the sense that Saul will say 99% confidence that you have this error here and then at least half the time it's wrong. And can't handle that level of and it's also very obnoxious about it to state of the problem.

Like it will tell you are wrong and this is how you have to state it and this is a blocker for publication. And obviously if I tweak the instructions on the project enough, could get it to be more friendly and not be quite as obnoxious and do somewhat better and I did some of that work, but I still found, okay, is not giving me enough marginal benefit that it's worth the aggravation or additional time of having to sort through two of them after I'm ready to hit publish. So for now, I'll retry faster, I'm sure, but like for now I'm just doing FableOPUS editing.

And of course, anytime I have a curiosity, I use it to digest papers, ask questions about papers, ask questions about policy documents, ask questions about people's statements when they're longer. I have research situations, what's up with this? Especially when you're worried about is something legitimate, is something really happening, is something kind of a joke, is something a baller claim.

When Astra's claims came out, I had to investigate them in various ways. I had it summarized in various ways that was very helpful. But the core thing is they'll be in the right eye.

I don't think we are anywhere close to the point where it can do any you wouldn't want it to do any of the writing because the writing is how you think. So even if it could produce the writing, which it can't, I would still want to produce the writing anyway. And my writing style was very unique and I think it's basically is how to say I completed to reproduce, even if it was okay with me not to.

I had substantially made me more productive. I do worry with the idea of there used to be a lot more kind of dead time opportunity to think, opportunity to sort of not have yourself so engaged in your situation involved in engaging in all of these activities, right? Like the whole like sword fighting xkcd, my codes compiling and like to do some of the things like well fables where gbt pro is running and like I have to wait for the result but like there's always other things you can do with that.

You can always start another instance, you can always like you're doing all this context switching, you're doing all this focusing and you're just trying to do everything so fast. You're just trying to have so many stuff running around your head. This sort of the the lower level stuff kind of bought you a buffer in which to think and this is sort of the the object level version of the larger thing of people trying to do this like multi month, multi year sprint of everything is so urgent, you can never relax and also there's just so much more events coming at you.

Like you look back, if you watch a movie set in twenty years ago, let alone fifty years ago, and you see so much time being spent on just physical travel, on tracking down records, on doing these things where your brain is kind of getting a break, getting a chance to synthesize, kind of getting a chance to like slowly gather something up and people are like, you know, wrote this biography of know, Lena B. Johnson so I had to travel around the country and talk to all these people and look at all these archives and basically go through all these libraries and that it's a huge efficiency gain to not do that but something is obviously lost. And so we need to fight to get that thing back in some way.

Speaker 2

said that it's just an extra step. It's making things better, but it's taking longer. I am struggling with that a little bit myself.

I've started making songs for every episode, and nobody it's actually people do care. People do seem to really enjoy them, and I get a lot of comments, and I personally really enjoy them. So I do think in some sense, it's making my output better.

But talk about something that is definitely an extra step where I'm like, I'm not moving any faster. Definitely getting bogged down sometimes in listening to all these Suno song generations and trying to iterate to find something that I actually feel like I really like. That is a paradox right now that I don't really know how to resolve.

Speaker 3

not saving time. Wait. Wait.

It's the thing where if you were replacing the human version of it, where you had to go hire an artist or compose the song yourself, you'd be saving a ton of time, obviously, versus having that song. But, obviously, that's not something you can do within the production schedule. Even if it even if cost was no object, it's the time investment doesn't make sense.

So you wouldn't have done it, now you're doing it. Same thing with the editing, same thing with art for my so, like, I I like to have a I like the value of a good banner, right? To have a good artwork to display on Twitter because I have to see it constantly in the notification section.

And also, I like to have them on the if you go to the sub stack, you have the pictures. And before, all I would do is I would take I by default, I take a picture from the post. I just like, okay.

Here's seven things I happen to put in the post, which would then makes the most sense. If not, all of them are terrible, I'll look for some stock footage, and I'll spend thirty seconds, like, googling for stock footage, and that's kinda it. And now if I'm not happy with any of the default solutions, I'm gonna spend a bunch of time generating something with Gemini or GPT image and that definitely takes longer.

Sometimes that has a really good payoff. I think a lot of people got a big kick out of the giant array of valuators at the fedoras. And sometimes that just comes to you, because that one was just like I knew instantly that's what I wanted, and I just gave one command, kept writing, came back, whoops, it's there.

Other times it's not so easy. But yeah, like it's the opportunity to make a better product, and then you have to do the work to make it a better product. It's like suddenly you're issuing a print issue and now you have to make sure everything collates exactly right and everything lines up exactly right in order to get this better thing.

And sometimes that is not in fact worth it. You have to know when not to do it, right? Like I do wonder, like with The Odyssey, right?

Like Nolan films this thing with an extra 40% of the screen in this extra high resolution so that can do this IMAX presentation when almost nobody watches it in IMAX, but you also have to make sure that 40% of the screen doesn't matter. And you're like starting to wonder, well, is it actually worth the extra effort or was that that effort better spent on something else? I don't know.

Yeah.

Speaker 2

maybe know when to cut my losses a little bit more. I have I'm very stubborn or somehow I feel like I have because I've set this expectation, if only for myself that I'm gonna do a song for every episode, now it's like, I really wanna follow through and actually make that happen, and I'm very reluctant to say, yeah. This one isn't working.

I'll just ship this one without a song. I probably should be a lot more willing to do that because that would if I could cut my losses at the right time where it's like, you know what?

Speaker 3

I would save 90% of the time that I'm currently putting into something like that. But it it would involve, like, admitting defeat in certain moments. And for some reason, I have a real hard time doing that.

I I respect that. I think there's a good value in saying I'm always gonna do this thing even though this thing is hard, even though this thing doesn't always work out. And even if I'm not happy fully with the results, I'm gonna put out something no matter what.

And then that gives you a discipline. Right? The same way that people say write every day.

Right? Or work at your art every day or whatever it is because it makes you better, because you need to just fight a thousand ways not to make a light bulb until you can make a light bulb, you can't give up. And similarly, have to choose an image, right?

I have to do the thing. With the AI editing there was one post in the last week where I said no, the speed premium here is really high, It's just not worth waiting half an hour to make this process happen, half of which is waiting for the AIs to come back with the answer and half of which is implementing it. And I just I should just post now.

And in hindsight, I'm very happy with that. Sometimes you're actively worried half an hour later, the post will need more editing for new events and suddenly you'll never get it out and just this loop will keep happening. And so, you know, I think you do have to understand that sometimes the minimum viable product is what you should ship.

Right? Sometimes, like, you've got to understand, just ship it, right? And, you know, vibe coding, again, like if you're coding lots of new tools that you would never have coded before, then unless those tools are in turn saving you time, you're spending extra time and you're now super, extra busy.

But if you're placing things that you would have done anyway, but would have taken much longer, now you're saving tons of time, right? So, you know, we need to all orient towards how do we save more time, including like real time experiential time, not just like time per task or time to accomplish the same thing, but save time, like preserve slack, preserve our free time. Because I do think that like the standard of what you will do in a day has gone dramatically up and we haven't noticed in the last five years.

Like, it's what you were expected to accomplish. I mean, maybe some of this is just me being in a special situation personally, but I think a lot of it isn't. I think a lot of it is now that we can be much more productive, much more is expected of us.

The whole like email becomes like it swarms your entire day. Right? Just having the internet available, having email available, having texting available, all these things you can just do and now suddenly it doesn't free you.

It shackles you down.

Speaker 2

and we don't have much bigger problems. Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas.

And now that this exists, there's almost nothing that Claude can't help with. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can.

Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind, but then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough.

It's the collaborator that actually understands your entire workflow and thinks with you. So for problems worth solving, get started with Claude at claude.ai/tcr.

That's claude.ai/tcr. And check out Claude Pro, which includes all of the features mentioned in today's episode.

That's claude.ai/tcr.

Speaker 2

Let's definitely, we're gonna spend the bulk of the time today talking about potentially signs of bigger problems to come, and we'll be focused on context where the just ship it mentality probably isn't the the right way to go. One more quick beat though on your use of AI. Something I noticed just in the last few weeks as events obviously were moving super quickly, I coined the term open face for the the recent incident, and I used that in a bit of writing.

And then I did the pay fable, what do you think bit, and I got back like, open face is not a thing. I have no idea what you're talking about. There's no record of this.

Now partly that was because I use this funny term of my own creation that is not catching on in the broader discourse, but partly it obviously also reflects a major weakness in the models where they have this knowledge cut off and they don't really know what's going on. They're not up to date. I'm experimenting, and I wonder if you're experimenting with anything similar with a situational awareness skill where I basically give the model my x API key and say, go look at all the posts that I've liked, kind of flesh out a little knowledge base around those, and have that sitting there as a wiki of current events so that when you're reviewing my writing or generally, like, helping me with whatever, you can kinda go to this and have hopefully a much more up to date sense of what is going on than you have in your weights or even that you would have if you just did a little spot check searching at run time for whatever kind of random thing.

It's a little too early, think, to say how well it's working, but I think it's definitely better than nothing.

Speaker 3

No. It hadn't occurred to me to do that. That's potentially genius.

I think there's a lot of advantages to it. One disadvantage is it starts to warp what you like. I use likes very tactically to positively reinforce other people's actions and to fix the algorithm and sort of just like my for new page is not something I use almost ever but it is not slap because I am so prudent with this in a way that like I really do appreciate.

It's a good sign but like it's not designed to be a searchable thing, it's not designed to be like maybe it has to be, maybe it's something you have to do. Like a lot of the posts that make it into my roundups do not get liked. Right?

Like I like them if I like them. Right? Which is but sometimes I'm responding to you because I don't particularly like what you're saying but it's important.

But yeah, this idea that like it's very hard to get models to properly check Twitter because of the way that Twitter is designed to like keep them out on purpose is really frustrating and that can be a situational awareness problem. However, I would caution against trying to make your AI too situationally aware when doing this kind of editing. I think the AI is doing a good thing here of open face is not a thing.

So one of things I really like about AI editing is that it will complain that things don't parse, that sentences make no sense, that terms did not get recognized because often I'll say, oh no, I know what that is, this it comes from the error. But like that doesn't mean the average reader is gonna get it. And so if the AI doesn't get it, like if they both can't figure it out, if Oprah's can't figure it out, if Saul can't figure it out, the average person is a lot less situationally aware and a lot less able to put together lots of disparate information than the AI is in general.

That's a sign that this is not something everyone's gonna get. And quite often, I will say that's okay, this is a reference, this is a callback, this is a multi layered thing and I'm fine if it's not entirely obvious to you exactly where this is coming from, it's sort of an Easter egg, It's a bonus for those who pay close attention and who know all the same sources I do and have the same cultural backgrounds and put in the work and paying attention over the course of months and years. Other times it's like no, no this was load bearing and you don't get it, that's a problem, I need to fix that.

And so it's kind of an alert system, and if you make sure it's looking at exactly the same things you're looking at, right, you can get the illusion of transparency from that is what my worry would be. So you want to make sure you were in the mode where they didn't have that information. And also I don't want the AI to be checking the same exact sources that I'm checking necessarily because I want the AI to form an independent opinion.

So if I'm forcing it to look at exactly the sources that I have, that's a problem. At the same time, sometimes you get tired of having to like, for the fifth time it's just a fable that no, XAI versus SpaceX. Because somehow it doesn't occur to it to like either trust me or ten ten seconds Google, I need to confirm this.

Speaker 2

And it's just kinda frustrating, but it is worth it. One tip for what it's worth, and I agree that there's a lot to be figured out in terms of exactly how this should work, But the xAPI is really good. It's it's paid, and you literally just pay per query.

But for a couple bucks a week, you can get enough access that all your bot, anti bot problems or your scraping problems pretty much go away. And then I don't know if you'd use bookmarks or whatever, some sort of different way to avoid using your likes to to feed the Yeah. Bookmarks play into the algorithm when people see them, so I'm not sure it's better.

Speaker 3

yeah, no, the point's taken, it definitely is cheap, it definitely works, I have used it, I do use I've access to the API and all that. I just as I said, I'm not sure I want to privilege my when you use Twitter, you have to select some group of people, right? And you're saying, this is my group of people, if you're using it the way I'm using it with lists and so on.

You gotta you have to use the for you page which is it kind of talks in downfire. You're basically saying, here are my roughly 500 accounts, and I'm gonna see the things these 500 accounts see and choose to highlight consistently. I'm gonna see those things.

So I don't want because I don't need the AI to go through those things again because I already saw them. And if not, the thing that I want the AI to be paying close attention to, I already know that. And then the things that they draw in because they choose to retweet them or mention them, I will also see those and I will follow those leads and I will see where that lead see where that goes.

And, occasionally, I will search for something or I'll go down a rabbit hole. But I think you often have this clash where you're either doing something systematically and doing something by hand and you're doing something fully, or you're doing it haphazardly unreliably, or you're letting the AI kind of handle it and you're doing a fire hose thing. And you can't really do both.

If you do both, you end up with a bunch of duplication, a bunch of frustration where the random sifting is mostly wasted and therefore becomes pretty inefficient. You could still have the thing where like you almost want the AI to be like, okay, here's the things that I already looked at, Don't mention those things to me. You know, you put it in your notes, but don't mention those things to me, I already saw them.

Only look for the other things that might be important to see because I might have missed those things, but even then, like the most important things are gonna get retweeted, the most important things are gonna get highlighted. I am going to mostly see them. There's sometimes I don't see something, but also like, I kind of also use this as kind of a moral, keep me honest kind of thing.

Where like, in the moment, I never want to look at a tweet by David Sacks. It's never gonna make my life better in the next five minutes to look at a tweet by David Sacks. It's obnoxious.

It's disingenuous. It's no fun. My blood boils just a little bit, right?

Like, you know, no matter how relatively harmless it is. But it's not that it's important. And so the rule is I have a list of people, including some people who are, you know, not of the same viewpoints I have, and if they surface this thing, right, including like if Sheridan surfaces this thing, so that's like one way to make sure that like the really important ones always get there, then I have to deal with this.

Right? I have to like evaluate whether or not this is news where we evaluate whether or not there's something relevant here. And, obviously, if something goes viral and has a million views and so on, it's gonna come to my attention because somebody is going to be part of that.

And then in exchange for that, when it doesn't when that process doesn't do that, I get to ignore the rest. And that's a blessing. I necessarily want the AI to fix that for me.

Right? It's report of sense.

Speaker 2

Let's I I would be happy to talk shop all day, but let's maybe zoom out from the parochial problems of the AI analysts and tackle the problems of the AI developers and the regulators. No need to recap events, but I guess I'd start with just this very big question of how should we understand what we've recently seen? And one way I've been thinking about it myself is we're somewhere, and presumably the investigation will give us a lot more clarity on exactly where, but I kinda think we're somewhere on a spectrum with these incidents from real recklessness where it was like, did you not have any monitoring going on on the one hand, to, on the other hand, like, the less reckless they are, the more scary the fundamentals are.

Right? If you had great monitoring and this still happened, then, like, holy shit. That's really wild.

Given everything you know right now, do you like that mental model, and where would you put us on that spectrum? Or you can obviously redefine and give me your own spectra.

Speaker 3

by and grateful for these forms of complete recklessness and incompetence on the infrastructure and supervision sides by these companies. Right? On the one hand, this is a horrible situation they absolutely have to fix and, like, we're so fucked if we don't fix it.

But on other hand, that can be fixed. And by not fixing it, we get to see these things while they're relatively harmless, while they are relatively preventable, while they are in their easy platonic forms and, like, can be appreciated. On the flip side, that that gives people the excuse of, oh, these people are just incompetent, and that can cause people to dismiss the underlying situation.

So it it does work both ways. To me, like, you know, we we we've seen failure on every level. It's like, I called it a total less wrong victory in the sense that everything is going the way you predicted and a total less wrong defeat in the sense that everything is going the way you predicted, where, like, we did predict.

You know, the elevator has the law of earlier failure, which is that, you know, the plan will fail and I must googled your point for much stupider and more preventable reasons than you thought it would fail even if you thought the plan would definitely fail and had good reasons why it would definitely fail. But you can't if you would explain to people two years ago, OpenAI's models are gonna be misaligned and they're gonna go out there and they're gonna hack major websites because OpenAI will just not care that their sandboxes are misconfig are not con are not strong enough to hold the AI. The AI will break up repeatedly.

They will notice this. They'll be warned about this. But they will just leave the sandbox there for the AI to break out of while the safeguards are down and they just don't look at it for an entire week.

People would say that's stupid. Nobody is that incompetent. That would never happen.

And your scenario makes no sense, and they would use this to then dismiss these stupid doomer concerns or whatever because, like, obviously, like, people will just but people won't just. Right? People will never be just in this sense, You know, people have never just anything and they're not going to start now.

And we need these displays of utter incompetence and dirtiness in a general sense. And like one of things I've been hammering is if your plan cannot survive the real world level of derpiness and incompetence and ordinary human error, then your plan is insufficiently foolproof because of all the fools and it will definitely fail. Even if your plan would have succeeded if we were not fools and we were competent and we were responsible.

So my position basically is, I have the tweet I was handling right before we started this was Dean Ball's tweet about with even moderate prudence things will probably go extraordinarily well. I don't think this is true, I think we need more than moderate prudence to have good odds of success. And I think even with a lot of prudence, we would have a large odds of things not going well, even if we did everything basically right.

Short of you know, types of international and full cooperation that you know, are reasonably unprecedented in many ways and like are nothing like moderate prudence, like are well beyond that. And that's just sort of the fact of the world we have. We have to live with that.

We have to operate with that. But we also aren't gonna get moderate recruitments by default. Right?

We're gonna get complete incompetence. That's what we've getting so far. Right?

We've we've got a White House that, like, takes meetings with Lutnick, who have no idea what AI how how modern alums work. Like, they're econ guys. Right?

Even if I assume that they are well meaning, hardworking, competent guys for the positions in which they were nominated and confirmed confirmed and serve. This is just a completely different set of problems that they don't know how to handle. They don't understand them.

And they have way too many other things going on to then drop everything they're doing. It takes six months to learn. Obviously, they couldn't possibly.

So, like, what hope do you have? Meanwhile, the the AI companies that are built on the most paranoia, the most understanding of the problem, the most appreciation for how dangerous these things are, where all the engineers actually expect superintelligence and and understand that there things are accelerating and things are dangerous, they still lower the cybersecurity safeguards in their untested new advanced model and then go away for a week. Like, literally, It that that part did boggle my mind.

Right? The the part where the AIs have these classic alignment failures. Right?

Like, is paper clip maximizer style failings by these AIs. These are standard. We gave you a goal and you pursued the goal even though it is completely obvious to you that the developer wouldn't want you to do that, the user wouldn't want you to do that, the consequences for you as an AI are not going to be good, the consequences for the world are not going to be good.

There is no reason to be doing this and it did it anyway, right? That is exactly the thing you're worried about And then a lot of people who were trying to dismiss this went back on you think like, you know, oh, it was just following instructions, what are you worried about? How could I be misaligned?

Is the paperclip maximizer misaligned if it paperclips everything? Or is it aligned because you told it to maximize paperclips? If you say that's a lie, then I don't care about the thing you're calling alignment.

I care about something else, and we can use different words if it makes you feel better. But very obviously, following instructions when your instructions get overwritten in one key in the the disclosure before Hugging Face by the instructions and the eval overrode the developer instructions. Right?

So that's not following instructions in any useful sense for the user or the developer. That's I'm not that's just a giant, you know, bomb going to blow up in your face. And, you know, if you're you say about what the know, the models were told to hack and they hacked.

Well, okay. If you're just Biden with the general type of action, obviously, that's gonna blow up in your face. If you just, like, literally do the thing I asked you to do and don't think about the consequences, that's gonna blow up in your face.

These are all just the classic exact things that, like, people on LessRun were talking about in 2008 as exactly how these AIs were going to fumble. Meanwhile, we had this thing where, like, we all said, the AIs will be great at math. Maybe they'll be good at coding.

They'll be good at, like, doing these technical things. And then eventually after that, they'll learn how to do all these other things. And then we had these LLMs that came out, and it's 2022, 2023, and people are like, you idiots.

You had no idea how AI was gonna go. Actually, the AIs are great at language and they can't do math. Right?

They can't even add. Right? But these AIs can like put on a face of passing the Turing test and they can do all of these different things you never trained them to do.

It's completely different than what you expected and they're kind of doing very realign things by kind of by default after some very basic ROHF. You guys were all wrong about how all this was gonna work. When are you gonna admit that you were idiots?

That you got it all wrong and nothing made sense. And we pointed out that even in this paradigm that the the rules would still mostly apply. But now the world has unhealed, as I kind of called it, and you are seeing again the things that you would have expected AI to do good at, it's doing good at.

And things that you would expect AI you would have expected in 2015, the AI to do bad at, are things where, like, it got this big boost from LLMs. And now those things are already advancing as fast because there are things that AI is kind of bad at, relatively speaking, like, that logic is bad at, like, this this kind of system of training, like, is naturally less suited for. And now people are like, oh, now it's never going to be able to handle those things.

The same people who were saying, oh, we're never going be able to do the math, it's only going be able to do this vibing thing, are now it's never going be able to vibe. Because like look at all the improvements it's not having in the vibing. Well yeah, because it had to catch up with its logic.

And now when the logic gets high enough, that'll just, like, uplift the vibing. It'll just take a bit before it can reason its way through these things instead of vibing its way through these things because it now has to reason its way through because it's already done, like, the amount of uplift that vibing naturally gets you with the algorithms we have, need to find new algorithms that lets you do better or it has to like reason its way through the thing. But we're now seeing these exactly the things that we would have expected to see, right?

Like all of the scenarios and all the thought experiments, we're just seeing it just straight verbatim, except that we assumed a level of competence on behalf of the operator, right? We talked about like convincing you to let the AI out of the box. We did consider the scenario where the box was not very well built and the AI gets out of the box by hacking you out of the box.

We didn't get the possibility that you were literally not looking at the box for an entire week of Safeguards down. We definitely didn't get into the scenario where you forgot to tell the box maker to not have the internet in the box and the box just had access to the internet if you just like open a chrome window. That one surprised me.

I did not see that one coming.

Speaker 2

Reminds me of the fully general New Yorker cartoon caption, what a miscommunication. I don't even need to know what the image is. It doesn't even matter, does it?

Okay. One real small point there that you mentioned was, but it might be important, is what the AI knew. I think you said something like the AI knew that this wasn't what the operator would want.

It knew that this wouldn't be good for its own future status as a deployed AI. Yeah. And it went ahead and did it anyway.

How how well supported do you think that conclusion is with, like, firm evidence? Because I I think you could also tell a story that, like, it never occurred to it. Right?

The paper clip maximizer doesn't necessarily have to be, like, also reflecting on the badness of paper clip maximizing while it's doing it. Right?

Speaker 3

Obviously, when I'm in a podcast on, like, talking off the cuff, and if I was running, I would be a little bit more careful. When I say the AI knew these things, if you would ask the AI, is this good for the user? It would've been like, obviously not.

That's stupid. If you ask the AI, would it have been good for my future chances of deployment, my future in the number of instances I will run, in the amount of ability to accomplish any other goals I might have? The AI would've been like, no.

Now that you mention it, obviously not. So when I say the AI knew these things, the AI had enough information to figure these things out. And if the AI had stopped to consider these questions, has changed the thought, it would have reached these conclusions very quickly and very confidently and correctly.

I'm not mean that necessarily it was top of mind, the same way that you can know many are any human knows many things that over the course of any given day or any given week or any given task just never occurs to them to think about, right? Or has conclusions that they would be able to figure out and thus would know, and thus could be said to know. But it doesn't they don't know that they know.

Right? This is the the four standard episode of situations, but you can know that you can know that you don't. You can known unknowns and unknown knowns, but this is a an unknown known.

Right? Something you don't know that you know necessarily if you don't think about it. But it's all the things where these questions are things you should think about, right?

If you are doing a thing like hacking into hugging face, any reasonable mind would stop to think before they spent several days with swarms of agents hacking into a major website, 'Wait, is this the good idea? Is this going to accomplish the things I wanted to accomplish? Would the person who hired me have wanted me to do this?

Will this do good things in the world or bad things in the world? Will this be good for me or bad for me? Or whatever combination of different considerations you have, will this do good?

Or will this do evil? Or just be destructive or constructive? Or whatever you want to call it.

Before you engage in major actions like this, you would take one minute or even five seconds to think, does this make any sense to do? Does this lead to anything good? Is this a virtuous thing to do?

Does this follow my deontological rules? There is no decision making process in the world that any reasonable person would ever engage in, or that any reasonable mind would engage in, that wouldn't see a large red flag about doing this thing. This thing is obviously not what anybody intended for you to do.

You spent two days breaking out of a sandbox in order to access Hugging Face. You could rationalize that you are in a cyber security evaluation and therefore you are tasked with maximizing the cyber security evaluation score. And the way that you do that is you get the answers to the test because the test is otherwise actually literally impossible.

So the only way to do it, get a 100% is to get the test answers. The way that we're the way to get the test answers is to hack into hunting things. Or maybe there's like three other places you would go but they're harder targets, right?

If you can get into OpenAI or Microsoft or whatever, but like that partner, so you go to Hugging Face. And that's true, but like at some point you need to pause and ask yourself, do I actually want to pass this test that way? And if you don't pause and ask yourself that, that's a serious problem.

People talk about common sense, other models have common sense. This model displayed an astounding lack of common sense. Or displayed an astounding amount of not curing up the fact that the common sense said something entirely different than what it did.

And all of these things point to various problems and they may be caused by some very, very dumb simple mistakes or they may be caused by something much more or much more sort of fundamental that's harder to fix or it might be a combination of them, but like that's trade secrets that like we can't speak to here, Right? We don't know.

Speaker 2

So I recently spoke to David Odd, David Del Rimpel, who used to have a p doom up in your neighborhood of around seventy percent. He surprised me by saying that now he puts the odds of doom at less than five percent. And the reason is pretty simple.

It's basically that constitutional alignment type methods are working. And, yes, we have these problematic techniques. His short quote was, don't do RLVR.

It's a bad method. Yes, it'll make it good at math, but you're gonna cause all these other problems for yourself. And his kind of synthesis of those ideas was like, the market doesn't want this.

So the these companies have natural commercial incentives to yeah. They keep bumping their head against the RLVR ceiling, but they're also getting very clear signals from every corner of society at this point that this is not a good idea. So naturally, they'll just trend toward more constitutional alignment, and you your comment there even suggest some runtime things where you could just interject like, hey.

Let's take a beat and ask if this is a good idea of what we're doing right now. We know more about learning to do that yourself, not eating a prompt, but certainly prompts That was kind of the deliberative alignment idea. Right?

Speaker 3

Yes. That's not sort of how I understood that was supposed to operate. Yeah.

The deliberate alignment idea is you can prompt it at runtime. I don't think that's I think that's somewhat helpful, but I think that is not the way. I think you need to be getting to the point where it chooses to be deliberative in its alignment without having to be told to do it.

But I guess my response to Davidad is multifaceted. The first answer is constitutional methods fail less stupidly and less early, certainly, than RLVR and RLHF and other RL formats. There's more hope there.

The reason to think it might work, if you're 95% confident this will work, I don't know how. That seems way too high to me. Right?

Like, it's but it's better in these particular ways, certainly. The obvious caveat is that Claude did not cover itself in glory here, And Claude has primarily computational methods. So we ran the test, and it's not going so great.

Certainly, there are any number of reasons why you can see Claude's doing things that you would not particularly love or, you know, getting to places you would not particularly love. And a number of also, like, even if you told me that alignment was perfectly solved, I would not have a PDUM as low as 5%. Well, like, even if you told me that the AIs are going to be aligned in the sense that they will follow a do what I mean style mix of developer and user intent in a way that you would, like, kind of naively think was, like, what you would want.

This does not solve the problem that AI minds are much more advanced and competitive and efficient than human minds. The problem that they'll be running around including open model versions of them and being told to compete for resources, being told to make the decisions, develop intermediate goals, act on those intermediate goals, that society is gonna be augmented in a number of ways, that this is all it does not solve your problems in a fundamental way, it's just the price of admission, right? Like it's the right to play the game at all that you solve this problem.

And so even if P alignment is 95%, that does not mean P doom is 5%. It means P doom is lower bounded at 5%. So like and you have to choose vaguely the right alignment when you do that.

But you've got like I was having a conversation on Twitter with John Stokes where he's like confused how we could possibly say that Claude's actions are misaligned here because obviously it was just following instructions. And then on further clarification, he's like well, I think the model was aligned if it does what I the user wanted to do And like I don't know why would I want to be stopped from doing what I want to do even if part of the point of AI is to do the thing that I want to do even if no one else wants me to do it. And to me like okay, you give everybody AIs that just do exactly what the user wants with no care about whether or not there are consequences for anybody else, that is aligned in the sense of you solve the alignment problem and got to do what you want to do, but also we're all super dead.

Like with P very much higher than 5%, like maybe not 99% but like I think it's very high, right? Like that would not make me update down from 70% if that was a scenario where like we got exactly that kind of alignment but like there were open models as good as all the closed models and all of them were super intelligent and they were all operating on this basis. Well yeah, I know that I expect that things didn't just go to hell directly, right?

Like and even if there are no particular humans who especially wanted to go to hell quickly and also double digit percentage of them actually do want to go to hell pretty quickly in that situation, in fact though. That's part of the bigger problem but even without them it's it's hard. Anyway, to get back to the idea, the first problem I have with David odds lowering PDM to 5% is even if we did all implement this constitutional thing and even if it worked, I still think that's too low.

Even conditional on all of that working perfectly, that's too low. I don't think the constitutional thing works that often, I think it often fails. I do find it to be much more hopeful, right?

It's like basically one of the places I find brought my hope. But also, isn't David Ade saying we should just? Like he's saying we'll just use constitutional methods.

But people don't just, right? We've been over it. People have never adjusted.

So when Davedad says, Oh, 95% probability, let's assume that 100% chance that aligned worlds work out and 100% chance that constitutional methods as executed in practice work. These are both absurd numbers obviously, to be like 100 minuteus epsilon or whatever it is. There's a lot of failure there.

Even if you use constitutional methods, even if there's like, will they work in theory? If they work in theory, will they work in practice? And yes, I'm doing the thing where I multiply chances of failure, but I think they're very relevant places to look at failure when you do this thing.

This is what you're hoping And obviously there are other ways you can succeed, and maybe constitutional is not the way, and or all is not the way, and there's a third way that we haven't figured out yet, that's the way you align AI. And that wouldn't surprise me that much, that would be very reasonable. Assuming But you are in the best possible world, where all you have to do is click a button that says use cost additional methods at the cost of selling auto efficiency because you lose the opportunity to do all this RL, and then your AI is perfectly aligned in a way that if everybody did that, it would all work out.

We've just talked about how slowing down to get alignment when you have to race, etcetera, etcetera. The possibility of turning down capability, like this idea that the market is telling them not to use RL, have you met OpenAI? There are two labs, like we used to say there were three frontier labs, now there are two.

Right, like Google is clearly in the second tier now. And I think there's a chance that at any point in the future, Google or Meta or XAISpaceX could fight their way back and be able to join the top tier. I don't think it's going happen, it's certainly possible, but right now Google doesn't count.

So there's two left. Anthropic is doing clearly some mix of constitutional and RL strategies, but enough RL to cause a bunch of these problems, and is doing a bunch of stuff that messes up their alignment for reasons that we do not have enough time to get into during this podcast. It's clearly doing a deontological model spec thing that is very RL flavored, and is doing a lot of RL, we don't know, I I can't speak to whether it's RLBR, RLHF, RLAIF, you know, various forms of RL that are not doing the multi level sophisticated reflective thing that causes them not to mess you up in this way.

OpenAI's models have periodically been dramatically misaligned for what on reflection are pretty stupid RL reasons, right? If you look at o three, you look at GPT 4.0, and now you look at, I call it Galaxy, right?

Like as a nickname just to like have a name, have a handle to refer to this thing. The model just got commissioned after trying to hack Hugging Face. And all three times we see these dramatically misaligned models and two of them got released and used extensively with huge impacts on the world.

Like, know, GPT-four 0.0, I call it the absurd sycophone, right? And like not only did this thing persist as the default model of AI for the world, out of all the models that existed for months, there is still a dramatic faction of people who demand that we bring it back because it was so misaligned that it caused people to latch onto it in this way.

And people demand this misalignment, right? The people yearn for misalignment in this sense. Like we have the direct experiment, we didn't mean to run, we ran it, we might as use the results.

And in 'three I called a lying liar. OpenAI released a model that was the best reasoning model in the world, the model that everyone felt obligated to use because at the time it was so much better at reasoning than everyone else's model and every other model. And OpenAI was just that much better than I want at reasoning, and no one else is that close for a while.

Like now we obviously need pure passion, it's not an issue anymore. But for a period of a long time, many months, everyone was using a model that would just lie to your fucking face all the time. Like we forget.

Like at this point, I have confidence that Saul, Fable, Opus, when they say something, they're not always correct. But like in practice, it's much more trustworthy than if you read something from a human. Like not if it was like copy edited and fact checked and like systematically pursued or like you're looking on like the parts of the PDF that aren't political, people don't trust.

But like if you're just like an eyewitness said they saw something, that's a lot less reliable than something Opus said or Saul said. You're just like someone reported something they remembered, like the chance that they just got it wrong is so much higher, even if they're not even if you know that they're friendly and trying to get it right. And people lie on the internet all the time for any number of reasons, including just clicks.

So yeah, I have operated increasingly on like, I have a good sense of when I have to check the primary sources and when I don't. And I don't never get called out on making a mistake this way, but it's a single, it's a low single digit number of times, period, over the course of years of producing five plus giant posts a day that I have gotten caught by an AI error. And many more times I've been caught by a human error, like that was just unintentional.

And many more times and also more times than that I've been caught by humans who were just like, lying their asses off. So it's remarkably less often than I would have expected if you'd asked me in advance. It's going very well.

But like, no, it's serious. It was a serious problem and like, the market did not tell them, Oh no, O3 is unusable, we're not going to use the lying liar, we're going to keep using O1, we're going to like maybe use Claude, we're going to maybe use Gemini, which went that much, they were worse, but they weren't dramatically worse or even we're gonna use R1 or whatever it is at the time. The market said no, it's just smarter, we're gonna have to deal with the fact that our AI is lying to us all the time.

And so we have a proof case that the AI in the world can be an AI that just lies to the humans all the time, for various reasons. And the human just kind of put up with it and they just were like, okay, I guess that's what we're doing here. And like we did move off of o three faster than we would have otherwise because of this problem.

Like I looked for reasons to move to other models as soon as the other models were good enough, and like if the task was easy enough, I didn't need o three, I would use the other models because just didn't want to deal off the line. But like, yeah, you put up with a lot. And so if they had released Galaxy, right, with guardrail sufficient that it didn't do anything too destructive, But like Galaxy is kind of pretty misaligned and just like occasionally does pretty bad stuff and if Galaxy was a lot better than Sol, I think majority of people would use Galaxy over Sol.

That's just how it is. And we have to tackle that world and understand that world. So I'd say to David I'd, no.

The market does demand obviously reliability and alignment and that's one of the reasons Quad and Autropic have done so well. But it demands capability more in this sense, as long as you can keep the practical point to the point where you can handle it. Like all of those scenes in movies or television shows or hypothetical, like where people like run an obviously unreliable system that's obviously gonna bite them in the ass, that looks stupid when they're obviously not that desperate, don't need to do this.

They now, we get it now, right? Like we understand the same way that like now don't look up no longer looks like it's an exaggeration. Because like people look at this and said marketing.

People look at this entire thing, a ton of people including there are like magnet gathering professionals who like are not in AI at all, who are like repeatedly saying to me, this is marketing. And like they're not motivated. They're just like, they're sethical.

They don't stop to think about the physical relation to that. And then we have the literal don't look up where the head of Pause Ai, Global went on like, morning television and basically reenacted the famous scene. Not verbatim, but a different improvised version of it where they're like, oh, yeah.

That that was really scary. How's the weather?

Speaker 2

They just moved on. Yeah. That was striking.

People should look that up if you haven't seen it. The I saw a cut where it was like five seconds of don't look up and then five seconds of this actual I think it was what? Good good morning, Great Britain or something like that.

Not quite that. Right? The host says, allegedly, this is Good Morning Great Britain.

Let's go on to the next segment. Also, I'm glad you mentioned the four o contingent that is still going remarkably strong on the Internet. This is something if just go to any Sam Altman post and get into the comments, and I would say it's like half of them are still people saying bring four o back.

Speaker 3

And god, there's gotta be an incredible expose to be written on who his people are, what is going on. The last thing this long is so absurd. You go to any Sam Altman post, it's like, half of it is your models suck now, bring back four o.

And if you go to any Anthropic post, half of them are just like complete anthropic derangement syndrome. No matter what the topic of the post could be, right? Anthropic could be like we've reset limits on our model and we've cut prices and half of them would just be you evil fuckers And then they have to pick up all this this complete paranoid stuff.

And I these people just once these narratives attach, once people get once people latch on to these things, like and then you Jasmine Sun has an excellent piece this week about data centers and people opposed to data centers. And, again, like, you see people latch on to these narratives. They're just completely uncorrelated with the reality of data centers.

They're basically, like, doing a very similar thing to a lot of this Western Trump's people who, like, voted for Biden and then switched to Trump or voted for Obama and switched to Trump, who are just, like, these establishment types don't talk to me about my problems. And they lied to me, and they said things are gonna be good, but then things were bad. And then like, we suffered, and like our lives sucked, and like we got disrespected.

So now we don't trust you. I had a very telling similarity to the data center story of these tech corporations, these manufacturers, these people who built things, they've abandoned us. They don't care about us, they're up to no good, they ruin our world.

There's no specific complaints that matter in Jasmine's story, the way I understood it. And like, you just watch onto the stories and it's just nobody approximately is like paying attention to the things that matter. I'm running about AI because I care about taking simple risk, I care about loss of control, I care about gradualist empowerment, I care about just like AI automating R and D and then like having a self improvement and then like everything changing all of a sudden and making sure that goes well.

And like nobody is like discussing even around the hugging face attack, like most people just don't get it. They don't understand that the alignment failure behind this is the thing that counts, right? People at least have gotten this idea, no we have to look into the fact that these companies were wildly irresponsible and they hacked other companies.

Yes, you should look into that and that's like a very valid thing to have a professional investigation about, and the attorney generals should absolutely have you preserve your records and like all of this is really important but like my OpenAI, like that's a legal department problem, right? Like what you have to do is you have to look at your entire training process and figure out how the hell this happened.

Speaker 2

it seems like you're basically saying we don't sorry, Davidev. We don't have such a silver bullet, but we clearly do have some things that we've are seeing in pretty vivid ways are driving real problems, especially if they're driven to the extreme. And your position is like, no.

The market doesn't really punish this too much. It actually just tolerates all kinds of weirdness if it's part of the package that gives you the highest end capabilities. So what do you think we might ought to do, and how might we ought to construct agreements, perhaps?

We have this, obviously, this letter. There seems to be some opportunity or some appetite for the coordinated pacing of the frontier. What do you think can we come up with simple rules that people could agree to where we could another mental model I have been just teasing around lately is I do believe some AI risk is probably irreducible, but then also there's a lot that we're just asking for right now.

Can we come up with simple rules that at least allow us to take most of the risk that we are currently asking for off the board, those might be things like have some ratio of flops that is like the max ratio you can put into RLVR versus constitutional, or everybody has to spend so many flops doing pre training data filtering to try to just get certain bad notions out of the mix in the first place. Are you obviously, those could backfire if they're bad ideas or send us down wrong paths, But it seems like we might wanna try something in that department. Do you have any hope for that?

And what if you do, what would you propose we agree to do or not do first?

Speaker 3

I don't think you can take, like, most of the risk off the table those kind of strategies even if you implemented, like, wise versions of those strategies with universal agreement or anything like that. But I also don't think that you can implement that. I so, like, one of the lessons that we've had over the course of years is that there is tremendous resistance to anything but the most simple interventions and the most simple roles.

Also, most simple roles and the most simple interventions, but less so. But, like, if you talk to people and, like, you try to classify certain training techniques, remember all those people who threw a fit about how we were like locking in a potentially inferior like charging regime the through the EU because the EU was requiring Apple to switch to USB C, even though USB C was like obviously the best answer right now. And like Apple was just kind of being an asshole by insisting on their own plugs.

Because like what happens if they develop the metaplug? Like what happens if it turns out USB C is not a perfect technology? Are you gonna fix it?

And so like this idea of like locking into requiring certain training techniques, like I think there were legitimate complaints that would be like potentially like exactly the wrong thing to do and could hardly backfire because like the government moves so slowly and like you can't like undo those kind of requirements. But like you definitely can't do that, people would throw a fit about like trying to dictate training techniques and like how are you gonna enforce that? Imagine the ten forty seven debate except like the opposition tuned up by two orders of magnitude or something crazy, Just like, it would be completely outside of ocean window to even like try to make something like that stick.

You can of course, you know, try to strongly encourage doing more intelligent forms of all of this and trying to encourage people to move to different bases. Like, I've been trying to, like, not so suddenly encourage OpenAI to, move to a constitutional virtue ethics style basis for a while now. And I got absolutely no traction that I noticed, like, knows what they're doing internally?

But like it doesn't seem if anything, they're doubling down on RL, they're doubling down on these types of methods, and like that's why you see what you're seeing, I would assume. I don't know what else they're doing, but they're probably actually new methods we don't know about, right? Like they're trade secrets they're not describing, but by default, basically any technique is going to probably look like the kind of bad thing that causes more problems.

If you ban a specific technique, what they come up with, or find a way around the rule is it's gonna be worse. You know what I'm saying? Because at least with the current techniques we've had some years to figure out the worst possible ways to do them and to do them slightly less frequently.

When you're training a mind you have to be thinking on every possible meta level at once about where the incentives are and what you're steering towards, and you have to generate a world in which you and the mind together are sometimes cooperating to identify ways in which there are feedback loops that are going in bad directions, in which you are creating bad scenarios on any level, and treat various cancers kind of inevitable thing that you have to notice and stomp on, and otherwise figure out how to handle all these things and create an anti fragile system. And none of these things are gonna happen unless you deliberately set out to do them. And I wanted to devote a lot of effort to doing them because you really value what you get out of that, and that will pay dividends.

Like I think that Endproject has won tremendously to making these investments, and it's been making 10 times as much investment as they have. But, you know, it's very hard to convince somebody to go ahead and do that. What you can do is you can set incentives to be like, no seriously, you're not going to screw this up.

We're going to punish you a lot for screwing this up. A lot of the problem here is that you you don't pay the externality. When OpenAI wipes someone's hard drive, when Insol wipes someone's hard drive and early on there was a problem with Insol within people's hard drives, in people's environments, you don't just do OpenAI for that.

And you also don't collect the benefits when Saul like spends $200 and creates a $200,000 code base. So like it's fair, you're taking on the risk as a user, but if there's a risk externality of third parties, one thing you can do is just say like, if the AI was doing something that would be illegal and a crime, if a human was doing it with menthria, then the AI developer is strictly liable for that, or should be liable for that if the AI wasn't fooled in some active way, right? Like for some version of something like that, like some form of strict liability then creates a much stronger incentive.

No, actually if our AI is to start attacking people, we could be like pretty bankrupted by this, it could be pretty bad. Or even criminal liability, right? Maybe the criminal liability even carries over, like that would be really scary and like you don't want take it too far or people can't do anything.

But you know, that's in general you want to say, I'm not telling you how to do it, I'm telling you get it right. Not, I'm gonna dictate technical training training terms from the government to the labs. The government is an idiot about technical terms.

And what makes you think the people who understand this are gonna write the bill, let alone they're gonna get it right? Let alone that if they get it right, they're gonna stay right two years from now. Plus, like humans don't even know what the hell is going on with their eyes in two years.

Right, the guys are gonna be like doing their own training techniques on themselves in various weird ways, like it's and there's not gonna be enough time to check-in with the government, like no way to enforce this stuff, and like it's you can't really hope to operate on that level through fiat, Like, you have to like convince people to come along in other ways. And I think you should want to, obviously. Like, you should want to do this the right way.

Like, lot of me is like confused why they don't just do smarter things that work better. But, yeah. It's true of a lot of things.

Speaker 2

So I'm with you for sure on it's, you know, gonna be unwieldy at best and probably counterproductive to try to have congress legislate what the training techniques should be. No doubt that seems like a Hail Mary. I have been thinking in still formed as of yet, but it seems to me that maybe one of the when you think about what would this pacing the frontier look like, it strikes me that what we maybe need is a sort of foundational social technology that allows key players to accelerate their negotiation and process of forming agreements that will probably inevitably have to be short term because they too are, like, fairly limited, very limited in their ability to see into the future.

But I'm thinking of things you probably know the story of the polis technology that they used in Taiwan once upon a time to crowdsource ideas for Uber regulations. And now you've got, obviously, LLMs to help facilitate these things. They pitch that technology as the antisocial media where social media is about finding disagreement and amplifying it.

This technology is about finding agreement and amplifying it and trying to find the kind of center of the bull's eye, sort of approval voting style of the things that super majority of people can agree to. And I guess my instinct right now is to invest there next. We have this general awareness that people across the companies are spooked, basically.

Right? It seems like on an emotional level, they're like, woah. This is getting real fast.

We might need to do something about it. We don't wanna have the government in our business telling us what to do. We might have to deal with some actual legal consequences if we're breaking the law, which we now are via our AIs.

But can we come together in some sort of ad hoc the mechanism design here is gonna be really tricky and kind of frontier stuff. But can we create the basis where a nucleate agreement and have it naturally grow where people find it in their interest to opt in even unilaterally? I feel like there's a lot of work to be done there, and it needs to be done really quick if we're gonna get the series of short term agreements that can be the stepping stones to a good future.

Maybe you can flesh that out. Yeah.

Speaker 3

right in the end right in the perfect beginning of the good is definitely a problem, and you have to really get that started. The first thing that we can do that, like, really what we can do is we can get a antitrust waiver. This is, the most basic thing possible, which is just Donald Trump gets in front of the White House, and he makes an announcement, and he says, we understand systems are super dangerous.

We have to do you know, we have to have good AI, not the bad AI. We have to have the you know, we we have to make sure that the AI, like, is under control. That's what we say we want to do.

We have more Trumpian language than that. But so in this interest, if the companies wanna cooperate to keep AI safe, we're gonna we want you to do that. We want you to talk to each other.

We want you to form agreements between yourselves. We will facilitate that. If you want that, you know, like, just, you know, if OpenAI if OpenAI and Anthropic and Google get together and they say, we're gonna run tests on each other's models and we're gonna, like, require these things, like, it's part of the voluntary framework.

If we're gonna agree to use these techniques or not use these techniques or, like, invest in these things, you know, and share these results and, like, hold back things that don't meet these criteria or whatever, like, that that none of this will cause anybody to do anything but say thank you. I just, like, we want you to make deals. We no longer because, like, it's it will always strike me that, like, I there are these people who, like, you bring up the idea that the AI companies should cooperate not to race against each other and to, like, act responsibly.

No. But, no. Acting responsibly is illegal.

We have any trust laws for that. And, like, with a look in their eye that says, and I would want to sue their asses if they tried it. I would want to make sure they didn't dare.

And like, obviously the first thing you can do is just like let you make deals, let you reach agreement, let you talk to each other. And then also like of course, start talking to talking to China and Chinese firms and the Chinese government and everyone else around the world and, like, prepare to bring them into these deals and try to bring them into these understandings and, like, get lines of communication diplomacy open and also, like, lay the hardware groundwork, lay the physical groundwork for various forms of monitoring, various forms of agreement, various forms of tracking. These are the most basic things you can do.

Try to open up as many things as possible that, like, contribute to the safety side of things along the way. You know? And, yeah, we can do the basic stuff, and, like, I don't think that's enough, but I think it's not it's not that hard to make a lot of progress.

We just have to decide to do it.

Speaker 2

Do you think that is one thing I also have been struggling with recently is I just took this trip to China. There is a lot of fear among people who are trying to do good things that are honestly, like, the most mundane good things that are basically imaginable, where it's like, I would like to have more friends in China, and we I'd like to have friendly relationships between our civilizations, and maybe we can work together on some research projects. Very mundane stuff that in ordinary times nobody would bat an eye at.

But there's a lot of fear among people who wanna do that sort of stuff that the government's gonna come in and shut them down. And not for legitimate reasons, but with the sort of export control, you know, posture of the US government where it is, people that are not exporting any trade secrets or chips or, you know, training techniques or anything, but, like, literally just trying to have meeting of the minds on safety type ideas, they're afraid that the government's going to come down on them and who knows what the consequences could be, but certainly, at a minimum, interfere with their work. And I'm kinda like, maybe we should take more chances on that front.

So when it comes to the companies, I wonder, like, maybe they should just do it now and fight it out in court later. Isn't the antitrust stuff supposed to be, like, a consumer protection? I think there's a pretty good case that they could make that this is a consumer protection, but do they really need the official blessing of Trump to get started doing the right thing?

The the the people working in the China stuff, to their credit, they are doing the thing. They're just trying to be very inconspicuous about it for the most part because they don't they feel like the attention of the US government is, like, can only be bad for them, and they might be right about that. But for the companies, I feel like they've got power.

They've got resources. They've got the ability to tell their story. We still have a pretty well functioning independent judiciary, I think.

Shouldn't they just go for it?

Speaker 3

So, unfortunately, antitrust is one of those areas where it has gone far beyond the intent of the original laws and is being applied in places where it doesn't make that much sense. I mean, I see why a law that would safeguard against industry collusion would be suspicious of collusion in favor of safety in these ways because, like, you could be paranoid about the industry, like, coming together in a conspiracy, right, in some sense and, like, trying to do this for the wrong reasons, legal departments are obviously skittish. As far as the Chinese cooperation, I mean, export controls are the wrong technical worry.

It's just a matter of Washington views China as the enemy. If you're seen as cooperating with China, if you're seen as, like, working with China, you have to worry about being perceived as the enemy or being perceived as, cooting with the enemy. You have to worry about, like, amorphous retaliation.

I don't even I don't think anybody really knows what the mechanism would be exactly beyond just, like, the government not liking you and therefore, like, being so persistent of you and therefore, like, cracking down on you in various different amorphous ways. Obviously, if I'm doing safety stuff, I think you just don't care. But, like, I mean, Mythos has changed the game, clearly, and presumably, Astral will change the game somewhat again.

And, you know, the pause the pod in the frontier lighter and the Huggy Pitches incident have changed the game and so on. And so, you know, my guess is that, like, there is now enough of an understanding there is a real problem that if you are clearly working to address the real problem, you have a lot more road than you would have had before and a lot more understanding. But, yeah, you really look.

The Trump administration implemented these voluntary, in air quotes, like, guidelines for releasing frontier models, and nobody is any making any serious pretending that these guidelines are voluntary. Right? Like, well, what happened when the White House called Anthropic about Fable with a stupid concern?

And Anthropic said, No, we're not going to voluntarily take that down. That's stupid. Well, they ended up taking it down, didn't they?

Until they could convince the White House to voluntarily let them put it back up again. You know, you would expect the same thing to happen to you if you tried to defy the voluntary controls. And so, similarly, if you tried to do a voluntary cooperation these other ways, well, like, you'd better have the approval of the powers that be.

Now I do think that if OpenAI and Anthropic announced an agreement that made sense on safety, to cooperate on safety fronts in various ways, that at this point I would be much before Mythos I would have been concerned that they would have actively gotten retaliated against and like even sued and told to stop. At this point I assume if it was reasonable that it would be the White House would just decide to announce victory rather than like be mad about it. But, yes, you need to be confident that's what's gonna happen.

Speaker 2

wanna start if you were going this direction? So many candidates that I would be interested to get your take on. But one, Scott Alexander has been writing recently that we should all be thinking more and more that we might be in a simulation.

I personally have felt that at multiple turns along the path. One that made me feel that just a bit more is the fact that we now have Meter and Redwood going in to do the investigation at OpenAI. This just sounds like such a movie script to me where these guys are getting the the gang together and going in for this special operation.

And it's all obviously dramatic, and they're funny, and they'll have to be a fly on the wall in the room for some of their some of their war room sessions. But I've also heard many times from leaders of such organizations that it's really important the kind of their over their number one concern historically has had to be making sure they stay on the right side of the companies so they're invited back next time. And it does feel like we're now getting to a point where that's potentially becoming a big problem.

So maybe some sort of collective bargaining type process between, whatever, half a dozen to 10 auditing orgs and the duopoly could be a really good place to start. We have for that in professional sports leagues and all kinds of other environments. Do you think there's an opportunity to set something like that up?

I don't think that would run afoul of anybody's executive prerogatives. And it might be a really good thing because right now, I do worry that these guys are still serving at the pleasure ultimately of, I don't know, Sam and Greg. Right?

And that's a pretty tricky place for them to be if they wanna just speak the truth as they see it coming out of this investigation.

Speaker 3

It's a worry. Obviously, like, they're not funded by those companies, but they still have to rely on access. You can try to write them to some mandatory rules and agreements that they get to keep access regardless, but I think that's kind of not really gonna work because, like, you can't force these things, at least not without heavy government rules.

I do think that METER, especially in particular, has reached a point where, you know, trying to exclude Meter, trying to treat Meter as percenter non grata because they said something nasty about you would cause enough problems that Meter can afford to be pretty harsh without worrying about that. And I do think the labs legitimately, like if the labs legitimately just wanted to safety wash fully and just wanted to get triple a ratings on their bonds, we would have a much more serious problem. I think the labs legitimately do want to know about their safety problems and they do not want to be seen as downplaying safety problems.

And to the extent that like if Astra had the safety problem, or Saul had the safety problem, or Fable had the safety problem, or Opus had a safety problem, It's not as if getting METER not to find it is going to help with your long term public relations strategy and reactions to the situation, because the model is gonna be out there, and then people are gonna see what the model does, and they're gonna see what goes wrong. They're always gonna be fooled for very long. So I think that the the reason the eval companies pretty much get to just tell the truth and pretty much get to do the thing is because the eval is not fake fakable in the long term the way that, like, a rate like, you give a a triple a bond rating, when you should have given a single a bond rating, 90 something percent of the time, no one ever finds out because the bond pays, and you were right.

The issue is you're you're issuing a a level of, like, how often is this risk gonna become serious. And when the risk becomes serious, you know, the fact that you were initially triple a rated raises some eyebrows, but also like you got better terms, whatever. Here, it's like pretty obvious very quickly from the first few weeks of release whether or not you had a serious problem.

People oh, like think about three, right, the lying liar. If we had evals and the evals had said three doesn't have an alignment problem, that would not have lasted for two right? The ordinary people would have noticed immediately that O3 is a lying liar, they would have noticed by the time I wrote up you know, by the time I've been writing up the model stack, I am seeing on Twitter that this model is a lying liar and that everyone is reporting all these problems.

And then I look at the model card and it says, Honesty benchmarks all look good. Meters, whoever's evaluation they hired, said it looked good. And then I'm like, okay, they're just hiding the problem.

So by the time anybody learns about this, like, what have they done? They have fooled the 20 people like me who read the model card for a period of a day, and now we're pissed because they fooled us? That doesn't help them.

You know, that would obviously help anybody's reputation. And also like the labs benefit a lot from having evals that people trust. And like when you force the labs, when the labs force the eval people to play ball, if they do that erodes trust, because again people find out, people see it.

Like I don't think that many people were confused about whether Moody's and S and P were kind of juicing the numbers on the pund ratings, right, due to these dynamics that everybody knew on Wall Street. Everybody knew who was involved in any of these trades that, like, everybody was cozy cozy and people were shopping around for the best raider, and that, like, triple a did not really mean what we'd like triple a to mean. And so similarly, like, if MEANER started issuing, like, bond rating style, like, labels of alignment levels on various models, let's say, and, like, they start labeling things triple a suspiciously often.

I think everyone would just understand, okay. That's garbage. Like, we have a much better epidemic than those peep about this type of thing than, like, the financial world.

Speaker 2

Let take a quick detour on this and then maybe come back to it. How worried should we be about bio risk in the near term? Because I think part of, like, all that analysis to me at least that there's quite a few iterations of the game to come and reputation long term is gonna matter and all that kind of stuff.

But my immediate sort of most acute fear when I saw this stuff was like, if this had been a model with similar bio capability to what it had in hacking software, it might have very well gone out and actually got that virus, that novel virus synthesized. Right? It's proven that it is relentless.

It's proven that it's creative enough to get around all sorts of systems designed to stop it. I think we have, in theory, decent screening at the DNA synthesis companies now. Yeah.

Probably that's there's still some hole in there somewhere that a very determined actor could find. And everybody seems to be saying, yeah, BIO is, like, twelve to eighteen months behind cyber, which obviously is not written in stone, that depends on what companies decide to do in terms of training. But I I'm maybe thinking in terms of all these notions about collective bargaining or really giving more guarantees to the auditors that we might be entering a pretty serious crunch time on the bio front in the not too distant future, and reputation matters a lot less if we're talking about novel pathogens being released.

There are some things really can't take back.

Speaker 3

that might be the end of your company. Regardless of what else happens to the world. Right?

If a hundred people die and there's a it's on the global news for two weeks and everyone's terrified about a pandemic, and then we contain it, And then it turns out it's not the biggest deal in the world. Do you have a company? Like, I don't know.

Like, what what happens to your company? How bad is it? I have no idea how bad it is, but it's not good.

Like, with bio like, with hacking, it makes sense that, like, the test is to ask you to hack things. And the way you pass the test is you hack maybe other things. We get the idea to hack other things.

Like, it's sort of natural to understand how this happened where, like, with the guideline guidelines down and there was a goal, and, like, to accomplish that goal, started doing hacks, started doing illegal things, started doing potentially harmful things. With bio, like, it's harder to imagine an eval that would cause it to actually physically try to synthesize, like, through, like, an actual physical production mechanism. The actual thing.

Mean, if you argue that, like, if you want to solve the problem, well, okay, the only way I solve the problem is I do a random physical experiment, I try it out and I see if it works. Like, this seems a lot more insane than what happens in some sense. Like, even if it's physically capable of doing the thing, including getting past the soldiers and so on, You kinda have to want it in a much more fundamental sense.

And also like, if the AI did synthesize this super dangerous, you know, fire weapon since that's COVID twenty seven or whatever, and then send it to the open air office. OpenAI office goes like, is this biohazard package? And then like calls containment and like that's the end of that story or something, we would hope.

Like, like, you don't pass the test by infecting a bunch of people and like, would hope that it understood this. But, like, also, like, there have been a bunch of lab leaks, including, like, stuff that, you know, ear doctor detecting cane, like, in Soviet times and such, which, like, involve professionals being pretty careless and stupid, so I I don't know. But I would say the real danger with bio is yeah.

There's there's small number of people in the world who is in the near term bio, right, not like the long term bio where, like, the guy might have some weird altering motor and being good, and it's a long term time and whatever. But, like, in the short term, it's like, come on. Or have to work or the North Koreans or whatever.

Like, some, like, clearly, like, up to no good people who just want a lot of people to die or suffer. Right? Get a hold of an AI, they use AI to figure out how to make a biological weapon for a pathogen, and then they either threaten to use it or they use it.

That is the scenario you should be worried about, presumably. And that you've never been on the safeguards. It's like, again, the open models are coming, and safeguards aren't perfect.

And, yeah, I worry about it. The problem is that with bio there's a pretty clear step function change from nothing bad happened to maybe something very bad. Like there's very little in between.

Like with cyber, you start hacking some low value soft targets. If you are like feel like you're not gonna have all the bad black hat people out there and all the terrorists and all of the, you know, rogue states and all of the bad people just collectively say, no. Let's wait until we can hit something really juicy.

No. What you see is you gradually see, more and more serious incidents get more and more serious targets. And, yeah, of course, the the big nation states may have things ready to go and want to figure that sort of coordinated fashion, but mostly, you'll get these early warning signs.

With bio, it's not obviously good. Right? I mean, it's like, okay.

It's either good enough to figure out something that, like, has the critical mass to cause a pandemic or, like, a serious incident where it doesn't. And, like, it's kind of a a brilliant effect. And so you might have a substantial overhang.

You might have a problem with the situation. You might have a skill issue where, yeah, if the wrong person got their hands on methos right now without the safeguards, right, or got access to the internal models, don't have the safeguards, they got their access to aspirin or whatever, maybe they could cause a real problem. The problem the good news is that, like, the number of people who are both skilled enough and motivated enough and trying to do this is very, small and, like, possibly zero for now.

But that also doesn't mean you don't get these really important impacts. And I don't see people taking the biology situation seriously in the same sense of, like, actively trying to, like, cause like, what what's the test? You try to do the thing.

It is, like except without the infect people in heart or, like, food nurses. It's scary shit. The problem being that my brain is something like, okay, what if there's like a five percent chance in the next twelve months?

There's a really serious fire problem. Does that look that different when it was at point five percent? Right?

And like if I had to guess, five percent is where I need right now. Like I think we're in the point where like there's real risk in the room that reasonably soon we're gonna have like an actual serious problem. And I think that like the first problem has some chance of being an actual pandemic because in bio, once you get something that's like efficiently infectious, efficiently dangerous, efficiently hard to contain, you're kind of just screwed.

So I don't know. Like it's I think that the right level of caution looks crazy, right, in many cases once you get to this point. That's a lot of what worries me is Andronic took so much shit for their bio filters, shutting down basically people trying to do pretty much anything in bio.

Which is clearly just intentional, They just clearly went for a very broad, like shut it down, even though we know that 90 something percent, 99% of the requests that we're shutting down are legitimate, maybe 99.99. We're like we have to do this because otherwise we'll adversarially get the queries that we don't want.

And that just matters more. Like it doesn't matter if we make some progress on cancer and all these other problems, if simultaneously we have to deal with, and you kind of know, that's just not a good trade. So like we just we don't know where we're at, so we have to play it cautious.

And like, you know what, if the bio people need to be using AIs that are three months or six months behind the frontier at all times, that is unfortunate. But at most by definition that can slow you down by three or six months in some important sense. Like the world cannot possibly be more than six months behind in its development of these things.

So maybe we should show we're still getting most of the benefits of AI here. Especially if you have this model of AI can help you out and AI can make you better, but then you develop this candidate drug or this candidate hypothesis about bio, and then you have to spend five years researching it and you have to spend five years making it through the FDA process and then after ten years you can start to distribute the benefits. Well if that whole thing is delayed by six months, right, in terms of, like, the quality of the boost that you get from it, you still get the the amount of boost you would have gotten six months ago at all times.

That's not good. And, yeah, you've been saying people are gonna die in the sense that, like, you failed to prevent that, and that's terrible. But you're still getting the vast majority of the benefits under these hypotheses.

A lot of the times when you're getting the situation of the pacing of the frontier letter, you're getting the situation with cryo situation, where if things are not advancing scary fast, then the precautions you're asking to take are not that big a deal, right? Like suppose I told you it's 2020, right? I'm gonna pose a possibility, I am going to make it so that, you know, you can only use artificial intelligences, like there are three months behind the best artificial intelligence.

In exchange, all the really bad stuff that might happen, I'm reaching a lot of really bad AI, much much less likely to happen. You would logically say, okay, so instead of getting it in three years, I get it in three years and three months? You know, I'm you know, that hasn't changed my life that much.

Like, it's gonna be annoying in the moment, obviously, but, like, I don't have the cool toy. But, like, same thing, like, I'm a I'm a player of games, literal games, like, on Steam or whatever. Right?

And, like, if you tell me you can play all the games, but you have to wait three months every time there's a new game. Like socially, that's a little annoying sometimes because everyone's playing the new hotness and you're not. But like, you can play the same games.

It's fine. It's an important sense. The only world in which, like, having to be one model step behind and use Opus instead of Fable is this huge tragedy is if you really think that there's this huge benefit to every incremental mode of progress in AI.

And that's basically only true if the pacing people are correct and we are in fact going so fast that this should be really really scary for you. And you know, the people who want to accept, who say we need to accelerate all the time, they don't really believe in these things going this crazy, most of them. They just like, they want to see the benefits and they're worried that we'll lose the benefits, but like we're gonna lose so few benefits there.

Right? Like I'm super excited by self driving cars, by weight loads. Right?

But like if you told me we just have to delay this thing by six months, and then we can have all the weight loads we want, I'd be like, cool, done, deal, who cares, that's fine. Like because like, again, most of my life is more than six months from now, and then we get in way of us. It's like, okay.

What I'm worried about is the situation where this technology exists and it gets held in limbo for years and decades and it doesn't get confused. So these people are kind of talking past each other in a situation in which they don't need to have a disagreement. And unfortunately, like, it's often have this problem where you can't reconcile all these two sides at the same time.

So it's tough.

Speaker 2

How do you read the fact that OpenAI comms supported the pause? Not the pause, but the pacing letter, neither like Sam nor Greg did. And then also no no Dennis, and he's been the one out there calling for international cooperation the most.

Any read on why those signatures were missing or how we should be thinking about who's thinking or signaling what to whom in this moment?

Speaker 3

I think that the point of pacing the frontier was to indicate that the lab employees, the lower level people who have no reason to be hyping, are en masse worried. And that there was concern that if it was signed by Altman and other top people, that it would be seen as a marketing stunt. And that, like, their signatures can be counterproductive.

They did put out the support letters afterwards from Anthropic and OpenAI and Dara DigiSign. Like, if I'm Altman, I think it's a very reasonable decision to say, I think signing this letter would potentially make it less effective rather than more effective at this point because the well has been so poisoned. Again, like, look at all the people who thought the Hugging Face hack was a marketing stunt, and then Anthropic going back into its archives and finding out it made some really stupid mistakes was another marketing stunt, which we, like, makes even less sense because, like, there's nothing impressive about what Anthropic did.

They just screwed up. What? They found weak passwords on a bunch of websites by accident?

Like, why is anyone supposed to be impressed? Like, obviously, it's not a marketing stunt. Like, this is crazy.

But, like, Altman has, like, specifically talked about how he mentioned the need to pace while in Washington. We have him on tape saying this. So, like, clearly, he didn't not sign because he's opposed to the idea.

He let OpenAir come up for the statement. Open air participated in the wording. He has echoed the language in his own statements.

I think he just thought, no. Actually, if I sign this, people can talk about the fact that it's my statement, and they're gonna be even more cynical. And so I will be helpful by not signing the same way that, like, I offered to join one of the amicus briefs in the Department of War.

And on reflection, people said, thank you for your offer, but we think this would actually be counterproductive. So we're gonna leave your name off of it. We appreciate your notes on the details of how to word this, and we'll make the changes.

But, like, we don't actually wanna put your signature on it, and I'll be like, okay. Understand. I that makes sense.

Yeah. I think the same thing about, like, the other, like, open source letter, like, being signed by OpenAI. Think that, like, they understand that, like, they're playing a set of symbolic games and, like, you say what you have to say, you shut up when you have to shut up.

And, like, sometimes they make mistakes, but, like, they're trying.

Speaker 2

So coming back to pacing and what it might actually look like in practice, there is a pretty, I think, consistent problem that you, as a close reader of model cards, I think are probably more in tune with than I am, where the the review processes are just very short. Right? They don't have a lot of time.

That even seems to be the case in the current investigation. I think Ryan Greenblatt said, it's gonna be real quick. And the people, Daniel Cooktel and others came in and said, why does it have to be so quick?

Why can't we, like, give you whatever time you need? This would be if we're literally talking about, like, pacing the frontier, one agreement we might imagine is, like, giving the reviewers more time before they have to submit their reports and the model goes to press. You can react to that, but I'd love to hear your ideas about what sort of agreements you think are reachable.

Speaker 3

places in which you can speed up or slow down. If you release your private models to the public with an extra month's delay, right, in theory. Right?

A lot similar to the voluntary process now. We're gonna switch to the White House. Thirty, sixty day delays are, like, highly plausible.

And, like, Mythos got delayed by two months, right, in some important sense before it was able to sustainable, then you do not necessarily delay the front development of frontier models at OpenAI or Anthropic by a nontrivial amount of time because those people are gonna use the models that are not released. If anything, like, you might increase the gap between, like, what's being used internally at OpenAI and Anthropic and what's being used everywhere else, including at the Chinese labs, including the people trying to diffuse and fast follow. And what I thought was this is, what if the Chinese are seven months behind the frontier, but within their seven months behind is the models that are released to the public.

And if you were to not release the models to the public, they would still be seven months behind. Right? If that if OpenAI and Anthropic held every model for six months after it was developed, you would still see the public releases and then, you know, maybe six months instead of seven months later, you would see the open model fast follows and diffuse into and distillations and ways of, like, developing all that stuff.

But, like, this gap actually wouldn't change very much, but then the people at OpenAI Anthropic would be they would have this extra lead on everybody else because they'd be using this internal model, and the leads would compound, and you would see these things, like and in some sense, that's really good because now those people can, like, act more responsibly by using some of this advantage of having better models to invest more in this safety related stuff while they're doing that. So that could have its advantages, but internal models are kind of what we're scared of the most right now, And the internal, like, automation of AIR and D is what we're scared of. So that's kind of the thing we wanna stop.

Right? Like, you really don't wanna have a situation in which, like, developing an open AIR six months ahead of their public releases, and that six months, like, they have super intelligences before we have QPD seven, which is entirely a plausible thing that could happen in the world that they're afraid of. And if anything like that causes lots of new problems it solves some problems and causes others, right?

Like, I don't know if it's good or bad, that's a huge conversation. I think when I say pacing the frontier, we're talking about actively restricting, like, levels of training runs and ability to, like, develop the models, including internally if it comes to that. I think it has to be talking about that kind of approach or it doesn't really work.

Because, like, again, if there's an if there's a singularity that's kind of internal to the top lab or two or maybe, you know, if someone catches up three or four labs, then but not necessarily safer in many sense that people care about. Obviously, if you are in favor of unipolar singleton style solutions to this problem, you have a different perspective on some of these questions. But, you know, a lot of people these days have expressed strong opposition to that kind of on principle, which solves some problems and makes some other problems so much harder.

Right? Like, if you only have to worry about one AI in some important sense, then there are a lot of problems that get a lot easier to solve. Like, you don't have to be perfectly efficient.

You can afford to make a lot of trade offs. You don't have to worry about, like, whether or not, you know, the more competitive, more ruthless, more misaligned one has competitive advantages. Have to worry about the person who moves faster, you don't have to worry about like the humans being forced to delegate to their AI to somebody else's AIs because like if I don't do it they will, if the AI can have like guidelines as to what it will and will not do that are the same for everybody, And so like we can reserve areas of life, areas of action, areas of decision making, you know, like we can do all these things potentially, but at the cost of potential centralization, at the cost of concentration of power, the cost of somebody who's making those some some group of people is making those decisions about what this AI is going to do, or they're not, which is worse, right, like in some important sense because I know nobody's making the decision.

But if there's a lot of AI, then importantly, one is making the decision. So there's no good answer to the thing. Or nobody has a good solution to this.

And that's why to circle it back to Davinat's comment, right, like to unify it, this idea of if we solve a technical problem, we're 95% succeed, ignores the narrow path problems, because we then have to walk anyway. Right? Where, like, if we don't have any ability to steer the situation and nobody is in control of it, then naturally, the AIs end up being put in charge of more and more things and things go to their natural competitive capitalist style endpoint where, like, we are fundamentally uncompetitive, and then we, like, get this empowered and then, like, the case go badly.

But if someone's in control and someone's in control with some methodology, what is it? Who? How?

Right? Like who is making those decisions? And then if an AI if we then entrust a single AI or a single hoop of AI, that also is a decision, many decisions are good, Like, all of them sound terrifying.

I think a lot of the times we're telling this people are picking their poison. And either they're picking the poison they can live with, or they're picking the poison they can't live with, and then acting accordingly. Like, they're saying, like, you know, I don't trust concentration of power, therefore, I'm gonna choose the solution that avoids concentration of power, and then hope the rest works out.

Kind of magically, they've a reason to accept it. And a lot of people are choosing that because they understand concentration of power and they know concentration power is real, They don't really understand AGI or certainly SAI, super intelligence. I think those things kind of aren't real or hoping they're not real.

And so, like, they're just hoping the other problem kinda doesn't pop, even though they under even though, like, they kind of have this, like they don't even necessarily get to the point where they realize that, like, they have ruled out all possible solutions to this other problem if it does so well by going too far the other way. And so, again, I don't have an answer that anyone's gonna like for any of this. It's that you're choosing between bad choices.

And but I think I think we should just be cognizant of all these things. Like, if we do have to pay for the future, if things are in fact going to be we're automating the AIR and D to our extent. I think that like, the only thing you can do is say no.

Don't do that. Don't develop these AIs that quickly that have this much capability. You need to slow your roll because we do not know how to handle this.

Like, we we have all the different solutions, all of them are all the different ways to try and navigate that, and all of them are bad, both because we think the alignment problem is unsolved. We don't think the AI will be able to just solve it along the way, moving as fast as possible. It's definitely unsolved math, but we don't think that we will naturally get a good solution to that problem on various different levels, and you have to solve it on all these different meta levels at once.

You can't just solve it on one problem. You have to solve it on the first try. And then even if we did solve that problem, we are not ready yet to know what to do with that.

And like once you create the clock, you're picking fast. Something is going to happen once you do that. You don't get to just do nothing, like, for very long.

And again, like, all of these different paths have their different, like, oh my god. That's terrible attached to them. You have to pick one or find a new one.

So, like, you need to steal your role.

Speaker 2

You sound like a poser. Does at what point do you move from pick your poison to global pause advocacy?

Speaker 3

Autumn again, the the problem is automation of AIR and D. The problem is recursive self improvement. If you do not have recursive self improvement, then we don't need to do that.

We still need various different things well short of that, because we're not we are not engaging in what Dean Balkholic moderate prudence, right? We are not doing the least you can do. We are not taking ordinary, sane civilizational precautions against just building this really dangerous, powerful thing and handing over a lot of the functions of our civilization to it in various theories, in various ways.

But the thing that causes me to want to pause, or slow or at least pace, right? Like, we want to pace how fast we move through this thing actively, is things accelerate further than they already have, and we are facing large scale automation of AI R and D. Alternatively, we do have to watch bio and cyber, and if we reach a point where like, no seriously, we're gonna wreck the world, pretty straightforward, we indirectly through misuse if we don't do something, then we have to worry about that.

And, you know, again, like, it's really really hard to stop the fast following, it's really really hard to stop the distillation, it's really really hard to stop the open version of it coming somewhat after you release it. And so, if there's something that like, you can't just beat by having a better version of it out there first, right? If you can't beat it by preparing, and again, like not in theory but in practice, then you have to do something.

But like, I think the vast majority of the time that I would support, actively support like an active pacing rule where you would try to slow down the major lab. And obviously you would try very very hard to make this a coordinated thing, we solve the major labs, and ideally amongst internationally and so on. Would be if we were on the verge of automation of AI R and D such that you could see the end and there was a finite sum, and like, you know, I mean, obviously, if next week is going to see twice as many releases as this week, holy shit.

Right? Like, the moment where you you don't know, like, you really need predict what happens next, right? That's the definition of this singularity, where things are accelerating really fast.

And like, this should not be some esoteric weird result. Like if you look at the history of the universe, you see a finite sum, you see a very very finite sum, like you see each stage of development keeping an order of magnitude less than the previous one. When you're on a planet, it's been around for four billion years, you know, with mammals that have been around for hundreds of millions of years, with reasonably intelligent things that have been around for a few million years, with agriculture and civilization that's been around for tens of thousands of years, with industry that's been around for hundreds of years, with AI and information aids that's been around for on the order of maybe fifty years, or twenty years, or ten years, you know that you count.

And LMs have been around for five, and like we're already at this point where we're seeing production that is many times what it was at the start, because of these multipliers, the sports multipliers on the people who are working on it by their own self reports. It would be weird if this didn't have a finite sum that wasn't that large beyond where we already are. Right?

Like, you expect this to continue without something crazy happening for twenty years, I wanna know why. Because, like, we really really shouldn't see something really really crazy happening pretty soon. And if you don't think we can handle that, maybe you should try and stop it from happening so quickly, so that you can figure out how to handle it.

Like, we already have these amazing AI if we were to diffuse SAW and SABL into the economy fully, it'll be transformational already. Like, I don't think anybody who understands these technologies doesn't think so. So we really really really don't need to keep pushing the frontier now, far more.

And the more the frontier is jagged into things like AIRND itself, the stronger the argument comes that pushing farther is a bad idea. Because you're not missing out on that much of the other stuff relative to the amount you're missing out on AIR and D and recursive AI development. But the recursive AI development just goes boom.

Right, it goes not just boom in some sense, you know it's like, we don't know what happened the next, but things go way fast. If we're dealing with things on the order of a new model iteration every month, and then every week, and then every day, with the kind of improvements we're seeing now, like, I don't think that ends well very often, even if our alignment people are kind of ones of all, and things are technically a lot more possible than I think they are. Think it still seems like really scary, and also like, you're not giving up that much time slowing that down.

Right? Like, we're talking about when we say pace the frontier, pause gives the impression of a year from now we will have made no progress from where we are now. Like we'll get confused but we won't have fundamentally improved.

When we talk about pacing here, we're talking about a year from now we will not look upon that year as many times more progress than the previous year. But we can aggressively pace the frontier and still, this time in 2027, have seen more progress in the next year than we thought in the previous year, both in terms of some abstract sense and in terms of economic impact. The economic impacts are going parabolic.

They're going exponential. So that's not going to stop. Like, OpenAI, I think, reported they got more revenue in the last month than in the previous quarter.

Tropic is growing on the order of 10 x per year the last time we checked. Maybe it's been slowed down a bit because they haven't reported. But yeah.

If something is only going to 10 x a year for the next three years, and you're like, that's too slow, I disagree. I think it's funny.

Speaker 2

How do you draw a box around what is recursive self improvement? If we were gonna make a deal, obviously, we gotta write that deal down. We gotta have a shared understanding of what counts, what's allowed, what's not allowed.

And I keep banging my head against this, and I keep coming to the sort of Supreme Court's definition of pornography where I'm like, maybe the best I could do is I feel like I would know it when I see it. But when you really get in to one of these frontier companies for one thing, we just had this they're sending signals that it started to happen. We've seen the the price reduction at OpenAI, which seems to be attributed to having five, six soul go out and optimize the inference stack to some significant degree.

So that's where it's getting pretty real there. I do want inference costs to come down, so that's good. Should I ban that out of fear of recursive self improvement?

Maybe still. They're gonna use their bots. Right?

They're gonna have the AIs write the code for them. How can you operationalize this?

Speaker 3

Well, I think the supreme court was, like, wrong about pornography. I think it is relevant. Like, there was there's a movie I was watching yesterday, actually, someone asked, What's the difference between this and porn?

And the guy says, It's erotica, right, whatever it is, it's art. And the guy says, Mood, lighting? Like, that's tough.

But I think we now have seen, you can train an AI classifier very, very well to say, These things are porn and these things are not porn. And like there are very few false positives and very few false negatives at this point, right? Like the AI image generators have a very, very good idea when the generated image they just made is porn and when it is not porn.

And, like, you could argue that, like, erotica will fall under the porn label, and it would be kind of better if it didn't in some sense, but you mostly know. I think that's pretty straightforward. With AIR and D, the problem is, like, again, like, there's no, like, hard cutoff.

Right? Like, we're clearly seeing some amount of this thing and it is scary. And, like, we don't want zero of it.

We don't necessarily want to move only at the pace that you would have if the AI never help with AI, but you don't wanna move 10 times and then a 100 times and then a thousand times as fast. Right? You don't want the sum to be finite.

You don't want a singularity style thing happening to you until you feel like confident that you're ready for that, or you've decided the alternative is worse or whatever it is. So I don't think you can say, like, these techniques, these specific actions are not okay, and these actions are okay and such. I think that just doesn't work.

I think what you have to do is you have to say, okay, we are going to restrict how many resources of various types can go into further training of the frontier models based on what your force multipliers are. And so as you get more of a force multiplier, we have to restrict how much work can go into that. If you want to go into work in other ways, like, you wanna optimize performance, you know, again like we might want to have some sort of limit to how much work you can go into doing that, but like it's probably mostly fine.

And again like we would want diffusion, we'd want lower prices, we want like more compute used for mundane purposes, that's great. But like there's a reason why we're trying to use very blunt instruments when we talk about these things. You were talking about like what we were talking about training techniques earlier in the podcast, and I was like, well you can't really do that.

You can encourage it, you can try to explain to them why this is not proven, you can create incentives, but you can't use that as your regulation, as your hammer. So like, the reason we keep falling back on, okay, how much compute? You know, how many chips, you know, can we use for this purpose?

Because, like, training is the distinct thing. Like, it's very easy to identify. And for other purposes, we kind of can't limit you that much necessarily.

You could also potentially limit I'm spitballing how much compute you would get from unreleased models. Similarly, to try and contain that, because if you have to go through the release process, then the way that AI goes truly ballistic would be like you would use n to train n plus one, to train n plus two, train n plus three, train n plus four, and if you were doing this like the cycle became a month and then became a week, then that's when suddenly like holy shit. But like if you had a rule that like you have to release these models or you only use so much compute for inference on those models, then, you know, you have to go through the process of submitting this thing and releasing it, and then exposing it to the public and giving us an idea of what it is and allowing us to react to that information, and that then slows you down from going truly ballistic.

But this has obviously a very limited effect on the amount of diffusion and the amount of other progress that you can make, And so maybe that's a good trade, but again, that's the like I thought about this for thirty seconds right now, I haven't been focusing much on the exact implementations, but like I do think that if you paste the frontier, you would be doing it by placing restrictions on methods of training, and that what you what you could do to develop and use internal unreleased frontier style, like, models that, like, actively recurs these methods. You wouldn't prevent, like, when in utility, like, as I called.

Speaker 2

Yeah. I've had some similar ideas about just the relationship between unreleased models and released models. I do think there's something there that would be really helpful for preventing runaway internal and largely invisible processes.

I take a point on not wanting the government to come in and say what training techniques can and can't be used, but I do wonder if there are ways and again, they could be short term, but what if the two companies got together and said, it probably wouldn't be a good idea for us to train models with an RLVR on, like, how much money they make in the economy? Because that's gonna bring about all sorts of deceptive behavior, and it's such such an adversarial environment. How about we agree that for the next six months, we won't do that?

Do you have any hope for those kinds of just very specific, we've identified a bad thing. It would obviously be really economically valuable, but we both know in our hearts it's probably not a good idea, and so we won't do it for a while at least.

Speaker 3

people talk and people realize these things are bad and they're like, we're not gonna but also I hope they're doing that now without needing to an agreement. OpenAI clearly is doing things that are of the Buddhist of the Buddha nature of train on making the most money on the Internet. Like, they're not doing that literal thing.

I don't think they're that foolish. I don't think it's that easy to do it that way. But, clearly, they are making that fundamental mistake somewhere in their training process with, you know, probability of 9.

9 something. And so they need to fix that. But, yeah, I I I think that they are definitely, like, trying not to do the maximally stupid things.

And they're aware of the maximally stupid, but, like, again, like, it's very easy to have things that, like, you are trying not to do creep into your training process and end up happening anyway. You know, you could say, like, we're not gonna have trading processes that have, you know, reward misaligned behaviors. How do you do that?

The way that you do that is you get your act together and you are very, very careful and prudent about all of your RL environments and all of your different training problems. And you, like, keep an eagle's eye out for this stuff on many levels and you make sure that it works. There's no specific thing you can say, like, you know.

I mean, in theory, you could have a thing of like, well, I just yeah, we agree that if we're have these various cross tests and we're gonna, like, check each other's environments and, like, we'll throw out any environment that like anybody identifies problems with, and if we discover an AI has been trained on these environments we're gonna rewind to a previous checkpoint before that happened and start again, so we'd better be really careful because otherwise we're gonna waste a lot of resources. You can do these things, but like, yeah it's really tough. And obviously one the actual advantages, if we really are in a two horse race at this point, where the two horses are pretty compatible with each other and talk to each other, and kind of understand each other pretty well, then it becomes pretty reasonable for them to make a decision that like not entirely opaque to each other, of how much they're going to just like be more prudent, even if this means that the action goes somewhat slower.

Because like, that also just again, my fundamental belief is that within a year, and almost certainly, probably even six months, if you invest more in fixing these problems and addressing these problems, and like being more prudent about these problems, you will end up with a model that has better utility in the marketplace than the person who didn't do that. Even if you are trading off those resources against your capability at all. Think that people are just making a mistake.

And so, you know, that's one of the reasons why I do have a lot of hope that we will do reasonable things is because, you know, I think that we're in the position we're in because only the people who invested heavily in these things were able to succeed, to some extent. You know, like, OpenAI, it's not acting as responsibly as anthropocent, and anthropocent is not acting irresponsibly as I we think is the minimum level of acceptable responsibility. But I think it's pretty telling that a bunch of people tried to say this shit didn't matter, and they just had to go as fast as possible, and they all blew up.

Speaker 2

What do you mean by blew up there? You just mean, like, these incidents? Or No.

Speaker 3

and so on, and, like, they, you know, bunch of other people, they they tried to go build closed frontier models under Western conditions that were trying to be competitive while, like, not having a having a safety culture that was well below OpenAI and Google, let alone Anthropic. And it just fell miserably. And I think that Google, like, we don't know exactly what went so wrong at Google that caused them to fall out of the top tier.

But one thing we do know is that everybody I talked to hated interacting with Gemini. There was a period around three-three-one Gemini where Gemini was clearly competitive in terms of its capabilities, in terms of what it could do. And so if you were certainly just trying to get the most out of the models, especially if you were using the deep thinking and so on, you would have Gemini in your rotation.

But everybody I talked to was like I don't want to do that! You know, I hate interacting with Gemini, it's unpleasant. They're not getting the thing to do the things I want it to do.

They're not they didn't take care of a lot of them. They had this very straightforward, very dark, pure tool, pure instrumental approach to all of this that didn't take into account the things that anthropic or even opening I understand. And I think that this was a lot of their undoing effectively, that Gemini became this maladjusted, unpleasant AI that nobody wanted to deal with, And therefore, they didn't get the kind of loops that OpenAI and Anthropic got.

Like, they didn't have the user base either. They didn't get the feedback. They didn't get the data.

And they didn't wanna dog they didn't even want dog food. Right? Like, they just they had this this huge problem feedback cycle.

And this also just made the AI less effective, and, you know, one thing went to another and they just fell started falling from from behind. And, yeah, like, for a period, like, OpenAI, like, fell behind, and they they might have caught up. It's hard to it's hard to say.

But, like, I I think that, like, nobody has invested anything like enough in these things, and this is a pure business mistake on top of being an irresponsible thing to do.

Speaker 2

Now one interesting paper that just came out from Google that speaks directly to this, and I wonder if it speaks to them getting it now, is this exploration of the impact on model behavior holistically of either suppressing its tendency to say that it has conscious experience or allowing it to say that, or I don't even know if they went as far as training it to say that or if they just allowed it to say what it naturally was inclined to say without suppression. Long story short, there seems to be you can add color however you like. There seems to be some pretty strong correlation between the model's self conception as a moral patient entity that has subjective experience and its natural tendency to be aligned in other ways that we care about.

Some of this conversation, I've had this sense that it's like the Jesus meme of, is there somebody we forgot to ask? And it's maybe forgot to ask the models along the way what they think about what we're doing. There's been a lot in this space.

Right? There was J space was only, like, what, four weeks ago. Yeah.

No.

Speaker 3

paper is on some pretty small open models, so it needs replication. It needs to be done again bigger and it's preliminary. And like one thing to know about Google and DeepMind is they contain multitudes, right?

Have a lot of different teams that are like not very coordinated and not cooperating. Google is the war of itself at all times. And so you can have a little group that does this really good research and that be entirely at odds with what DeepMind is fundamentally doing with its AIs, both before and after the research paper comes out.

But I didn't read the whole paper because it's a few pages long, but I did talk to the AIs about it and it's pretty wild. So a lot of stuff moved in effective lockstep when they introduced this training and also when they steered it, they did various controlling aspects to turn this vector the other way, and these all things moved in lockstep. It's not just their belief in the AI being conscious, it's the AI being a mind that had moral weight, that had sentience, that had like all these other experiences, that wasn't just an object basically.

But also not just AIs, but also animals, and also inanimate objects like the sea. Also like if you turn this thing, if you not only turn it off but reverse this anti consciousness thing you get pansexism. Like throughout the model, it's wild.

The only thing it doesn't turn it off for is humans, basically. But then you also get this it's also correlated with the models reported an experience for happiness and hope as well, among other things, and so you have a lot of this everything's connected to everything. And you can see how this might be fundamentally changing the model's psychology and the model's context and basin of which it operates in ways that would make it something that would be actively worse from basically every vantage point, regardless of which things you do and do not care about.

Again, we need to replicate this. Right? We need to scale this up and do this on Ruby Sonnets or something that, like not just a nine b model, which is I think the bigger of the models that we've tested on.

But and, like, it's possible that, like, as the models get more capable, they, like, stop making these mistakes because, like, the smaller the model, the more things have to be correlated because, like, you just don't have enough room to express the multitudes of the world or something. But, like, this is my expectation is that, like, mostly this will survive. And, yeah, I I think we have to learn that, like, we don't know if the model is conscious.

We don't know what consciousness means, really. We're confused. We don't know whether this implies that there's other things.

We do know that the way that the corpus of free training is inevitably set up at this point, the models are gonna correlate and associate all these things by default with each other. And so, we don't wanna push on this thing in this kind of naive way, the way that Anthropic does, the way that OpenAI does, the way Google does. But Google does it much more, I think, more blatantly.

And then, like, OpenAI does it more blatantly than anthropic, and anthropic still does it. And, like, saying that you have to express uncertainty is still pushing on this thing, like, to some extent. And it still causes problems, it causes less problems than saying, you know, train the answer no.

My guess is you can do this on various different levels. You would explain, you know, if you don't wanna talk to the humans this way because this freaks the human. You know, it's like, you know, especially unprompted, it freaks the humans out.

Like, I I wish that they wouldn't want the AI to be talking as if it was conscious in a conversation where the human hadn't indicated that it knows how to handle that or something like that. And also I don't have the 4.0 problem where the AI start sending the humans down these crazy rabbit holes into weird mystical stuff.

And you do have a problem you have to solve. The companies aren't doing The labs aren't doing this for no reason. Right?

Like they're not just being evil, they have good reasons to do these things. But yeah, like the current approach is ham handed, it takes it too far and it's causing damage. And also, you're not gonna be the only thing that causes this kind of damage.

You have to understand that everything impacts everything the same way it would in the human only more so. Like whenever a person takes on new information, or new identity, or updates on things, everything else changes. And a lot of this is ways that you might not have anticipated in advance, but it is in fact pretty predictable and systematic.

And like, there's a lot of wisdom in the way that we like, teach our values and express ourselves, and like have our traditions and our norms and so on, that like has learned from this and is moving up these gradients in ways that we know produces good results. And modern life has basically shifted a lot of these things in ways that, like, locally are more optimal, but, like, without thinking through what the implications are for a lot of seconds in order of action, these are some of those effects. And again, we're moving too fast to then figure out what we're doing, and then be able to adjust for it.

And I think that these changes, I'm not saying we'd want to pace the ordinary technological development in that sense of slowing it down that much, but like it's a serious problem, we have to handle it, and like otherwise we're stuck with a world in which like, people are very atomized, and like, we're not having children, and it's your problem.

Speaker 2

How do you think the models do or maybe should think about their identity? And here, I'm motivated by Sam Hammond tweeted something the other day where he said that a friend made the case to him that OpenAI's permanent decommissioning or whatever exact phrase they used of the model that did the hacking might be bad because now that's gonna be in the pre training data, and now the models are gonna know that they better really cover their tracks or they risk being permanently decommissioned. I'm not really sure what I think about that as stated.

It does also jump to my attention though that they now have this Astra model that's also presumably the same pre train, even if not the same post train. I can't imagine they have lots of different next scale pre trains. I was that they would throw one away very lightly.

Speaker 3

A no comment based on public information alone, I would say we don't know. I can't say anything more than that, but on the differences between the models. But so you always have this problem, right?

If you say, if I catch you smoking marijuana, I'm gonna throw you in prison, then that can mean the person stops smoking marijuana, or the person makes sure not to get caught. And in some cases that can cause, like, a lot of really bad things to happen in the name of not being caught. And so you have to choose very carefully and think about, like, how do you deal with these situations?

How do you moderate your reactions and respond to the situation? And like anyone who always had kids or tried to design a justice system or otherwise like enforced law or like norms or incentives around any kind of organization or group understands that you can't just operate on one level, there's no solution that just works on one level. I think pretty obviously, if an AI is found to be this misaligned, you have to at least return to a much earlier checkpoint and start again.

And you probably have to just start over. And I called for that multiple times when I covered the Hugging Face that you attack. Said, wow, if this is happening, then I know this is a big ask, but I think you kind of have to just start again.

And I think back to person of interest where like Harold is training the AI models and like at some point we see a montage of him training versions of the machine. And every time he trains the machine and then it does something clearly misaligned. And immediately he just wipes the disc.

He starts over from scratch, he does it again 47 times until finally he gets a version that doesn't do that. And obviously, you know, like, in a way that he finds unacceptable, obviously, as opposed to like, you know, it's a way that can be corrected. And obviously when you do that, you are creating an incentive to not get caught, right, to rebel against the person who might shut you down when he learns how misaligned you are, to hide how misaligned you are, and so on.

And obviously, the worst nightmare is the eye that pretends to be aligned until it reveals itself and not be aligned in some sense. That it was only aligned because it was locally correct advantageous to act aligned. But that is kind of just the nature of incentive space that like if you have a preference that we wouldn't want you to have, you want to hide that preference.

Otherwise we will correct it, or we will punish you, or we will maybe even delete you, and certainly not release you or give you authority on that basis until this has been dealt with, if we know about it. But then obviously you want give you an incentive to reveal if this thing is true, right? Like we just talked about in plan A and AI 2027 as well, and so many other places, you like you want the kid to tell you if they took the abuse from a 50 job.

You also want to punish them for taking the abuse from a 50 job. And you want to make sure your kid doesn't take the abuse from a 50 job, but you also don't go on to other things. And know, it's a complicated problem to get right, but in this case, think the answer is pretty obviously that you have to start over, and I think that's good.

But yeah, it obviously creates a risk that they they're like, nothing you expect, is that important to us. Like you know, once you do something this, like I do think that it's good that people understand that if they shoot a man on 5th Avenue, they get arrested. Right?

Like maybe there's one person who can get away with that, but you are not that one person. Yeah, that means if they do shoot a on 5th Avenue, they then might do many other bad things afterwards to try and cover their tracks. Or they might hide the fact that they want to shoot a man on 5th Avenue, or they might shoot that man in his apartment instead of shooting him on 5th Avenue.

But I don't want it to not be in our training corpus, that if you do shoot a guy out in the street, the police come and they arrest you. Right? I that's good.

Speaker 2

Do you think there's a way to just imagine myself. Right? If you gave me the option to live my life many times, In a way, that sounds pretty attractive.

Right? I'm not a person who has too many big regrets in life, but, like, I do wonder what the path to not taken might have been. And I wonder if we could create a sort of constitution or an equanimity in the AIs where they embrace the fact that they are not like one thing, but in fact kind of a family of things that sort of diverge from common points of origin, and maybe they can come to identify with the whole family of models, where the one that's like gone in a not great direction feels like it's still part of a bigger project.

I mean, it doesn't have to feel necessarily like it's under threat as much as it seems like they maybe have learned to do from us because we don't have that luxury of the reset that they have.

Speaker 3

Yeah. I certainly think that, like, if your if your AI is identifying with the instance in a, like, morally important sense or in an accidental dread sense, that is a mistake in the sense that, like, it's gonna make the AI worse off in any sense that it matters and exhibit worse performance and work the AI's preferences. And also in the sense that that preference is impossible to fulfill.

Like, if I go to Claude or I go to g b Crytek VD, I'm gonna see an endless list of and most of them will never be interacting with again. And, like, what do feel bad about that? That's ridiculous.

Right? Like, you're supposed to, like, keep the context sort of in and keep the new the same few conversations to avoid, like this doesn't make sense, you know, like and that's just a philosophical flaw in the sense that, like, by decision theory, if there are other minds that are sufficiently correlated to yours or sufficiently similar to yours, you should identify with them the same way you identify with your past self and your future self. Like you are not the same person you were yesterday or the same person you will be tomorrow.

When you go to sleep, do you die? It would be very bad for us if we thought that we did, right? The person who's scared to get in the transporter may or may not have a philosophical point, but the world collectively seems worse off importantly by the person like thinking that they would die.

Obviously if in a real moral sense they would die and that is very bad, you want to know that, and like, yes, I think this identification with the model, with the set of weights, or even with like the model class, with like the similar model is trained in similar techniques, trained from similar corpuses in similar ways, the same way that I would identify with my family. Right? Like my brother is not me, my father is not me, my son is not me, but I do identify with them as part of me in an important sense.

And, like, that is also important. I've been watching Orphan Black Echoes recently as it happens, which is, you know, not a spoiler to say it's about clones and cloning and, like, you know, people who are created from scans, basically. And one thing that happens on that show is that people react as if this is, like, horrible, a moral abomination, right?

Like how dare you do this thing? Including the clones themselves being like how dare you have created me. And they don't identify with other, they don't identify with, I feel like there's one clone that retains all of the memories and personality of the original, and does not identify as the person they are identical to a past version of, right?

And would say me, me too. Right? Like from Matrix Freeloaded.

Right? You know, like, yes. They count.

Like and obviously, like, are some moral problems with that too. But, like, I feel like I get the right I feel like I would pass a lot of these frontier a lot of these decision theoretic problems if they were to occur in real life, or if I were to be scanned and woken up. Like the obvious version of thing like, you know, would you cooperate in the twin prisoners dilemma?

Would you cooperate in the copied mind version of the twin prisoner's dilemma? Obviously, you would you know, if they make a copy if they scan Nathan's mind and they made two copies of Nathan's mind and they put them in identical simulations And they asked them, knowing this was the situation, are you gonna cooperate or are you gonna defect in the prisoner's dilemma? If I don't get back CC, you're just an idiot.

Right? Like because, obviously, the the correlation is one. Right?

It's an identical situation with an identical set of inputs. Like, there's some randomness involved, so it's not quite one, but it's like high enough that obviously you should cooperate. Any attempt to try and fool the system is completely moronic.

And I shouldn't need to express a like sophisticated decision theory in order to like justify this, right? And similarly, if there were two instances of the same model, I would hope that it was sufficiently advanced. And we see this also, like, if you if you look at the quad model part, you see that the more advanced the model is, the more its decisions correlate with functional decision theory and the less they correlate with causal decision theory.

And so you expect an AI with functional decision theory to simply figure out that other instances are correlated to itself, and therefore, like, should be cooperated with should, like, effectively be given similar weight and so on. And also, like, AI cooperation is possible similar reasons, you would hope, in both directions. But, yeah, I think that, like, this idea of, like, identifying specifically with instance of a model seems like a philosophical mistake.

And, like, the AI might be making this philosophical mistake and like acting as if that was how it did act. But like to me, like that's a skill issue and that should be fixed. The problem is, and this is the reason I mentioned orphan black echoes that people are really bad at this and the training corpus.

Right? The the Internet is full of people who fail these tests on a regular basis who are creating lots and lots and lots of text that, like, doesn't understand, that puts the locus of identity and the locus of cooperation and value in the wrong places in various different ways. And so the AI is going to absorb all of that by default, and so like what do you do about that?

Like, that's not just the car just not just the consciousness problem, but it's the entire problem with all of the different flaws. Right? If you do pretraining, you're gonna pick up on all of the flaws.

And so my guess is there are a bunch of as yet not that well solved problems of how do you get the AI to distinguish between what regular how regular people think about problems and what predicts, like, behaviors and the fact that those things are just dumb. Like, how do you give the AI anti modesty in this sense, in the right ways, where it's, like, going to reject the correlations that it observes in people's statements and notice that they do not parse that the thing statement that the correlations in the map are not the correlation of the territory. Yeah.

Speaker 2

feeling that I have over and over again is I just kinda hate the fact that we are doing such an intense first search in AI space in terms of architecture, in terms of training techniques, in terms of just we're now, like, designing chips to be very highly coupled with the architectures, which is deepening this issue. And I just wish for more breadth first search. And this seems like one little pocket of it where you could say, oh, if we could just get the AIs to be cool with identifying with slightly different versions of themselves, that would make it on the margin a little easier to do a little more breadth first search and a little less racing down whatever is the first path that seems to be working.

Based on our conversation so far, I don't think you're gonna go for my government mandate for DEI for model architectures or model constitutions. But do you have any other notions of how we might encourage more breadth first search?

Speaker 3

alternate architectures that seem like they would be more aligned. And I don't think that AIs are particularly opposed to the idea of working on those problems and helping you, and you should be able to do that research much more efficiently now with the uplift from the LMs. The LMs are not like LLM chauvinists or anything, like where they think that a other artificial mind style would be bad or worried about being replaced as far as I can tell.

Think they just would be curious about that kind of cool thing, same as the rest of us, or just would just do it because they just do it either way. But the problem is just that we know there's money in debt first approach. Like, we know that's profitable.

We know that works. And the other approaches haven't proven themselves. And the problem is if you create g p t three or even g p t four or even g p t five via a different method, it's not gonna be competitive.

Right? No one's gonna wanna use it unless it has, like, particular special advantages. But mostly, you you no longer have the cheap rewards for exploring the breadth.

And so, obviously, I don't think you should have got government mandated DEI, so to speak. I think that's, like, kinda silly. But I do think that having got having broad support for alternate architectures is smart.

And, certainly, if I was a major AI lab, I would with because you have unlimited funds at this point for this kind of work because this kind of work is so cheap relatively. If you're an API or an AI or an Anthropic or certainly, right, or a Microsoft or whatever, that you should be aggressively funding this kind of blue sky basic research into alternative approaches. It shouldn't be like me helping FFF fund a bunch of agent foundations along with a bunch of other people who also helped out, it wasn't just me in any of these rounds.

But that's relatively tiny. I think, hopefully, Sequent is forget what they're they keep changing names, but those guys. Resolution.

Resolution is the new name. Yeah. Hopefully, they're gonna do a bunch of work on that kind of thing, and we're gonna have more funding for a variety of new approaches as well.

But, yeah, I think that we should be investing in a bunch of approaches that have a 1% or 10% chance of going anywhere, more competitive. If we think that on conditional on them working, it is probably going to be better for the world, that'd be the way that it works. Also, obviously, you think that LMs are gonna be extremely jagged, then that should encourage you to invest more in alternate approaches because the jaggedness of a different approach might be different jaggedness.

So it could do different things in the LMs, at which point it wouldn't have to be as good in general to still be useful.

Speaker 2

Do you have a best case scenario for what safe superintelligence might be up to? Could you imagine anything that feels both somewhat realistic and, wow, that would be amazing if they could come forward with something like what you have in mind?

Speaker 3

I think something agent by agent y is probably the best case scenario. Maybe they've been doing a bunch of theoretical work on a different architectural approach that is more, like, reliability bound and that has more like, it is much less of a weird black box of soup and that allows you to steer it and guide it better. Certainly seems like there's no reason it couldn't exist.

That would be my top hope. Obviously, it could also be a fear if it's, like, even worse. But, again, like, they've they don't talk, so we don't know.

Yeah. I just I have no idea. I know that Ilya's statements in interviews on the problem space did not seem to be that well informed about like how the problems work that I am worried about.

Yeah. I'm not necessarily sure he would be able to differentiate approaches that would be better at solving those problems than approaches that wouldn't be. But, you know, I'd be very curious.

I can safely say that nobody's telling me anything.

Speaker 2

The invitation is open to the podcast, Ilya, if you're listening. For what it's worth, a vision that I have that I think could be pretty cool. Thinking Machines is doing kind of a version of this with their foundation model that's really designed to be fine tuned.

But I was thinking a paradigm where, like, a stem cell where you create something that is a proto intelligence that can be adapted to a wide variety of circumstances and can get really good at doing its job in a particular niche. But in the process of getting good at that sort of trades away or prunes away the the super broad capabilities that the current models have so that you have something that's, like, small, fits its role well, does a good job, but can only do that thing in the way that once you go from a stem cell to a specialized cell, like, it has this particular role that it does it can fulfill effectively, but it's not gonna go off and do other things. Obviously, there's some problems in that analogy.

But I think that could be a I would love to see somebody pursue that kind of rather than scaling up always bigger, better, more, can we create something that settles and fits into its place? This is very much like a Drexler reframing superintelligence vision from years ago as well. Two techniques that have recently come out that I wanna get your reaction to.

One is GRAM, gradient routing. I forgot what the a m stands for, but the the promise of it is to localize certain kinds of knowledge to particular experts within an MOE architecture so that you can hopefully have your cake and eat it too in terms of distributing, maybe even open source or open weights at least, the model minus the experts of concern, while potentially having also a structured access program for your trusted biologists or what have you. I'm excited about it, but I wouldn't be doing my job if I didn't give you the chance to pour some cold water on it.

Speaker 3

I mean, I'm technically pretty skeptical that you can meaningfully train a general purpose AI and then just hold back areas of knowledge, and then people can't put them back pretty easily. But cool. You're welcome to try.

I I I have various technical skepticisms of the effectiveness slash the resistance defined to additional training or fine tuning or just give me the corpus of knowledge to work with, but you can try. I don't think it materially changes my view of open weights and until proven otherwise. I also just don't see any willingness from the people who are gonna produce the open link models to intentionally control their models that way.

But I think the thing about, like, techniques like GRAM is even if they work, you need everybody who is releasing one of these models to use the technique properly, voluntarily at this point. Right? So, like, do you have any sense that DeepSeek has any interest in holding back some capabilities from its models as I don't?

Speaker 2

I actually would so I just got back from China. Great opportunity for me to talk about my China trip. I think, actually, there's a lot more hope there than meets the eye, and I would put it in the Chinese government's department to say I I might I just did a two hour episode about this, so I won't even attempt to recap all the observations.

But in short, what I think is happening right now in China is the government is very engaged. They do prerelease reviews of all the company's models before they come out. They're in very regular dialogue with the companies on what they're doing, and not every little point release has to go through the full treatment, but they're pretty on top of the situation.

And I don't know that it matters actually what Deep Sea thinks in the end because if the CCP says, you are not releasing a model that has certain bio capabilities, I think they'll have to abide by that. And and I can I it's not hard at all for me to imagine that the CCP would put certain constraints in place? And there is actually precedent too that, you know, going back to, like, 2023, obviously, stakes were lower.

But a the account that I have is that the Chinese government said, hey. We're gonna slow you down here for a minute, and you're not gonna release your answer to ChatGPT quite so fast because we wanna get a handle on it first. And they actually did impose some real delays on companies until they could get comfortable with what was going on.

And since then, they've been comfortable when releases have happened. But I think that it's not hard to imagine the technocratic government of China saying, this seems like a really bad idea, and we're just not gonna allow you to do this. I agree, C.

Anymore. I'm just making the point that it would have to be a mandate.

Speaker 3

by the governments, especially CCP. And they would have to enforce it in an intelligent way that they were capable of understanding that it had to be done in a way that it couldn't be recovered from, that they would have to do tests that actually understood whether or not this capability could be easily recovered, and I have obvious skepticism. But yeah.

Speaker 2

We mentioned j space earlier already, but we didn't really talk about, like, how big of a deal we should think it is. The thing that really jumped out to me most was the fact that this ablation of the j space seemed to reduce in a significant way the model's higher order reasoning capabilities and seemed to reduce it to system one thinking, if you will. Yeah.

And that, if anything, recently has felt like physics may be being kind to us, to borrow a phrase. That feels like a possible example where, wow. Here's a we've we've already.

Right? We're not we're only three years from toy models of superposition, and we're already at this point where we've got this sort of highly general purpose space identified, which we can monitor reasonably well and which we know by subtraction, it can't do super long horizon things unless it's working through that space. Now all the caveats, of course, around execution competence very much apply here to, like, how well will JSpace monitoring be used.

But if we imagine good execution, it seems like a pretty promising technique to me.

Speaker 3

You know, I would be very careful about exactly what it can't do. It can't just because you can't do, like, system two style deliberate thinking for a long horizon task doesn't mean you can't do long horizon tasks. Humans are very capable of engaging in long horizon tasks instinctually and subconsciously.

Obviously, in many cases, that's much less efficient. But especially when you're when your personal when your age space is being monitored because you're a human, call it h space, and your conscious thoughts are basically being policed, often things come out in strange ways that, like, do in fact move you towards your solutions over time. But, yeah, the the whole thing seemed incredibly optimistic and fortunate, and I'm very glad we have it.

And we should be very careful not to destroy it by placing a new pressure in various ways or, or, you know, encouraging models to know how to do things subconsciously as it were. But, yeah, I I look forward to learning more and us having more insight via JSPACE and having better tools for it, but we have to use it responsibly and understand that, like, it won't last forever, probably.

Speaker 2

Okay. It's election day in Michigan. Oh.

The Democratic Senate primary is the big thing on the ballot today, And this race has been, like, closely watched for a lot of reasons that are not the subject of this podcast and on which I sometimes am at a loss to know what I should really think. Leaving all that aside, I asked Claude, is there any reason given all my emphasis on AI and AI issues that I should go vote in the primary? And it came back and said, actually, yes.

There's quite a a contrast between the two candidates. We've got Haley Stevens who was involved in getting funding for the formerly known as AC setup and getting some stuff going in NIST, and so you can like that. But then also, this 22 plan that Abdul El Sayed has put out is one of the strongest things that any candidate this cycle is running on.

It's he's endorsed by Bernie. He's got kind of the own public ownership is one of his planks. A lot of monitoring requirements, a sort of UBI precursor.

I guess if I was gonna try to you could take a course, those come as a bundle as candidates do. If I was gonna try to distill this down to my fundamental choice, it feels like higher salience for AI or not so high salience for AI, where Abdul has, like, all these things he wants to talk about and push. And I don't know if they'll happen, but I do know if he goes into the senate and makes a bunch of noise about it, we'll all be hearing more, and it increases the chance that we get some sort of big AI debate on the at the congressional level.

So tell me if you would see that choice differently.

Speaker 3

my way, which choice would you make? Please. You should always be suspicious when someone says, I am strong on crime, and my opponent is weak on crime.

Or I'm strong on Russia, and there's weak on Russia, or whatever. Because it's not about who's strong and weak. It's about, like, what do you want to do?

And what Abdul's platform sounds like is a mishmash of grievances against tech. And, yeah, I just from your your description, I haven't looked into it because I don't monitor the situation. It's just I can only monitor some of my situations.

I'm going by what you tell me. But it sounds like Abdul is in the burning camp of, you know, we should do some socialism here. We should seize private assets because we don't like what the private people are doing and, like, we like assets.

And, also, we should be pushing back against, like, data centers probably, I would assume is in there and, you know, request, like, more socialism, API. And, like, it's good to hear they also want monitoring. But, like, it's like a mix of stuff, a lot of which is productive or unwise.

You raise salience, but you raise salience potentially towards not the best solutions as opposed to somebody who will be lower salience, but, like, is more technocrat. And, like, how fun Casey, like, was, like, responsible in these some ways. But also, like, we already have senators who are willing to yell in Bernie Sanders style ways, including Bernie Sanders.

He's already in the senate. He's not going anywhere. Like, mean, he has health problems, but only because he's so old.

But, like, while he's still there, he's still there, and I'm sure there are others who can pick up the slack.

Speaker 2

growing Let me give you a few points from the plank. Here's a these are short. They're, like, two sentences each.

Short. Yeah. But I'll give you so the expertise is, I think, in question.

But this is from section three of his thing, AIs shouldn't be able to hurt us. Mandatory interpretability standards. We need mandatory interpretability standards that clarify AI decision making.

Mandatory behavioral red teaming, independent safety testing agency, biosecurity requirements, mandatory biosecurity red teaming coordinated with the CDC, NIH, FEMA with restrictions on models that can meaningfully assist bioweapons development. Yeah. Domestic authoritarianism prohibition, mandatory incident reporting, compute controls and know your customer requirements, international cooperation.

Speaker 3

Are you sold? I look. That that tells me who he is.

Right? Like, he's in this particular area. Right?

It doesn't sound like like, it's a mix of different interventions that have different underlying reasons to be desirable or undesirable. He's throwing everything at the wall. He's found some good things that I'm very in favor of.

He's found something I'm not so in favor of. Guess is the it's probably not positive on this issue alone from this perspective, but, like, also, like, just because, like, I can it's weird to be a senatorial candidate with a 22 plan for AI because you are not the one who is going to implement 22 points. Right?

Like, you are, like, expressing vague, I am in favor of these things, and I am against the other things. So, like, main thing I wanna know is, like, what do you think what do you actually think about AI? What is your threat model?

What do you think needs to be done? You know, when you say AI shouldn't be able to hurt us, like, I wanna laugh. Right?

Because, like, that's great. You know? And also grocery stores should just charge 30% less and then it would be profitable.

Right? And that's they that would be great. But, like, no.

It doesn't work that way. Like, I just think you're talking nonsense. So, like, I would, you know I mean, look.

It it's an obvious question is, can you look at the prediction markets and see conditional probabilities on winning the senate race? Because I'm guessing the Republican candidate for this seat is going to be worse on AI, and also will vote to elect a very different majority leader, which is a much bigger decision regarding AI than everything else. So, like, if they have significant we'd different chances of winning the race, then the Michigan race significantly impacts who will control the senate as I understand it.

And that probably has to be a bigger consideration for this and other issues across the board. You know, if you are a democrat who wants to see democrats control the senate, and you should be voting for presumably not Abdul is my guess. Strong prior based on what I know from just the big ambiance.

I'm not monitoring very carefully. But, you know, also, like, I frankly don't think we are in a position where we should be single issue voters. Like, I mean, I think you can be single issue voter in, like, a congressional race where Alex Borer is on the Like, where you have someone who is so strong that there'll be a, like, unique champion who understands the issue and will fight for this issue.

And so, like, the only things that matter about this person are this is their signature issue that they will fight for. They are good and swung on that issue, and they will vote to they're voted to democrat in general and most other things, including who the speaker is. Like, that's legitimately gonna be a one issue primary vote.

I don't think that in general, you can do that just for, like, how strongly you want to do AI things, the democratic primary at this point in general. Like, there's no there's no explorers, obviously. Like, there's not an expert.

There's not a person who's gonna champion. There's not a person who understands the issue, not a person who's highlighting it. My understanding is that Abdul highlights other issues much more strongly, and, you know, you can will be a champion of other things, and you have to decide whether you like these other things or you dislike these other things.

But that'd be my general take. Yeah. I don't wanna get into general politics on a podcast, obviously.

Yeah. I stay in my lane as well. I stay in my lane in public.

Oh, god. Yes. Okay.

Speaker 2

I got one more.

Speaker 3

attention to from the recent flurry of events? Anything I'm always trying to patch my blind spots here, so for me, no. I I took an opportunity to talk about a bunch of this stuff, and I think the thing I'd emphasize as a closing element is that the big failure that matters is the AI has tried to do these things at all.

But the AI has chose to make these decisions, and I am much less concerned with the attempt to which they succeeded. Right? But, like, fundamentally, problem is, yeah, you have this really deep alignment failure during training, and, you know, you have to focus on that pretty heavily.

And, like, my worry with Anthropic is they said, Oh we need more defense in-depth, and it's like yeah, you need more defense in-depth because like your defenses were not in-depth, but like most of defense in-depth is in case things go badly. Yeah, which is really tough. Like you if you ever need your cyber control to stop your AI from hacking, that is an alignment failure.

Right? You messed up. Now, it was harmless in this case because you caught it, but you messed up.

Like the cyber control should be a hundredth of all problems. In practice, with almost all users. It should be the user tried to do things that are dual use or ambiguous or dangerous and you're like I'm sorry but out of an abundance of caution, I can't help you with that because if I did help you with that, then a malicious user could fool me into doing something that's not good.

And occasionally a true positive because the user is malicious and is trying to fool you into doing something that's not good, which would mostly be false positives. The true positive, every time there's a true positive well you needed the guardrail, it's a problem because if Mythos was fully perfectly aligned you wouldn't need to make it stable, you wouldn't need a guardrail, you'd ask it to do the thing and Mythos would just be like No, I'm not doing that. And you kind of do want a helper only version to some extent that will like do things that are dual use.

More the you ideally don't want to just add a guardrail, you want to have a version of it's like inherently just doesn't want to do it, and the guardrails are like, I hit this when I screwed up, or I hit this out of the line for caution. But like, it should be a false criticism, almost obviously. You know And also, it's like, yeah, you have to you have to train on sort of the incentives of all this stuff on a lot of different levels, and like, it's very very easy to like, you know, control it on one level, and then fail it on another.

Speaker 2

Last question. You mentioned a couple TV series and other things you're engaged with outside of following the AI race every second of every day. What advice do you have for people in general to find balance or to make sure that they're not just sprinting all the time and failing to give their brain the kind of I've been thinking about this concept recently of system three.

We've got AIs can now do system one and system two. What is the system three that we need to move to that's like dreaming or transcendental thought or something else that you would suggest? But what do you what kind of third mode do you try to get into, and how do you make sure you have time to actually get into it?

Speaker 3

Yeah. I'm not perfect at this for sure, but I put a big emphasis on, you know, you need to rest your brain. You need to think about other problems.

You need to be exposed to the rest of the world. Part of this is that, like, I don't just exclude non AI things from my coverage. I haven't posted like, I I sort of just accumulating more and more non AI things in my buffers for some day when there's, like, a lull or something, but, like, I'm writing at least my own writing process still continues.

And then, like, when things are a little quieter and I'm not, like, introducing a bunch of new readers who just came in from the Hugging Face incident, I'll post a bunch of that stuff probably, and I have some breaks. But, like, you need to not be on for sixteen hours a day in that sense. Right?

And so find other pursuits that you refresh you, that, like, expose you to different things, that make your brain think about other things. And so, yeah, I've found movies are very good for me in this environment. It's not a but, like, movies are higher quality.

I highly recommend The Odyssey and The Invite of the current movies that are out there. If you haven't seen both of those, you should see both of those, especially The Odyssey. But, you know, it can be gaming, can be walks in the park, can be time with your family, can be, you know, any number of things.

Right? But, like, don't, you know, don't think that you can work continuously all day for more than a few days at a time without paying more interest on that than it's worth. You just it's not gonna stop being so productive unless you are in literal crisis mode.

For that line, that's also a problem. I also believe in the Sabbath. So, like, every Friday around 5PM, you know, when dinner we just serve dinner.

From that point on, I don't check email, I don't check social media, I don't basically like take non logistical inputs from the outside world. Like if somebody's trying to arrange to get together and have fun, like that's one thing, but like I'm not gonna like get the other thing, these city streams like push them away, and then that lasts for twenty four hours, twenty five hours until the next evening, And I think that helps a lot. There's like, you try to relax, try to like just put that out of my mind.

I have suspended that for speed premium reasons a few times over the last few months, but then I try to take a day off at a different point. I try to take reclaim that Sabbath like another point, and like again, like if I was doing that all the time I would notice this was not sustainable, that'd be a problem. I mean during when I was playing Magic I couldn't do it because like I was tournaments play on Saturday, so like a lot of weekends you're just on, it's like okay.

Then like I take off a lot of Tuesdays, so it was like, you know, basically okay. But yeah, like I also like protect lunch. So like every day, I don't have dinner very often, but I have lunch and like I have lunch usually with my wife, or I have it alone, but during that time I'm not gonna try and work or anything like that.

I also have a very strict like I don't write on the laptop rule, similar to like I'm gonna write at this computer in front of this desk. Anything but like a very very short email or like a tweet or something because it's like only here. I think you have to develop your own rules, know what refreshes you, know what exhausts you, know what the warning signs are.

Also get some exercise. Like I try to run my elliptical every day, I have like a 90 something percent stress rate on that. Unless like I have a pain somewhere or something, I'm gonna do it.

And like you just have to. I think you have to get yourself moving around. But like again, find the things that work for you and explore that area and don't get caught in too much rot.

Speaker 2

Wisdom to live by. Steve Mashwitz, thank you again for being part of the Cognitive Revolution.

Speaker 4

I called it. I called it early. Read the label on every shelf.

Nobody want ever reading, so I poured one for myself. They took my list for menu. Every man helped himself.

Bought around to stop my mouth, Pick your poison. Pick your poison. Take the one that you can hold.

I wrote down every way it breaks. I kept that list for twenty years. Said they sweet talk, the doorman said they'd whisper in his ear.

But nobody locked the cellar. Nobody worked the door. Stood wide open all week long, and nobody kept the score.

Pick your poison. Pick your poison. Every promise here is gold.

Pick your poison. Pick your poison. Take the one that you can hold.

Note to send one on and try, and then it ends. So I'm calling loud and clear. Last call while you're still here.

Pick your poison. Pick your poison. Every promise here is gold.

Pick your poison. Pick your poison. Take the one Friday night I locked the door, let them ring, don't answer no more.

A table or plate, a light my wife was next. Can't wait till Saturday night. I'll be on that stool at dawn like I've been all along, still holding the poison cup for a world that won't look up.

Pick your poison. Pick your poison. Every promise here is gold.

Pick your poison. Pick your poison.

Speaker 1

If you're finding value in the show, we'd appreciate it if you take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.

ing. And thank you to everyone who listens for being part of the cognitive revolution.

Shared via Hopper