Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Machine Learning Street Talk (MLST)
13 July 2026 55 min
0:00 --:--
Episode Description
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstBritain's most capable coding model can't be exported, and that ban is the whole reason Cosine set out to build one from scratch. Alistair Pullen, CEO and co-founder of Cosine, sits down with Tim Scarfe to explain how a frontier system he calls Fable, locked behind US export controls, became the founding case for a UK sovereign model trained on the Isambard supercomputer in Bristol.T

Summary

Alistair Pullen, CEO of Cosine, discusses the UK's initiative to build a sovereign large language model (LLM) in response to US export controls on frontier AI like Fable. He explains Cosine's strategy to compete with larger labs by focusing on efficient model architecture, advanced data curation, and innovative reinforcement learning techniques, leveraging the Isambard supercomputer. The conversation also delves into the challenges of AI-generated code quality, the evolving role of agentic harnesses, and the importance of synthetic data generation.

Chapters

Cosine's UK Sovereign AI MandateAlistair Pullen introduces Cosine as a frontier lab initially focused on coding agents for highly regulated environments, now tasked with building the UK's first sovereign LLM due to US export controls on models like Fable.
Leveraging Compute for Sovereign AICosine's ability to pursue sovereign AI is bolstered by government backing and access to the Isambard supercomputer, which provides crucial compute resources that would otherwise be a significant financial barrier for a startup.
Competing with Limited ResourcesAlistair explains how Cosine aims to build a competitive frontier model with significantly less capital than US labs by focusing on licensing technology rather than inference costs, and by employing novel algorithmic and architectural approaches.
Model Architecture and Data StrategyThe discussion highlights the importance of architecture (total and active parameter count), and data (pre-training, mid-training, and especially post-training with RL at scale) as key differentiators for achieving frontier model performance.
Mitigating 'Slop' in Coding AgentsAlistair addresses the problem of 'slop' (functionally correct but poorly written code) in AI agents, detailing Cosine's RL process that incorporates specific rewards and credit attribution to reinforce good coding practices beyond mere correctness.
Agentic Engineering and Swarm AIWhile Alistair believes agentic harnesses are becoming less critical as models improve, he emphasizes the value of sub-agents and introduces Cosine's 'swarm' architecture for orchestrating hundreds of agents to tackle complex, multi-faceted engineering problems.
Challenges of Agent Memory and DataThe conversation explores the inherent difficulties in implementing effective agent memory, noting current approaches as 'hacks' and highlighting issues with querying, relevance, and keeping information up-to-date, suggesting continual learning as a potential solution.
Synthetic Data Generation for RLCosine utilizes a sophisticated pipeline for synthetic data generation, creating ground truths and graders for real-world software engineering tasks beyond simple bug fixes, enabling robust RL training across diverse programming languages and stacks.
UK's AI Sovereignty and Future RisksAlistair concludes by acknowledging the commercial advantage gained from US export controls, expressing determination to succeed despite potential supply chain risks for hardware, and highlighting the increased urgency and support for UK sovereign AI.

Topics

Sovereign AI developmentLLM training costsModel architecture designReinforcement learningAI coding agentsSynthetic data generationAI ethics and qualityGeopolitics of AI

People

Sierra (mentioned) Indeed (mentioned) Alistair Pullen (guest) Tim Scarfe (host) Donald Trump (mentioned) Andrej Karpathy (mentioned) Francois Chollet (mentioned) Ben (mentioned)
Key Concepts (15)
Sovereign LLM — A large language model developed and controlled by a specific nation, free from foreign export controls, ensuring national security and economic independence in AI capabilities.
Inference vs. Training Costs — The distinction between the computational resources needed to run a trained model (inference) and the resources needed to create it (training), with inference often being a larger ongoing cost for large-scale deployments.
Model Architecture — The structural design of an LLM, including factors like total parameter count, active parameter count, and whether it uses a Mixture of Experts (MOE) or fully dense approach, which significantly impacts performance and inference ability.
Mixture of Experts (MOE) — An architectural approach where different 'expert' sub-networks specialize in different parts of the input space, allowing models to have a large total parameter count but a smaller active parameter count for faster inference.
User Trajectories — Sequences of interactions and prompts from users with an AI model, which provide valuable data for understanding real-world usage patterns and improving model training, especially for realistic prompt distributions.
Slop in AI Code — Refers to AI-generated code that is functionally correct (passes tests) but is poorly written, inefficient, or creates technical debt, often characterized by excessive lines of code or bad abstractions.
Credit Attribution in RL — An advanced reinforcement learning technique that aims to identify and disproportionately reward or penalize specific, important decisions or tokens within a long trajectory, rather than equally weighting all actions, to make training more efficient and targeted.
Epistemic Problem in ML — The fundamental challenge in machine learning of verifying whether a model's output is truly 'correct' or 'good' in domains where objective, verifiable rewards (like passing a unit test) are not readily available, such as law or abstract reasoning.
AI Psychosis — A term used to describe the phenomenon where humans lose understanding or competence when relying heavily on AI, leading to reduced ability to maintain, evolve, or ask the right questions about the AI-generated output.
Runtime Exploit Validation — A method used in cybersecurity scanning where an agent not only identifies potential vulnerabilities in code but also attempts to exploit them in a virtualized runtime environment to confirm true positives and reduce false alarms.
Agentic Harnesses — Software frameworks or systems that orchestrate and manage the interactions of AI models (agents) with tools, environments, and other agents to complete complex tasks.
Tokenomics — A term referring to the economics of AI token usage, particularly the cost and efficiency of tokens consumed by LLMs, which is increasingly important for enterprises.
Swarm Architecture — Cosine's hierarchical orchestration system for AI agents, where a top-level orchestrator decomposes a problem into sub-problems for 'sub-planners,' which then delegate tasks to multiple 'workers' (sub-agents) running concurrently.
Agent Memory — The ability of an AI agent to store and retrieve past information or knowledge relevant to its current task, often implemented using tools like VectorDBs, but facing challenges with relevance, freshness, and avoiding 'reward hacking'.
Synthetic Data Generation — The process of creating artificial data, often by AI models themselves, to augment real datasets or generate specific types of training examples, particularly useful for reinforcement learning where ground truths or graders are needed.
References (28)
Sierra company
Indeed company
Notion tool
Fable project
Cosine company
The Telegraph article
Isambard supercomputer project
Anthropic company
Colossus cluster project
Mistral company
Mistral Three Large (675B) project
DeepSeek company
Vertex tool
Sonnet project
Opus project
DeepSeek v4 Pro project
ClaudeCode tool
Outpost project
GPTOS s120b project
DevStraw 2123b project
Llama 70b project
OpenCode tool
ARC challenge by Francois Chollet project
Codex project
Kimmy k 2.6 project
Gemini 3.5 project
Mythos project
ChatGPT tool
Transcript (50 segments)
Speaker 1

Sierra has all the best active and outdoor brands for the super athletic stuff, like running gear for cruising up the trail. Whoo. And the super athletic stuff, like fishing gear for chilling by the creek.

Nice cast. Fitness apparel to push for higher reps. You got this.

And golf balls priced so you can afford to lose one. Or a few. Head to Sierra or sierra.

com for the brands you want at the prices that let you do it all. From athletic to athletic, Sierra's got it.

Speaker 2

When you need to build up your team to handle the growing chaos at work, use Indeed Sponsored Jobs. It gives your job post the boost it needs to be seen and helps reach people with the right skills, certifications, and more. Spend less time searching and more time actually interviewing candidates who check all your boxes.

Listeners of this show will get a $75 sponsored job credit at indeed.com/podcast. That's indeed.

com/podcast. Terms and conditions apply. Need a hiring hero?

This is a job for Indeed sponsored jobs.

Speaker 3

I was horrified as many were when Fable suddenly,

Speaker 4

got banned.

Speaker 3

The UK's first sort of sovereign LLM. How can you do in millions what they are doing with billions?

Speaker 4

There isn't a huge amount of room for error or a huge amount of wiggle room. It was not something that was on my bingo card in January. The numbers like 10 trill being knocked around.

We we we're compressing most of the Internet at that point. It's really funny.

Speaker 3

it it it'll say, oh, you're right to point that out.

Speaker 4

are getting less important over time. A model can probably do with bash only basically any task these days. Also, thank you, Donald Trump.

Yeah. Mean, for the first time, I feel like a second class citizen because they are going faster than I am, and I I really hate That boils my blood more than anyone else, and I we are going to do everything we can to pull this off. We have no choice but to make it happen.

Alastair, it's it's great to meet you, mate. We are here in London, where it's customary to say hello, Geezer. Hello, Geezer.

So where in London are we? We are in Hoxton right now. So we're in Shoreditch.

We're about half a mile away from where Cosign started in my apartment, which was in Hoxton Square, just over there. So we haven't come very far, but we have expanded a fair bit since then. And what is Cosign?

Cosign is a frontier lab based here in The UK. Prior to about three months ago, we built best in class coding agents specifically for highly regulated and high side environments, So think things like financial services, insurance, defense, and so on.

Speaker 3

which is a much more ambitious vision and something which is on a scale much larger than we've done before, but it's very exciting to be working on it. So tell me about that sovereign AI piece. Should say, by the way, so I read your, you know, the article about you in The Telegraph.

Marvelous. And I was horrified as many were when Fable suddenly, got banned Yes. Because of this export control.

And now everyone suddenly is thinking about sovereign AI. So tell me the story.

Speaker 4

well, it it it ties into it ties into a bunch of different things. It ties into, like, the backstory of cosign. And one of the reasons we're fortunately placed to be able to do sovereign AI is that we have a lot of expertise around model training, model building, all of the infra algorithms, data, people that you need to do that kind of thing.

And we've been doing that for some time. So we already had all those things in the organization. And then probably, I'd say nine, ten weeks ago, we were inducted into the government's sovereign AI units or backed by them, should say.

That is something that they have put out to, you know, increase the number of sovereign AI initiative, you know, companies that are that are being built in The UK. And what that looks like in practice for us is an allocation of compute on the ISMPAD supercomputer cluster out in Bristol. And that has allowed us to well, honestly, it's one of the things that's unlocked our ability to even might have the ambition to do something like this.

Right? Fundamentally, of the biggest blockers for, like, a startup of our size or smaller, to be honest, on being able to approach work like this is a huge part of it is the compute. Like, if you raised, you know, 50 to $100,000,000, like a good chunk of that would go on compute on doing a project like this and to have allocation come from the sub AI unit is is huge because it genuinely does enable it.

And we still use some private compete on the side, but fundamentally, all of it will be done on Izambard, which is which is super cool. And and to be honest with you, it it was some not something that was on my bingo card in January at the beginning of the year when we started out. We I mean, we still do our our, I'd say, our conventional business of, like, coding agents and the models that we already have built.

But we've been able to take that vision and really take it to the extreme in a way that we wouldn't have been able to otherwise.

Speaker 3

So I'm not I'm not being funny, but the million dollar question is, well, you actually I'll give more than that. I'll give you more than that. Yeah.

So, you know, folks over there in The US, they they've got probably on the order of hundreds of of billions. You've got Mistral, which is on the order of, you know, let's say, 14 Single to double digit billions. Yeah.

Something like that. How can you do in millions what they are doing with billions? Yeah.

No. It's it's a very fair question and one that I probably get more than anything else.

Speaker 4

and that sort of ties into the kinds of deployments in the way that we sell our product. So if your viewers, I should probably give a bit of background. Given the fact that we predominantly deploy into highly secure, high sided environments, Most of the time, nay, nearly all of the time, we are not hosting the model ourselves.

A customer isn't hitting, you know, cosign slash, you know, API slash v one and then hitting a chat completions endpoint or something like that from us. They are either taking the model weights that we give to them, deploying them on their own GPUs. We have a lot of that.

That is, like, the most air gap, the most secure deployment we do, or they are renting GPUs in some hyperscaler cloud they're already a part of. So whether it be like Azure or AWS or whatever, and then they'll run the model there. What that means in practice for cosign is that we license the technology that we built.

We don't actually make a margin on tokens or anything like that. And all of this ties into your question, meaning, we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now that's not to say that you don't also need a huge amount of compete for training.

Obviously, you do. And a huge amount of the infrastructure they have in The US will also be used for training, but I I think that one of the biggest reasons you've seen people like Anthropic struggle recently, and the reason they've signed the deals they have with, like, the Colossus cluster and so on is because of inference and not because of training. Yep.

So you do need significantly less resource if you're not gonna do the inference bit, and we're fortunate in the the way that we sell the product. It means we don't really have to. And then on the other side of that, we are taking some interesting research approaches in terms of, like, how you pull something like this off, and we can talk more about that in a minute, I'm sure, in in terms of how we are architecting the model, how we're training it, some algorithmic stuff.

All of that's to say that we we do have, a credible shot at at pulling off the full run, including, like, the the continued pre training, the mid training, the post training, all of those bits. But to be completely transparent with you, there isn't a huge amount of room for error or a huge amount of wiggle room. There are obvious places where we have had to make trade off decisions.

You know, that includes that that that that extends, like, the scope of RL. I'd always like to do larger RL runs, more generations, more, you know, inference time compute during the RL process to get more variety. We can't do as much of that as we would like to if we had, like, 10 times more compute, for instance.

Right? So there are trade offs, but fundamentally, I think given the way that we scoped the project, we can get into more detail, I'm sure. But given the way we've scoped the project, I think it is it is viable in that in that very narrow scope.

We have some of the largest companies in The UK all feeding use cases and, you know, their desires for what they want the model to be able to do directly into us so that we can train a model that's really good for them.

Speaker 3

a feedback loop that hasn't really been explored as much in the space. Obviously, you don't want to say bad things about some of these companies, but yeah, the models from Mistral, from Coherent and so on, they're not competitive. Even the Chinese models, arguably, they've only really started getting competitive the day before yesterday.

You know, GLM five point two. So the vibes are good, although on the ARC challenge, it didn't do very well, but maybe that was just a red herring. Well, I can talk about that in a minute.

But yes. Yeah. It's very cool.

But, you know, like, so is it because because you were you were almost implying, oh, it's because they weren't trying to make it better. Like, what how can we make models that are as good as those frontier models? Crudely, I think there are a couple of key things.

Speaker 4

Well, maybe three key things. I think one is architecture slash raw model size. I think the second is active parameter count, and the third is data.

I believe, and correct me if I'm wrong, I believe the largest model that Mistral have made to date is the six seventy five b Mistral three large. Right? Sparse MOE, very, very similar to DeepSeek's architecture, if not the same, I think.

As a result, like, fundamentally, the the that that model is exists in the way that it does because it fits a use case that they have seen. It probably fits a GPU deployment profile that they have seen in the enterprises they're trying to sell to in France or Europe. And as a result, pragmatically, they're like, right.

This is probably the biggest that we can get away with given what what certain companies have access to. As a result, you're obviously gonna cap out how far you can go in terms of model performance. So there there was a very interesting analysis done in a blog post, which I can't remember the name of, but I can send you post hoc because I just found it very interesting of a breakdown of the the probable sizes of models like Sonnet and Opus.

Oh, yeah. I saw that. Did you see that?

It was so cool the way that it was done, right, through vote through Vertex and figuring out, okay, given what we know about open weight models and and latency times and stuff like that? What I saw was it was they came up with a bunch of questions, and they could infer based on the general knowledge. That one as well.

Yeah. But that that was a little bit sketchy, wasn't it? I've seen a more empirical one.

I'll send it to you because it it it it it that that blog post alone played a large part in, like, our decisions architecturally we took for the sovereign model. One of the things that was very clear from that is that the likelihood is and, like, again, no one really knows outside of the of outside of Anthropic That something like SONNET is in the, I believe, like, one point one point three to 1,500,000,000,000 total perms, MOE, and has probably a 100 a 100 plus billion active. Right?

And then an OPUS is probably in the 1.5 to 1.8 range and probably has a 150 to a 180,000,000,000 active depending on the d type we're talking about, whether it's f p a or f p four, whatever.

What about Fable? That that wasn't up for long enough for for for me to try to work, because I basically, what I wanted to do is give Fable that blog post and then point it and then point it and be like, right. Do the analysis on yourself and tell me how big you are.

Never never got that far, though, because I don't think it was up for long enough. It was it's a ridiculous uplift, though, isn't it? It's I saw I'm sure you saw on x, like, numbers like 10 trill being knocked around.

I don't know if that's true or not. Right. I have no idea.

I think it's obviously bigger than the other ones and obviously has way more active premises, But I I couldn't speculate. I genuinely don't know. But I yeah.

I think that the the net of that of that article was that, you know, Opus was in the region of 1.5 to 1.8 with around a 150,000,000,000 active.

Obviously, there's a lot of algorithmic and data work that goes into it, but I think if you don't at least match that architecture, then you're you're already gonna struggle to to sort of reach that ceiling. And I think an example of this, and it's way more nuanced than than this fairly basic argument I'm gonna make. But if you look at like a DeepSeek v four Pro, right, 1,600,000,000,000 total, can't remember the exact number, but it's gonna be in the region of 30 to 50,000,000,000 active.

Yep. Right? And and my view is that the reason for these architectural decisions is largely for inference ability of the model.

Like, it get like, sure. Great that you have those top line 1,600,000,000,000 parameters, but, like, if no one can run it cause they need two nodes of b three hundreds just to be able to, like, fit it into memory and actually, you know, run it at decent TPS, then, like, how many people can actually take advantage of that? I think that is one of the reasons that whether it be, like, Chinese open source or European or American open source has has not reached closed source performances.

Like, there is that pragmatic question of, okay. Well, the labs have a huge number of GPUs, and they have enough inbound demand to make sure those GPUs are utilized to a level where they're not that worried about having them up. Whereas, you know, the the the open source community doesn't really have the same the same argument if you're running it yourself.

If you're gonna if you're gonna run a model of that scale on your own hardware and it's not really being utilized that much by your organization, you're gonna worry about how much money you're spending on just those GPUs being idle. Yeah. So I I I think that on that first point, architecturally, I think that overall parameter count is obviously important.

I think active parameter count is also incredibly important. And then the last bit is data. I think that I think that obviously the labs have some element of edge in data both because they are able to procure so much from the brokers who sell it and also because they are able to they they have internal data functions which are very mature at this point.

I'm not saying the other labs don't have that, but they're definitely not at the same scale both in terms of spend or maturity. Particularly, I I I think I think to an extent, and this is definitely not true, but I think to an extent, like the pre training corpora, Is there gonna be that much difference once you're in, the 30,000,000,000,000 token range? Okay.

You kind of all have roughly the same stuff. You we we we're compressing most of the Internet at that point. Mid training, similar story.

I think post training is, like, has been and remains one of the most interesting areas for these labs. I think that is having having worked with some of the labs on post training data because obviously we're we're very good at coding. It's been interesting to see, like, how they've been procuring even the formats they've been using, like the raw data and the different use cases that they're interested about when they're putting requests out for we want this, we want that, we want the other things.

I think that is one of the key areas they're differentiating. Having really good post training data and also just being able to run RL at ridiculous scale.

Speaker 3

Quick pause. Agents are getting smarter every day. But even the smartest agents get stuck without the right context and the right tools.

That is where Notion comes in. With the recent launch of Custom Agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now their new development platform is turning that workspace into infrastructure developers can build on.

Now this is exactly how I run MLST. The whole show lives in Notion, my guests, the publishing calendar, the commercial side, everything is in there. But what's changed is that it's now agentic.

I just talk to my agent, it can be Claude or any agentic harness, and then it then talks to Notion via the MCP or the CLI, and it's just done. And then I can access it on my phone. It's an absolute game changer.

We'll get to the RL bit, because I know you've got an opinion that, you know, RL training is really, really important. But I mean, there's a few things on what you just said.

Speaker 4

having an MOE? I think at least like my opinion is if if if if we could had trillion parameter dense models, we would. It would just be strictly better.

Think it would just be better. Yeah. I think but I think, like, again, that that is taking it to the logical extreme of of could you what would you need to inference that?

What would you need to actually run a parameter you know, model of of that of that size? I think what like an interesting an interesting example of this is and and it's not really a fair fight. But if you take a small scale, something like a a GPTOS s one twenty b.

Right? That is an MOE. I believe it's 5,000,000,000 active.

I can't remember. It was a while ago, but it was something like that. And then you take a, like, Devstraw two one twenty three b, which is, I believe, like a similar architecture to Lama 70 b, so it's fully dense.

I don't know if you've used them back to back before, DevStraw feels so much better than the GPOSS model. Yep. Right?

Yep. And that that they are architecturally quite quite different, particularly in like the attention mechanism. But fundamentally, I think that a huge part of that just comes from the fact that one, you have a 120,000,000,000 active parameters, again, the other you have five.

Speaker 3

and the difference was night and day in terms of how it felt and what they got out of the coding agent when they were running it. Oh, interesting. Yeah.

Because on the other stuff, so you were talking about like the data, the pre training, the algorithmic stuff. Mean, maybe the algorithmic stuff is kind of converged only because we now have this basin of attraction where there's kernel optimizers and entire ecosystems around this and maybe it's converged. But the data thing is interesting, right?

Because there must be an insane amount of engineering. And with LLMs, it's a little bit like, what's the magic word? If you frame the question in the right way, it has that representational friction and it does interesting things.

So I'm guessing Anthropic do a whole bunch of date data curation and pruning. But another thing that Anthropic do is they have Claude code. Have the ecosystem.

Speaker 4

Have trajectories coming in all day, every day. Exactly.

Speaker 3

you've got this big thing, which is that it's not about where you end up, it's about how you got there. And I think software engineering is the it's not about writing code. It's actually just creating- Creating mental abstractions, doing experiments, refining those abstractions, showing them with the team.

Process, so iteratively, I'm running code, I'm testing things, I'm refining my abstractions. And Anthropic have access to so much of that data. How much of an advantage is that?

Obviously it's a huge advantage.

Speaker 4

I don't know what their terms of service say. I'm not saying that they train all that data, I don't know whether they do or not. Certainly, one of the interesting things that having trajectories means and something that, like, obviously, internally, we, you know, collect our own trajectories of our own use of cosign.

Obviously, don't have anyone else's because the way we deploy it. But for cosign employees using cosign at least, we do still get a fair number of trajectories, nothing like the order of magnitude that Anthropic get. But it the most useful thing for trajectories for us is seeing how users prompt models.

So like, there there there's one thing when you're putting together RL datasets, and you can come up with these like beautifully formed like problem statements that get fed into the model that are very well structured and very well formed, and they're very clear about what the expected outcome is. The reality is that users don't prompt models in that way at all. They're like, f you.

It doesn't work. Why the hell? Like, are you dumb?

Like, why have you done this, that, and the other thing? And and the reality is is is that's the distribution that the model's gonna be exposed to in the real world is you need training data that looks like that. And I think that I would speculate there's a lot of a lot of the alpha that Antarctica gets out of their, trajectories they get back is obviously obviously, they're helpful.

I don't think that they canonically know. It's not like they there is a grader attributed to that trajectory. Like, they don't it's not like an RL thing where you know that this trajectory was good, like, absolutely.

I'm I'm sure that you can do some judging stuff, and you can see how the users responded to be like, okay. Well, at this point, the user was probably satisfied that the job was done. But it's also fundamentally seeing, like, what is the user saying during the conversation, and how are they saying it is incredibly useful.

And it's and and and even even our scale, we have there's an argument to be made about how representative it is for, like, the entire development ecosystem. That's a different question. But we we do have, like, an idea of of generally speaking across across months and months and months of usage.

Like, how do engineers interact with these products and how do they respond and all these things. That's very useful to make things more realistic at training time. Yeah, 100%.

Yeah, because I mean, you've spoken about SLOP online. It's one of my pet topics as well.

Speaker 3

quite often it does the thing that you tell it to do and the tests will pass. But actually you're building a spaghetti monster and you're throwing more bad spaghetti after good spaghetti. And instead of doing a one line fix that it's supposed to do, it'll give you like 200 extra lines.

There's the whole understanding debt thing. We can talk about that. But, so what we actually wanted happen and neural networks do do this to a certain extent, learn statistical inferences that represent some kind of abstract structure and that helps them to generalize.

And obviously what we want them to do is to learn problems in the abstract so that they can generalize to new novel problems that they've never seen before. So the idea is we capture the thought process and then we capture that generalization if we do it well enough. That's the rough idea.

It I mean, how do you think about slop?

Speaker 4

slop is also one of my pet hates. When you are obviously, any any of your viewers who I'm sure use ClaudeCode or OpenCode or whatever agentic coding harness they like, See Ad Nauseam is is is like both a combination of model vibe and also slot problems. And you hit the nail on the head, I think, in your question, which is when you said that, okay.

Yeah. The test pass. Right?

Okay. Is it is it technically functionally correct? Sure.

Right? But to what cost Yeah. Is is the big is the big question.

And I think fundamentally, I mean, like, when when you think about how these models are trained to do software engineering, and this is the realization that we sort of had when we when we built Outpost, and sure you've you've seen the blogs about how we did it. Within that realm, normally, and again, we don't have that much insight over how this works in in the big labs. But normally, when you're training a model to be better at software engineering, you have some kind of software engineering problem.

You have a problem statement. You have some kind of test, often a unit test, but not always depending on the task type that needs that it that is either failing in the early in the prior state and passing in the after state. Then, you know, you you you give the agent the problem.

It goes and does its trajectory, its rollout. And then at the end, you you run the unit test, and if it passes, then, okay. I'm like, you got it right.

Great. You get rewards, and then the the the weights update. Obviously, the the problem with that is is is a few fold, but fundamentally, like, it could have come up with the most insane way of doing something, whether it be in terms of, like, commands that it ran that could have been unsafe in terms of code that's absolute crap compared to what it should have actually done.

And all of those things get reinforced whether you like it or not when you give that reward based on purely correctness. So there are a number of things that we have done and are continuing to do in the RL process to try to ameliorate this. And in terms of slop specifically, there are a couple of key things.

One is that, like, correctness gates everything else. So, like, if you get the problem wrong regardless of whether you did it in an elegant way, like, you don't get rewarded. But beyond that, we do have other rewards that target the exact things that we've been talking about.

And also, in some cases, but not all, we do have reference implementations for these things. Like, if if if you are using a a pull request from, like, a a permissively licensed open source repo as, some seed data, you do have the original patch, the human made, and you can actually do some level of, like, okay. Let's compare what the agent wrote.

Let's compare what the human wrote. And does this seem reasonable to be, you know, 500 lines longer? Probably not.

Maybe we shouldn't give the full reward for this. So stuff you can do on the pure reward level, but we're also doing stuff on an algorithmic level, which has been looked at more and more in the space, which has to do with credit attribution in trajectories. One of the big problems with RL as it stands, I think is that, and and I I I'm by far not the first person to say this thing.

Andre Kapathi said this, like, over a year ago. But fundamentally, this notion of you have a rollout of maybe 256,000 tokens, right, in some extreme cases. And that culminated in, a one or a zero depending on what the model did.

And what we're saying at the moment in many cases is okay. All of those tokens are equally weighted in getting us to that answer. Yep.

Yep. Which when you think about it is insane because that's clearly not true. And in so many cases, the that there will be small or, like, important decisions in a trajectory that were sort of, you know, forks in the road that would have either that could have resulted actually a bad outcome, where the model decided to get on the right one and ended there.

And what we and others in the space, I guess, trying to do right now is, okay, if you can find out what those ranges of, like, I guess, high end well, high entropy tokens or places where you know a decision was made, finding that is half the problem. And then figuring out once you know that this is an important thing that happened, determining whether it was good or bad relative to the final outcome is a different story. Yeah.

But if you can do that, a, your RL gets significantly more efficient because you're not relying I mean, analogy I always come up with is like, okay. Say you were doing your English a level, and you'd written you you were practicing and you'd written like an essay for your teacher. And you'd written a 2,500 word essay on a question, and the teacher just gives you, okay, that's a b.

Thank you so much. Yeah. And you're like, I don't what made it a b?

He's like, I'm not gonna tell you what made it a b. It was a b. And then what you're gonna have to do is you're gonna have to write hundreds of essays, and you'll get an a on some, you'll get a b on others, you'll get a c on others, and eventually, you are gonna be like, okay, so when I do this, I tend to get an a more.

So I think this is probably a good thing to reinforce. But it would be far easier if the teacher just send you know, circle the sentence and be like, this is rubbish. Don't say this.

Right? I know. And that is fundamentally the principle we're trying to bring into, like, RL across the board because you get so much more performance, you get more out of the flops that you have.

Speaker 3

And also, you're teaching the model to learn the things that are actually important and not just the filler. Right? I I know.

I mean, the the great thing about machine learning is that it just generalizes low down the abstraction mountain. So from very superficial statistical generalizations, the bad thing about machine learning is it generalizes too Exactly, yes. So yeah, I completely agree with you.

We have a huge problem with benchmarks in machine learning. So we are obsessed with pass that one, pass that five accuracy. And we don't seem to care about reliability, about consistency, about security, about abstraction forming.

That's clearly the most important thing. I mean, Francois Cholet, he did the art challenge and unfortunately they were brute able. Now he's got this new version, which is so difficult to brute force.

You have to form abstractions to get any kind of good performance on it. But so you're saying there's a new form of RL, perhaps different from the deep seat type of RL, which is like, rather than just being rewarded for getting the right answer, you're actually forcing it to form reusable abstractions and go higher up the mountain. Yes.

Speaker 4

In short, you've explained that far better than I did. But yes, that is that is that is essentially what we're trying to get to. One of the nice things is that I think that essentially what we're talking about here is credit attribution within a within a trajectory.

That ports quite nicely to a bunch of different RL algorithms that are in vogue at the moment. You can use it with a gRPO. You can use it with, you know, GSPO and all the different flavors of the algorithm.

But fundamentally, having a rigorous and importantly unopinionated way of pointing at ranges of work that an agent has done and being like, this is good, this is bad, and so on.

Speaker 3

either appear more frequently or less frequently. Do we still have an epistemic problem? Because you know, the the one problem with machine learning is it it doesn't really have, like, the notion of true and false.

So, we can do feedback from code execution feedback. We can do a whole bunch of abstract lenses on actual processes that engineers are doing. But don't we still have this gap though, that we don't really know whether it was correct or not?

Speaker 4

It's one of the hard yes, it's one of the hardest things. Particularly, I mean, one of the obviously, one of the reasons that coding has taken off as a use case is because you have verifiable rewards in some guys. And I think one of the things that, I mean, I've just said is that we are trying to bring some of the more fluffy taste related things and make them verifiable, but on a on a more floating scale.

I think for obviously, for things like law, for things like other nonverifiable domains, that is way, way harder. That is way harder. And I think that's one of the reasons one of the key reasons that we haven't seen, like, the same revolution in other industries as we've seen in coding or maths or or physics because you can't just statically, like, compile some law and see whether you get, you know, a one or a zero.

And I don't know I don't know what the answer to that is. I really don't. Maybe maybe there are new ways of codifying those domains, or maybe you just bring everything in distribution, which I think is what's happening these days.

I think if you just make everything in distribution, you target every use case, and you and you have some level of whether it be human judge or good LLM judge that has been either trained or well prompted by a human, you get close.

Speaker 3

the dream of generalization across the board, I I I think, has already been sort of shown to not really exist that much. I know. But we're in such an interesting time because every new model comes out and, Fable was so much better, but we want to have systems that do more with less, which is what you're saying.

So we want them to acquire these abstractions. And it's a really, really weird situation, right? Because do you think we'll ever get to a point where we can remove the human from the loop?

Because right now I think the basis of AI psychosis is that you have very, very talented humans, and they know how to ask the question, and they have taste, and they go in the right direction. And there's this virtuous co creation cycle. It's very, very good.

We're now in the realm of, it's not vibe coding, it's agentic engineering. It's very exciting. But do you think it could ever be done without humans?

Speaker 4

Yes, I think it can be done without humans. I don't think we're anywhere near there yet.

Speaker 3

you mean in well specified? I'm not talking about style transfer. You know, there was that Anthropic, they built a C compiler.

That's a well specified problem. I give you a novel application, and I can only vaguely specify it. We're gonna need a human for a long time, aren't we?

Speaker 4

If you want it to be good and maintainable and actually in the style of something that a senior engineer would write, For now, yes, you definitely need a human there. Probably post hoc to be like, okay. Here's the mountain of stuff that you need to change, and all the design decisions that, like, in your chain of thought, you thought were good but actually weren't because of, you know, real reasons.

But I do think that I do think that we will get there.

Speaker 3

and also harness engineering. I think a combination of all those three things. Yes.

And we'll get to harness engineering. But okay. But in the meantime, we've got the spaghetti monster mitigation strategy.

And and I think one of the big problems is code review. Right? Yes.

Like the AI psychosis is manifested as understanding that. Increasingly, I become less aware of what's going on. That's actually really bad for me maintaining my competence and for me being able to evolve the software going forward.

Ask the right questions. Yeah, absolutely. Exactly.

So how can we do this? Because now we are generating ridiculous amounts of code. Right?

Is this a case of let's use more AI to do the review? Or do we still need humans in the review?

Speaker 4

I think that for I I think we need more, runtime validation of what AI is producing. And what that looks like is I I I think code review will evolve somewhat. I think it will be more like proof that the thing that it says it's doing is actually doing that thing.

Obviously, like AI reading get diffs is is is not not useful. I think it can catch things, and I've seen it catch things in the past. But one analogy and and it's not actually for code review, but one analogy that I can I can tell you that's really good when it comes to, say, our cybersecurity scanning products that we have, which I guess is analogous to code review because, like, it's reading basically a whole code base using a swarm?

But one of the key things that we saw with that was that it would go through that process, and it would pick up, like, so many things across large code bases. Like, this could be a problem. This could be a problem.

This could be a problem. And I think one of the things that we see in that and in AI code review is, like, sure. Like, if you just look at that code in isolation and that function definition, it can look quite dodgy.

But in reality, the code path is never hit or the there is, like, this other function that's called first that mutates this variable that then means that that doesn't do what you think it does, or there's an environment variable that happens at runtime that means that this doesn't happen. All of these things that, know, it's just task analysis. You can't do it.

And one of the best ways that we mitigated that problem in that product was that we had what we call exploit validation, which is literally like, okay. The the agents, the swarm comes up with its list of things. And before it makes it to you, every single one so we spin up the application in a virtual machine, in a in a in a runtime, right, and in as club close to production way as possible.

And we tell the agent, right, well, you've seen the source code. So if it's vulnerable, you should you should be able to figure out how to get through it. Right?

You should be able to, like, craft your your horrible zip file to exploit this thing that you think exists. And if you can't, then you just take it out the list because it's clearly a false positive. And what we have been working on is applying the same logic, but to just PRs.

Right? And and we're not necessarily looking for, like, cyber vulnerabilities, but we're looking for, okay, you have allegedly built out this feature. You have this new screen that has this table in it or this form in it that does this.

And when you click on this button, it should result in a new entry in the DB, and it should show up in the all the stuff you'd expect. Right? And instead of, like obviously, you can look at the diff, and for the most part, you can get a lot of mileage out of that.

But also just, like, show me it's doing that. Prove that it's done that in some reasonable way before before it even makes it to me. Because otherwise, you end up in this situation.

Was just talking to a customer earlier today where they were saying when they first started adopting, like, agentic coding tools, they were in the spot where either they would end up with this enormous backlog of code review. Right? And and and no one likes code review.

Let's face it. No one actually enjoys that. Or you get people being like, after a minute, looks good to me.

Merge. Thank you so much. And you just get this YOLO merging into, like, your main branch, and that's bad as well.

So I think there is gonna be way less cognitive burden if in however you're doing a review, you can see the code, but you can also also see canonical proof that at least on the happy path, the thing that it's it's saying it's doing is actually happening. And the other half, I think, at least in the present day to this is also really comprehensive end to end testing of everything you build. And that is something that really sucks to have to build out.

But once you have it, it saves you from so many problems because I'm sure as, you know, you've seen I've seen in, personal projects and so on. You vibe code for an afternoon. You built 10 new features, and all a sudden, the other five you had before stopped working.

Like, why has this happened? It's like, oh, well, okay. The abstraction you had, I've just messed with it.

And now it doesn't work for that thing. So, yeah, it's a combination of, like, defensive stuff, the end to end stuff, and also just proactively lifting mental burden from people by being like, look. Here's, like, either a screen recording or some screenshots or whatever of of me showing you that this is what I think it is.

I I know. I know. It's it's really funny.

Speaker 3

it it'll say, oh, you're right to point that out.

Speaker 4

Oh, I'm so sorry for doing that. No. Sorry for doing that.

There's nothing you can do about it, by the way. But yeah. I This is another alignment problem though, right?

Speaker 3

we are accumulating understanding that the functional descriptions themselves are going to suffer from that because we don't understand the functional description anymore. And it's not just functional descriptions, there are intents, there's behavior, there's all of these different levels of describing a system, user stories and stuff like that. And unfortunately, these are different views of the blind elephant, right?

They don't necessarily have friction with reality. And you see the problem here? We're just losing touch with what it's supposed to be doing.

In many cases, actually, Claude and the models, they're writing the functional tests, and then they're kind of hacking their own Oh, it didn't pass. Okay. I'll just change the test.

Now it passes. Great. Here we go.

You see it all the time. I know. So, you know, it's a very difficult problem, but one that we need to fix.

And, you know, for me, I think a lot of it is to do with scoping and constraints, know, to at least cut down the size the problem. But we should move on. What are your thoughts on Agenic engineering?

You guys have got an Agenic harness. If I understand correctly, a couple of years ago, you actually forced everyone to start using that because you really want to optimize the hell out of it. You were talking about the RL piece, maybe there's some co evolution with the Agenic harness and the RL.

What's important in this agentic harness?

Speaker 4

I think that so broadly, I think agentic harnesses are getting less important over time. Oh, interesting. Why?

The models are just getting so good. Oh. I think I think that I think that you can get the proof point is that, like, a model can probably do with bash only basically any task these days.

More slowly and with more tokens, but it can probably still do it. That's not to say that agentic harnesses aren't important, but I think, like, over time, where's the value coming from? It's coming from the model and not from from the harness.

I think in the way that we built ours, and and we we have been building agentic harnesses for a very long time. Like, the first agentic model that we had was a fine tuned GPT four turbo model that we trained in, like, January 2024. And the coding agent harnesses didn't exist at that point.

Claude code didn't exist. None of this stuff existed, and we had to build one out. And we we did that symbiotically with the design of the model, which is something we still do today because you do get way more performance out of, like, tightly coupling the two fundamentally.

Harness engineering is still something we care a lot about. These days, we actually care more about efficiency than anything else. In a world where token costs and tokenomics, which is a word I heard for the first time today, awful word, tokenomics is becoming increasingly important to particularly enterprises who we sell to.

We want our harness to use as few tokens as possible, full stop. That's what we're trying to do. Yeah.

Speaker 3

you could, in principle, allocate a budget of tokens to get certain things done. But but even before we get there, I mean, I wanna push back on on the harness engineering because one thing I have found is sub agents are a game changer. Yes, absolutely.

And that is because a lot of problems are too complicated for an LLM to do in a single pass. I'm sure you've had a similar experience. As a problem becomes more specified, as you reduce the ambiguity, the entropy goes down, the models get better, and the models are better when they have less rot in their context.

So what happens is the engineers, after doing a bunch of engineering, they find is that they decompose problems into agentic subtasks. And then they have a fresh agent and the agent has a clear specification and only do this one thing. Exactly.

But, know, so what you're doing logically, as an engineer is you're factorizing a problem into smaller sub problems, you're getting agents to orchestrate and you're not rotting the context in the main one. I mean, how how do you see that evolving over time? Because now it's quite a manual process, but do you think I mean, you've got this swarm thing maybe in the past.

Thank you for bringing that up because that was exactly what I gonna want. So you're on.

Speaker 4

to me, is sub agent orchestration taken to the logical extreme. We're kind of lucky to be able to do it because I think that, obviously, ClaudeCode has, I believe, workflows, and Codex has sub agents as well. A swarm is what it sounds like on the tin.

A swarm is genuinely like a a a real a lot of a lot of sub agents running at the same time in a hierarchical way. And because cosign isn't trying to serve hundreds of millions of people a day, we are able to serve a swarm like feature. Whereas I think if Anthropic had a swarm for, like, Opus, I think they would even Colossus Swarm would run out of of run out of tokens.

Swarm does exactly what you've just outlined automatically, basically. So that there is a there is a video on my on my Twitter and also my LinkedIn of me taking our Lumen Outpost model, which is post trained Kimi k 2.6.

So definitely not an Opus or a Mythos or anything like that. And I asked it, okay. I want you to build me a mechanical watch compiler was the use case that I came up with.

I am Swiss by birth, so I have a reason to do this. And I basically asked and and actually someone did this exact same task with Fable. I don't know if you saw it on Twitter when Fable was out.

But essentially, like, I want you to build me a SDK in Python for me to be able to, like, specify mechanical watches in code because I don't know horology, but I want to be able to do it anyway. I want it to be physically congruent. I want you to use some kind of physics engine.

And then I also want, like, a three d viewer to be able to, like, see the thing running. And it all needs to be possible in real life. You can't have things like intersecting which other wouldn't be possible and so on.

And that is something that out of the box, Kimmy cannot do. There's no way. Not even close.

Like, it would be terrible. In fact, even Gemini 3.5 and Opus and 5.

5 can't really do it. But as soon as you put them in a swarm, and for what swarm looks like for cosigners, you have one orchestrator at the very top. It breaks down a problem into sub problems for basically product managers or or or whatever you wanna call them.

We call them sub planners, but they own verticals of this. So, like, within that task, you would have had a a sub planner to do the SDK. You would have had a sub planner to do the three d view.

You'd have a sub planner to write the documentation and so on. And then those sub planners could then dedicate to workers, and then they have, like, a flat layer of of as many workers as they like. And for that problem, we use 253 sub agents, which I think is more than you tend to see in a Claude code session and so on.

I think I think you'd you'd probably hit your usage limit pretty quickly that way. But when you do that, it is possible. And you can do that entire project in one shot.

Speaker 3

And and and I am contradicting myself quite badly because I've just said harnesses don't matter, but in that respect, they do obviously matter. Oh, indeed. I'm very excited about that.

But, you know, it raises the question that when first of all, when when you start to have loads and loads of agents, you have more understanding debt and less interactivity because Oh, yes. You know, for me, the the the lack of interactivity is part and parcel of the understanding debt. And sometimes you want to interject, want to say, you've gone slightly wrong there, I want to change what this agent's doing.

And what many folks have found when they build these agent systems is that the agents interfere with each other, they kind of overwrite its own. They go into deadlock. How are you dealing with all that?

Speaker 4

it's a hard problem, and we experience all those things. Yep. In terms of you can so one of the key things that we did is we made you we gave the ability to interject to, like, an agent on the lowest level.

So say you had a worker that was, like, two levels down from the top one. You can actually, like, talk to that one, which is an important thing. And with regard to, like, other problems in terms of, like, treading on each other's toes, you can put write locks on files so that you only one agent can can edit it at a time.

You also can provide context to agents when they are using say, like, an agent is reading a file that it just read. We have stuff in the harness which says, okay. Another agent has just edited this file, so don't be surprised if you see it slightly differently as to how you saw it last time.

All these things help. They are not like a panacea, but they certainly help. And they they they make sure the agent is less surprised when they're like, oh, where did that come from?

That wasn't in my last edit. But fundamentally, it it is it sort of comes back to the point I made earlier where, okay. Yes.

You will run this thing, and it will provide you a huge amount of value very quickly, but you are still gonna have to comb through afterwards to be like, actually, the reality is I don't like this abstraction you've done. I don't like the way you've done this. There is going to have to be some sweeping afterwards, I think.

Yeah. What are your thoughts on memory? Very hard to get right.

Okay. Tell me more. Very hard.

Speaker 3

Similar thing actually to the RL, because memory is not about the destination, it's about how you got there. Yes.

Speaker 4

Memory is very hard to get right. We've tried a bunch of different approaches. Fundamentally, I think every approach to memory that exists right now is a bit of a hack.

Right? It's like a tool, and it's it's, in many cases like a VectorDB or like an embedded version of of, like, some tidbit of knowledge, but they're very hard for the agents to know when to query. It's also fundamentally quite hard for the agent to know whether something was useful enough to write to memory.

And it's it's it's also difficult to to keep these things up to date. We've had many situations where an agent's been doing a trajectory, and when it's been doing something that was genuinely the right thing to do, it's like used its memory, and it's like the memory's been old and out of date. And then the agent's like, oh, well, the memory says you should do it this way, and it changes tacking.

There you're like an engine engineer. You're like, no. Please don't do that.

There are things that we're looking at internally with regard to, like, continue continual learning and stuff like that to try to avoid memory being a tool and for it to just something that for it be something that's just in the latent space with the model. That is also very hard. But I think it is a more intuitive and elegant solution to it just being a tool.

It's also a tool that's very hard to get right during RL because it is a huge surface area foot gunnery, in terms of reward hacking, in terms of leakage, in terms of it's being able to clear some query something from the future that it shouldn't have access to yet despite all the guardrails that you can put in place. It's it's just hard. Yeah, exactly.

And in a sense, this is another area for AI psychosis because I've written a memory C line. And I would almost argue that now you don't even need vector databases and so on.

Speaker 3

just to see called right, because the models are so good at asking So in different yeah, there is a huge problem it needs to know to retrieve, that's a big one. But it actually works, but it only works for me because it creates this fractionated spaghetti mess again. There's another spaghetti mess in the memory CLI, but it works really, really well.

But it doesn't work very well at the organizational level because my Spaghetti Monster doesn't play with John's Spaghetti Monster. So if we could solve that problem, and you're drawing an interesting picture as well of how we can actually optimize the different layers of the sandwich together.

Speaker 4

Yeah. I think that as soon as someone gets it right, you'll just know when you're using it immediately. Like, I I haven't seen a single implementation, whether it be in Claude code, to be honest, whether it be ours or chat GPTs or any of them where I'm truly like, oh, no.

This isn't just a hack. This isn't just this isn't just rag. This is this is like the it's the one remnant of rag that still exists really in in in the more traditional sense.

Speaker 3

I I am certain that there is a better way out there somewhere. Oh, definitely. But I think another thing you've said is specialization or generalization.

I mean, for for me, lot of agentic engineering is like emergent specialization. So it's it's like, let let's take a big intelligence to crystallize a small intelligence to do the particular thing we're doing. But as we're nearly out of time, final question is synthetic data generation.

So what are you guys doing about that? Tons.

Speaker 4

I don't know how much of it I can get through in five minutes, but I'll do my best. So there are a number of areas that we have done synthetic data generation in in the past. We're doing a lot more of it in the more forward looking sense for, like, the sovereign model, because the sovereign model can't just be good at software engineering.

It has to be useful across the board. And that and that that means that we have to get good at that synthetic data generation, particularly in the RL realm in things that aren't just coding. The bread and butter though is coding.

And the way that we have done this in the past is with a very cool and sophisticated pipeline that we have built out over the course of like a year and a half now. But that's one of the big cold start problems in CodingRL, particularly if you're using like open source repositories as a sort of seed data, and even if we're using closed source, actually doesn't matter, is that nearly all of the PRs or commits or whatever you wanna refer to them as don't have a built in grader. Mhmm.

Right? The ones that do are often bug fixes. The the like, the the classic example, I guess, you're if you're nerdy enough to be in the space is like a sweet bench style problem.

You have a git issue. You have a PR that fixed it. And then because it's open source, you have some sort of regression test that was added, and then that's your, like, seed data.

The real world doesn't look like that, unfortunately, and and software engineering in the broad doesn't look like that. Meaning that to do RL well, you still need to fundamentally be able to do tasks that aren't bug fixes, which is the vast majority of of of what of what engineers do. And you still need to be able to tell whether, like, the agent got it right or not, broadly speaking.

Yeah. And what we do at cosign is we take, like, real work that was done. We're not we're not magicking up made up problems for the problem for the model to solve because there is fundamentally, if the model can come up with a problem, it can probably solve it.

But we take real problems that were solved, whether it be feature work, refactoring, whatever it is, And we what what we're synthesizing is we're synthesizing ground truths or grade as always of measuring whether that thing has been done. And it is a bit of a minefield because particularly with an RL, obviously, it's not supervised. So fundamentally, we need to be able to test these things in a way that isn't too tightly coupled to, like, the original implementation.

Like, there are many ways to skin a cat as we know. And what that looks like in practice is you need a implementation agnostic enough way of testing it that's still rigorous rigorous enough to check functional correctness. Mhmm.

It is a very fine line to tread, and we have done a lot of work around it. We have a long pipeline built on it. We have custom post train models that live inside that pipeline that essentially have gotten good because we had to do a lot of manual labeling in places and stuff like that.

But what it has allowed us to do is is have a autonomous pipeline where and this is particularly important for enterprises. We can point to essentially any programming language, type of task, any stack, anything like that, and say, okay, I want RL data for this problem set and we can get a, you know, like a good chunk of it. And that is why for, like, our Outpost model, when we were coming up with the the languages that we wanted to get the model good at, things like Java, Fortran, c plus plus, I could go on.

We use that pipeline to to gather the RL data for these things. And in many cases, and in some programming languages, there aren't even test suites. That's where it gets really hard.

That's where you have to get a bit a bit inventive in terms of, like, measuring whether the agent has gotten something right or not. Like, I think it yeah. It's it's it's things like Verilog and Sysverilog where you have to actually run, I think, what's called a synthesizer in your environment to be able to tell whether, this chip actually works or not.

Speaker 3

Yeah. An EDA. Yeah.

Yeah. Exactly.

Speaker 4

Fortunately, I don't run this pipeline. A chap called Ben does, and he knows far more about that than I do.

Speaker 3

other use cases that aren't just software engineering as well. Very cool. So in in closing, you might argue that the US government have handed you a commercial advantage here because now it's so it's more important than ever to build sovereign AI.

You guys have this model coming out towards the end of this year. So do you think you're going be able to do it? But also, are you still at risk from a supply chain point of view?

Because so much hardware is controlled by America. How's how's this gonna pan out for you guys?

Speaker 4

Naively, I think that if the hardware is already in The UK, I don't know how much there is they can do about that. Like, the all of the all of the infrastructure the model is gonna be trained on already exists and is up in The UK because we're doing it very soon. Right?

In fact, like, experimentation already is happening upstairs. In terms of is it is it has it been a bit of a gift? What's happened recently?

Given our positioning, absolutely. Yes. Without a doubt.

It has been probably the busiest I have ever been since founding the company in terms of people coming to us saying, okay. We we now realize what you're doing is really important. How can we be involved?

Like, the consortium of companies that you read out is is growing by the day, And the involvement and the urgency importantly from those companies, from government, from just, like, citizens as well has gone through the roof. And and and I feel very fortunate and to an extent lucky that that that obviously, like, we were well positioned to take advantage of this early. Also, like, we did put ourselves in that position, but also thank you Donald Trump.

How did we not see this coming though? Because it was such an oh shit moment for so many people. I think we did though.

I think a lot of people I mean, we certainly did, but it was always it was always it was always fobbed off as like, oh, sure. Okay. I suppose that could happen.

But, like, I personally didn't expect it to happen as soon as it did. Like, I I am I am also freshly surprised by, like, the 5.6 news that I'm sure you've seen as well where that's gonna be rolled out.

And there there is even for me, like, huge part of me being like, oh, man. That really sucks. I wanted to try that model, and I don't know whether I'll be able to now.

And that might just be the existence for us now unless, like, we and others do work to get that level of performance out in some other way, and that's like our job now, I guess. Yeah. I mean, for the first time, I feel like a second class citizen because those those folks over there in America, they've got better AI than I do.

Yes. They are going faster than I am, and I I really hate that. Mhmm.

And and and and believe me, that boils my blood more than anyone else. And I we are going to do everything we can to pull this off. You mentioned, like, how are you gonna do this?

Like, we are just going to we're just going to make it happen. We have no choice but to make it happen. Please do.

Yes. We are going to do everything we can to make it happen. On behalf of everyone in The UK, please do.

We'll do our best. Thank you. Oh, so it's been a been a pleasure.

Thank you so much. Very much for having me.

Shared via Hopper