This episode features Yann Lechelle and Guillaume Lemaitre from Probable, a company stewarding open-source data science projects like scikit-learn. They discuss Probable's unique governance model, which prioritizes open-source contributions and sustainability, and highlight scikit-learn's foundational role in machine learning, contrasting it with generative AI and deep learning. The conversation also covers scikit-learn's widespread adoption, its utility in real-world applications like fraud detection, and future plans for the project and its ecosystem.
Welcome to Practical AI, the podcast that makes artificial intelligence practical, productive, and accessible to all. If you like this show, you will love the changelog. It's news on Mondays, deep technical interviews on Wednesdays, and on Fridays, an awesome talk show for your weekend enjoyment.
Find us by searching for the changelog wherever you get your podcasts. Thanks to our partners at fly.io.
Launch your AI apps in five minutes or less. Learn how at fly.io.
Okay, friends. I'm here with new friend of ours over at Timescale Avthar Suathan.
So Avthar, help me understand what exactly is Timescale? So Timescale is a Postgres company. We build tools in the cloud and in the open source ecosystem that allow developers to do more with Postgres.
So using it for things like time series analytics, and more recently, AI applications like Rag, Search, and Agents. Okay.
AI application development, what would you tell them? What's a good roadmap?
getting tasked with building an AI application or you're interested in just seeing all the innovation going on in the space and want to get involved yourself. And the good news is that any developer today can become an AI engineer using tools that they already know and love. And so the work that we've been doing at Timescale with the PGAI project is allowing developers to build AI applications with the tools and with the database that they already know, and that being Postgres.
What this means is that you can actually level up your career, you can build new interesting projects, you can add more skills without learning a whole new set of technologies. And the best part is it's all open source, both PGAI and PGVectorscale are open source. You can go and spin it up on your local machine via Docker, follow one of the tutorials of the timescale blog, build these cutting edge applications like Rag and such without having to learn 10 different new technologies and just using Postgres and the SQL query language that you'll probably already know and are familiar with.
So yeah, that's it. Get started today. It's a PGAI project and just go to any of the timescale GitHub repos, either the PG AI one or the PG vector scale one, and follow one of the tutorials to get started with becoming an AI engineer just using Postgres.
Okay. Just use Postgres and just use Postgres to get started with AI development, build Rag, search, AI agents, and it's all open source. Go to timescale.
com/ai. Play with PG AI. Play with PG vector scale all locally on your desktop.
It's open source. Once again, timescale.com/ai.
Welcome to another episode of Practical Practical AI. This is Daniel Witenack. I am the CEO at Prediction Guard, and I'm joined as always by my cohost, Chris Benson, who is a principal AI research engineer at Lockheed Martin.
How are you doing, Chris? Doing very well today, Daniel. How's it going?
It's going great. I was saying that I'm really pumped to be talking about something that's near and dear to my heart, over many, many years Because today, we have with us Jan, who's the CEO at Probable, and Guillaume, who's an open source engineer at Probable.
Welcome. Thanks for having us.
Well, Jan and Guillaume are working on data science that you own, including projects like scikit learn, which is, of course, very near and dear to me along with other data scientists, all around the world. So, Yan, if if you could, since you're coming from the CEO perspectives, help us understand a little bit. Maybe for those that have heard of scikit learn or some of the other projects that you're involved with, but they haven't heard of probable, if you could give us a sense of what is probable.
As you mentioned kind of in the lead up to this conversation, it's a slightly different kind of company that came about in different sorts of ways, than other types of startups. So, yeah, if you could give us a little bit of context, that would be great. Well, very glad to be on the show with you tonight.
Probable is a company that is typically known as a spin off from a research center in France called INRIA. And INRIA is the place where this technology, Psychiclearning, has been developed over the past ten, fifteen years. Not many people know that, and the project has been somewhat protected and sort of incubated within that research center.
And after all that time, as you know, Psychicarn has been adopted or even probably participated in creating the field of data science because it is applied math and essentially has created a sort of paradigm for how data scientists approach data science, typically through two functions, fit and predict. And the French government has a national strategy for AI like many many countries. And the government decided to double down on Psychic Learn.
And, they came up with a budget. They entrusted the research center with that budget. But then they also asked for the project to be breakeven at some point.
And the team said, okay, breakeven is fine, but we don't do that in the research center. We don't breakeven. So why don't we call an entrepreneur to try and help us figure it out?
And they called me. So I have a track record as a software engineer and an entrepreneur in tech for the past twenty five plus years, but I'm not a data scientist. So I did my due diligence, and I sort of, you know, dug deep to find out what this project was all about under the hood.
Is it any good? Is the community any good? And of course, PsychicLearn is this quite amazing jewel of the technology that every data scientist on the planet uses that discovered that it was downloaded 1,500,000,000 times, cumulatively 80,000,000 times a month, 22% in The US, only 3% in France.
So this is a project that is used all over the world and Probable is essentially the spin off that takes all of the team, including Guillaume here, from the research center and turns it into an open source company that inherited the mission that was initially given to the research center. And the mission is to build a suite of open source technologies, including Psychic Learn, but above and beyond Psychic Learn as well, for data science. So the scope is large.
The mission is noble. And this is what we're building, essentially. So Probable is a one year old company that has already started doing many, many things.
used again by every data scientist on the planet. Well, this has brought up a lot of interesting questions on on my end, and I I really love the part of your pitch and at least how you framed it on your website and in your materials online about data science that you own and the open source side of this, which I know from experience, there can be some interesting challenges around finding business models that really work with open source technologies. And we've seen technologies where companies start with a posture towards open source and then gradually become more closed over time.
So I'm wondering from the leadership perspective, it sounds even in the way that this company was formed that there's a posture towards stewarding the, you know, scikit learn and these types of projects. But from your perspective, what what is your posture towards stewarding these projects in the open source side? And how do you view the business element of of this to make it sustainable in the longer term?
So that is the hard question, but it is the one that is important here.
Psychicarn is a technology that is, again, applied math. It's not rocket science, but it's applied math, and it's quite intricate. The thing is, the scientific community uses it, day in day out, and everyone depends on this.
So typically, when I when I discovered the scope of the project and the mission that was entrusted to the research center, I realized that this project is bigger than me, number one. Number two, the mission is to actually create more open source. In other words, you know, in 2024, it's even more acute.
Typically, big tech keeps on amassing so much power concentration, and we could argue that they do not distribute as much as they should. So that's not a judgment, but it is a fact. And Psychic Learn is precisely the contrary.
It actually enables so many companies to to do data science. So with that in mind, before creating the company, we decided to craft sort of architecture for the company that would respect that. And so, you know, before we created the company, before Guillaume joined as a cofounder, before even incorporating the company, we had a template that actually created the the governance, the shareholding structure, and also leveraging a new law in France that allows us to do a sort of b corp.
So a company with a mission where the mission is clearly stated in the bylaws, and that mission is to create open source for data science. So in a way, we've created a sort of constrained environment that is unlike many companies because it's by design. This company by design has created guardrails so that the governance cannot take this company too far on the right, let's say, proprietary technology or even changing the license.
That's not in the cards. And we've created a sort of mechanism where, you know, if we do not uphold to the mission, then we can actually lose some of the assets such as the brand. We are the official brand operator, but the brand belongs to the research institute.
Right? Still. So there are many mechanisms, trigger mechanisms that force us, including shareholders that we would would bring in, to actually bind with the mission long term.
Gotcha.
raised so many questions for me that I wanna ask. I actually wanna take just a moment and kind of go back, because it occurred to me as we're talking about this, for some folks listening who may have never even used scikit learn, they might have heard the name stuff. And you talked about it being applied math.
Could you guys expand on that a little bit for somebody who hasn't had a chance to ever actually utilize it themselves, in terms of what it's doing, and kinda catch them up to us in the conversation a little bit? And then I'm gonna pepper you with a few more questions because you've got me really interested. You hit so many topics on that last.
Yeah. So maybe I can give a bit of background.
So second, basically, the tagline is machine learning in Python. So it goes back to the, let's say, statistical rules. So the simple answer is we try to make pretty predictive modeling.
So try to use mathematics to form data in the future to give answer to a specific question to a specific paradigm. The big difference with generative AI or deep learning is just that all the statistics that you have there are simple steps. So they are fundamentals.
And deep learning build on those, but are just like much more, let's say, costly to train, costly to in the inference states, or not in the same scope as well. And Cycogen is like the de facto choice when you want to have Tableau data, so Excel spreadsheets data structuring this way. So that's a de facto way of training, like to be able to eat those spreadsheets and give back some labels or some regression, let's say.
And whatever is like image or NLP or like, this is like, let's say, deep learning and transformers is more in that area. So we are more like back to what was machine learning like a few days ago, but that have many, many, many applications specifically.
Well, considering how incredibly popular and foundational in the data science world, could you kind of give me a little bit of a landscape view, and I'm not sure which of you would be the right one to answer, so you guys pick between yourselves, but a little bit about kind of how that fits into the data science landscape with AI coming in, just so that with people listening, they can kinda go, ah, I see how it fits into the many organizations and tools that are out there. How do you think about that, for that? And then I'll get back a little bit more to the organizational stuff that Jan was talking about a few minutes ago.
maybe I can answer partly, which is by giving use cases and to see that, for instance, with a partner as well that we worked over the years, to give where do you find machine learning? And for instance, machine learning can be found in health care, where you want to know if a drug works or not. Then if you want to find diseases as well in some type of data, It could be as well like fraud detections in banks, in insurance, predictive maintenance, and those type of all applications that we have since many years.
So let's say the use cases are very, very large. And what brings Cyclone on that is that this is not for one of the use case. I mean, it was thought from the beginning to be general enough so that you can apply to any of those use cases and to come back to, let's say, classification and regression programs, let's say, or unsupervised learning as well.
in that field. So maybe, Jan, you have something more to add. Perhaps also at the macro level is to say, you know, scikit does a lot of things, including deep learning.
But in to be to be frank, when you wanna do deep learning, typically, you'd go to PyTorch or TensorFlow. But for everything else, you know, scikit learn. In other words, in the great AI family of algorithms, there is machine learning.
And within machine learning, you have deep learning. Within deep learning, you have other categories of algorithms such as, you know, transformer based models that lead to LLMs. So it's basically, you know, Russian dolls of sorts, and psychic learning is the biggest provider of algorithms in the machine learning space.
And in fact, if you look at the downloads, typically, Psychic Learn is downloaded as many times as PyTorch and TensorFlow combined, which is crazy because now everyone is talking about, you know, LLMs, of course, but also deep learning because deep learning is currently in a spring state, not quite a winter yet. So, of course, deep learning and Gen AI is a wonderful breakthrough. That being said, you know, I like to simplify sometimes the eighty twenty Pareto distribution.
So I had the intuition that 80% of the use cases out there use Psychic Learn when it comes to machine learning. And people actually tell me, Yan, you're wrong. It's more like 95%.
Right? Because in terms of, you know, technology that is robust, that is tried and true, that is used to actually, you know, turn a profit or return on investment. Banks and insurance companies, right, Guillaume was mentioning fraud detection.
Fraud detection typically uses psychic man, and that actually saves money. That actually you know banks would be losing money without that. So it is actually quite essential but again it's applied math right so scikit learn is only a facilitator to this category of problems.
What's up, friends? I'm here with a friend of mine, a good friend of mine, Michael Greenwich, CEO and founder of WorkOS. WorkOS is the all in one enterprise SSO and a whole lot more solution for everyone from a brand new startup to a enterprise and all the AI apps in between.
So, Michael, when is too early or too late to begin to think about being enterprise ready?
It's not just a single point in time where people make this transition. It occurs at many steps of the business. Enterprise single sign on like SAML, auth, you usually don't need that until you have users.
You're not gonna need that when you're getting started. And we call it an enterprise feature, but I think what you'll find is there's companies when you sell to like a 50 person company, they might want this. Especially if they care about security, they might want that capability in it.
So it's more of like SMB features even if they're tech forward. At Worker West, we provide a ton of other stuff that we give away for free for people earlier in their life cycle. We just don't charge you for it.
So that auth kit stuff I mentioned, that identity service, we give that away for free up to a million users. 1,000,000 users. And this competes with Auth0 and other platforms that have much, much lower free plans.
I'm talking 10,050. We give you a million free. Because we really wanna give developers the best tools and capabilities to build their products faster, and to go to market much, much faster.
And where we charge people money for the service is on these enterprise things if you end up being successful and grow and scale up market. That's where we monetize and that's also when you're making money as a business. So we really like to align, you know, our incentives across that.
So we have people using Authkit that are brand new apps just getting started, companies in Y Combinator, side projects, hackathon things, you know, things that are not necessarily commercial focused but could be someday. They're kind of future proofing their tech stack by using WorkOS. On the other side, we have companies much, much later that are really big who typically don't like us talking about them, their logos, you know, because they're big, big customers.
But they say, hey. We we tried to build this stuff or we have some existing technology, but sort of unhappy with it. The developer that built it maybe has left.
I was talking last week with a company that does over 1,000,000,000 in revenue each year, and their SCIM connection, the user provisioning, was written last summer by an intern who's no longer obviously at the company and the thing doesn't really work. And so they're looking for a solution for that. So there's a really wide spectrum.
We'll serve companies that are in a, you know, their office is in a coffee shop or their living room all the way through they have a, you know, their own building in Downtown San Francisco or in New York or something. And it's the same platform, same technology, same tools on both sides. The volume is obviously different and sometimes the way we support them from a kind of customer support perspective is a little bit different.
Their needs are different, but same technology, same platform, just like AWS. Right? You can use AWS and pay them $10 a month.
a month. Same product. Or more.
Yeah. For sure. Or more.
I don't know. Well, no matter where you're at on your enterprise ready journey, WorkOS has a solution for you. They're trusted by Perplexity, copy.
ai, Loom, Vercel, Indeed, and so many more. You can learn more and check them out at workos.com.
That's workos.com. Again, workos.
com.
So, Yan, you were kind of already going there, and I I love the direction that you're going with this. But I I think maybe I could tee up a a softball for you here because I'm I'm personally passionate about the the answer to this question, and you probably have a a better view on it. But there's there might be people out there maybe listening to this podcast who are thinking, well, now that we have Gen AI, we have large language models, I could put in a prompt to one of these models to do fraud detection or to find entities in text or to make some prediction of a classification.
And, you know, sometimes that works. And so maybe there's people thinking, well, there's these general purpose large models out there. How does that change the way that something like scikit learn plays in an industry?
And I personally would argue and and think that this actually makes scikit learn more valuable, if anything, rather than than less valuable in terms of the ways that it can be combined even as a tool that's orchestrated with GenAI models. But I'm I'm curious your perspective on this from the business side, and maybe Guillaume has some ideas on the technical side. Yes.
Psychicarn typically is this one technology that is patrimonial. In other words, it belongs to everybody. In fact, there is another stat when you look at, you know, the the the figures that are public, by the way, the number of dependencies.
So Cycoturn is is actually used by nearly 900,000 projects on GitHub. So there's nearly a million projects that depend on Cycoturn. And there's a new law that I discovered recently, someone mentioned that, Linde's effect, which means that something that's been used long enough will remain important for long enough.
So not saying that Cyclerean will go the way of COBOL, but Cyclerean is here to stay, and we are with the community, the guardians of that. So we're gonna make sure that Cyclerean remains there forever for companies that actually need it in a stable version. And, of course, Guillaume and the team are building up new features as we go.
Right? So there's a a dedicated effort, and and I should say that we have carved out nearly, you know, 10 people in the team are doing only that, contributing to Psychicarn and the other associated libraries. Now your question, Daniel, is is, you know, whether Psychicarn will be obsolete in, say, a number of years because general purpose technology has made it, you know, irrelevant in some ways.
Number one, Cyclean is extremely frugal. It actually works on CPUs, and it is well controlled, well understood. It's actually quite predictable in some ways, whereas deep learning is usually known as a black box where it's really, really hard to introspect.
And so Cyclone does produce for certain categories of problems things that are actually working quite well, more so than large language models for sure today and and more so than any sort of deep learning based technology that we understand today. Now it is possible that with additional data, additional training and techniques, and even evolutions on on the transformer based model, we could improve and probably render obsolete cyclotron. But to us, and then Guillaume and I, we talk about that and with the team, we also experiment with other lens.
And we are also trying to figure out how we can use these new technologies to actually help our first persona. And that is with this data scientist. So we are a technology provider to help data scientists, and increasingly so, the data scientists in enterprises, because we will be creating value adding services and solutions so that we can generate revenue to sustain our mission.
So the goal for us is to actually project ourselves while contributing to open source, but also create a sort of business value proposition not dissimilar to Red Hat, because that is the closest type of company that we identify with in terms of spirit.
you know, I'd like to get back to something that you said earlier that that feels like you're kinda tying back to it anyway there. And that's that you talked about the mission to create more open source, and the mission that you're trying to create this environment that you're describing by design, you said, which is that with scikit learn here to stay for the long haul, it's gonna be something that is not going away soon. It's solving such a high percentage of the problems.
Could you describe a little bit about kind of what you're thinking around that in terms of further developing this particular set of software and the ecosystem around it so that we have the benefit for many years to come? How are you approaching that?
the company is built with multiple business units, if you wish. That's that's a big word for a start up. Right?
But we have multiple, revenue lines and multiple activities even within the open source team, which is dedicated. So Guillaume perhaps can elaborate on some of the other libraries that we support that complement Cyclone. So, you know, that's one way to answer the question.
But also, we are building a new product, which I call reversible SATs. So we are building a product that will provide additional value to data scientists, and the goal is to create a sort of I don't want to use the the term copilot because that is too close to LLMs, but it is the spirit. We are building a companion to augment the work of data scientists all the way to teams.
So that is an additional product on top of SecuKern because SecuKern just works. And so we don't wanna change that. And contrary to a company that would build a SaaS solution with a proprietary approach, we wanna say, okay, whatever you guys use is fine.
We need to find a way to add new value, and some of it will be open source, certainly modular. But for those companies that have more money than time, that need more service than beyond their own, we'll have a solution for you and will make your life easier. And, data scientists are a new breed.
It's a new type of job. It's not been around for very long. And in a way, when I talk to people so I've been in code forever, and you know this.
Right? The the developers, when they get hired, they are turnkey in some some ways. Right?
They have their Git environment, and they know how to pure code, and that's all pretty standard. But when you talk about data scientists, it's actually quite artisanal. It's an art and a science at the same time.
And you're you're manipulating two objects, actual code, but data scientists are not coders, and you're manipulating actual data. It's not code. It's patterns.
And so data scientists are have a difficult task, which is to combine these two things and create value for the enterprise. And then they talk to business units, and they're like, what do I do with this model? How do I put it in production?
Right? So there's a huge conundrum to solve, and that's what we're gonna do additionally to building open source that are modules that people can use. Maybe Guillaume, you can elaborate on some of the other libraries that are key, to actually help.
Yeah.
Within Probable, so we have the open source team. And so we worked for many years on Saketon already. But we see the importance and as a community, we see the importance of putting models into productions and as well getting closer to the data sources.
So we are just working on libraries that should make those come together. So for instance, we have a library that is more on the MLOps side that is called SCOPS. We work a bit to make, let's say, persisting more secure in some way.
But we look as well on how to bring databases like SQL words into like closer to the machine learning models. So like how can you transform data with states, with different tables, and how you can be in your Python words. We were caring so much about like SQL, for instance, and how you can bring this into scikit learn.
And we in scikit learn as well, we want to improve like whatever is visualization evaluations, inspection of models, which is on the top of just like training an algorithm. Because this is a so we want to augment all those aspects like beyond us. And either it's in Cyclone or either this is like library connected to Cyclone, let's say.
So the one before is called Scrub, by the way. So it's like scrubbing data. So it's Scrub and Scops are two libraries that we look at.
you have this robust open source contributor community built up around scikit learn and the various projects within it. How does Probable work with those? How have you guys set up that relationship?
What does the governance look like on that? Because you have both your core team that you alluded to earlier that's working at Probable on this, but you also have that larger open source community. How does that all work?
Can you kinda tell us how that's evolved? I imagine it's quite mature by now. And that's the point.
means that by design, we decided to not affect the license of SecuClient. We're not branching it out. We're gonna care for it.
And so the governance of scikit learn being so sane already means you don't touch it. If it ain't broke, don't fix it. So the the governance is unchanged.
So the the center of gravity was at INRIA, the research center, but also involving people all over the world. I I don't know. Guillaume, many contributors?
Maybe 200?
Oh, even more. I think, like, in a year, you have more more than that. You have maybe, like, maybe three, four hundred.
And the core team is, like, let's say, of it might be around France, around Paris, around Probables, but then there's, like, another half of 10 ish person around the world that contributes very like almost every day, let's say, by communicating with the community. And as Jan mentioned, we didn't want to change that. Nothing changed in that regard.
So the only thing that we actually did more we did more to bring transparency. So to explain to people. So now that we're in Probable, we feel that because we are a private entity, we need to communicate what are we doing and what are our roadmap and which committee items are we going to work on just to like bring more trust such that, I mean, we don't go like in the dark and that nobody knows now what we're doing.
So we try really to pay attention to every six months to mention which of the items that are defined by the community, that are not defined by probable, but from the items, which one we have the capacity to work with the human resources that we have, let's say, at hand. So we really want to show that.
And then by design, the open source team that is full time on Segiturn and other open source libraries means it's a cost center to the company. So that cost center is by design, and we know that's a cost we have to cover. So we we will cover it through different types of activities.
So for instance and this was something that was done in the past where brands were sponsors. So either they hire someone that becomes a core developer, and they're naturally sponsoring someone to build up this technology, or they were giving money as a donation to the research center. But now that the team is with us, we are translating this to into a contractual sponsorship framework.
And so, you know, brands who wanna contribute to Psychic Learn and help us compensate for salaries will get something in return, exposure. And and and if they actually put more money into it, then we'll have a conversations around the road map. Find a way to make it converge in a win win kind of way.
Because Guillaume, for instance, can say, you know, this brand wants us to do something, but it makes no sense for the community, then we won't take their money for the sponsorship type of business. Right? However, if companies want to pay us to do a certain type of paid for software, we'll do we'll look at it.
But that's a different branch of the company. So we've we've really clearly separated. And by design, we know there's a there's a cost to it.
And that cost is actually, if we are doing well, it's compensated by the fact that we have done good by the brand. In other words, hopefully, the the community will actually resonate with what we're doing. And so they'll pay us back by actually appreciating what we're doing, which will carry the message further.
So we think that there is a self fulfilling prophecy if we actually keep adding value to the whole scheme as opposed to removing value, and I will not name certain projects that have chosen a different way. But on the other hand, going back to the governance of the company, when a company flips and becomes VC funded or only VC funded, VCs require a sort of return on investment that is too radical. And so that sort of forces a change of posture with vis a vis the community and the licensing scheme.
In our case, we've actually created a structure that is balanced in terms of shareholding groups and so we will ultimately have or that's the goal of the structure of the architecture is to have as much money from public support than from private support. So it's again sort of balanced.
You know, when we started podcasting back in 2009, an online store was just the furthest thing from our minds. Now we have merch.changelog.
com, and you can go there right now and order some T shirts, and that's all powered by Shopify. What do we do before Shopify? I'll tell you.
We did nothing. We couldn't sell. There were other ways, of course, but they were very hard, very difficult.
Shopify let us build out an entire front end, obviously branded like Changelog is. It's amazing. Merch.
changelog.com. And our favorite feature is we use their API to generate a new coupon code, a personalized coupon code for every guest that comes on our podcast, and they get a free t shirt from our merch store.
And that's so cool. They choose the shirt they want. They use the coupon code.
It arrives free of charge to them, and life is amazing. But, also, you can go there right now too, merch.changelaw.
com, and buy some threads yourself, and that's awesome as well. So upgrade your business and get the same checkout we use with Shopify. Sign up for your $1 per month trial period at shopify.
com/practicalai, all lowercase. Go to shopify.com/practicalai to upgrade your selling today.
Again, shopify.com/practicalai.
So as we come back out of break here, I wanna turn to kind of a fun question for you, and I'd like each of you to take a swing at it because it's not specific to being the CEO or doing the technology itself. If each of you could describe kind of a cool use case, something fun or interesting, or that's really captured your imagination with scikit learn, and kind of share that with the listeners in terms of something that just kinda really took you as your thing. I I'd love to hear I'm expecting it to be a bit different coming from each of you and your different roles, but I'd love to hear kinda how how you see that and what's the thing that sticks out in your mind.
Jim, you start because I have to think about it now.
So it's a very technical one, let's say. But so during my PhD, I was doing classification, which is something that I was trying to find people that has a specific type of concept, so prostate cancer, versus people that didn't have it. And inside that space, you had one fairly specific frame, which is called imbalance data.
And it's what introduced me basically to Secuen, because I had that frame. I was using Secuen for the specific issues and how to tackle down those type of issues. And what is really funny is that it's how I got introduced to Cyclone and speak, for instance, with the developers.
And I developed one library, which is called Imbalance Learn that is merging as well with scikit learn, and it's compatible in some ways. And for many years, I maintained that package even when I was mentioning scikit learn. And over years, years after years, we did everything by the book, basically.
In title library, we implemented the algorithms that were inside the literatures, and everything was fine. Until that, as part of Innoea and now Probable, we have as well time to educate ourselves and to try to as well then bring through the recommendations of Cyclone to explain some concept to people. And by doing this, we find out that most of the research there didn't look at the prem properly.
And by communicating with other Core Dev, we just found out that a huge part of this thing was just wrong and that you should look at it in another way. And then it's pretty funny because with this, we found like some useless stuff that was, for instance, inside an imbalance run. But then now we have like better content.
We went to conferences to explain these problems, and people start to tell us, oh, yes, actually, that's right. And it's fun that you come and say that whatever you were doing like five years ago or ten years ago is actually like obsolete or not good or, I mean, wouldn't expect from there. And it's something that I find really like fun when you do open source because you can like, you are just here to contribute to something and just to bring, like, the best of what you do to everyone.
And everybody will be, like, thankful for that. Even I mean and you are not defending your own, like, let's say, scientific paper or, like, that's all what is true. And and for me, that's, like, one experience that come from my PhD from, like, now eight or nine years ago to up where I am now.
And then, like, I see an evolution where I was with very good people and you could correct errors that you do in the past. And actually, that will benefit everyone afterwards because that's landing inside the documentation of Securian or even inside the library. And then everybody will just, it says, millions of users will be affected and say, oh, actually, that's good.
And this is something that I would have stayed in academia, for instance.
the books, pursue books and go like this. But that's one anecdote, let's say. That's good.
I might be the CEO, but I do have the imposter syndrome because psychic learning is so impressive. It's day in, day out. I mean, that team and Guillaume is very humble and very discreet, but the amount of knowledge and the amount of technicality that is trapped inside this library is mind blowing.
And you haven't met the other members of the team. It's pretty much very, very hard to compete in terms of the amount of CPU cycles that go in there. So scikit learn is the gift that keeps on giving in some ways.
And the team is just out of this world and nice, and it's just a pleasure to work with that team all the time. Now the more I discover Psychic Learn and the more I find it amazing because of what the brand means to people. And so last week and and today, actually, we just released, and if you allow us to actually put the the link in the notes Of course.
Absolutely. We released the very first official Psychicarn certification program. And what's amazing is that we so this is the first time.
So we're doing it step by step, and the system works. People can register. They can pass or fail the test.
But without advertising, we had, like, within a couple of days, 600 registrations all over the world. A lot from India, actually, because people in India, they do work also remotely for other clients across the world. And so they do need a stamp of approval to showcase their ability to provide a service.
So very interesting that this brand almost instantly can promote a sort of service that is value adding. So that's the one thing. But then on the more technical level, I fell in love with one new feature that came out with 1.
5 of Psychic Learn developed by another co founder and and core developer, Jeremy. And that is the callback feature. Why?
Because Psychic Learn, in fact, is a platform. It is a platform. And the callback feature allows us to to provide extensions, if you wish, where people can hook into their inner workings of as they are building new models.
And in fact, I find that to be essential because we are entering an age of liability with regards to AI. Companies need to be able to introspect. They need to actually find out why the model is producing such and such results.
And so introspection is is critical. And as I said earlier, deep learning is sort of a black box type of approach, which I love, by the way. Again, in 1992, I was building deep learning models in the middle of the winter of AI at the time.
But psychic learning is actually quite introspective, quite transparent, frugal, as I said. And so callbacks are yet another feature that provides actual introspection into how we build models. Because talk about insurance companies, fraud detections, you've got human beings at the end of the of the spectrum being handled by algorithms.
And so that is critical. And I think we we fulfill a very important need with these features.
with this tool. And, you know, speaking of of this, team, Guillaume, I'm gonna throw a question at you. People out there have been listening to this.
They're they're kinda going, okay. I wanna I want to dig into this. So you're gonna get some new developers that are gonna come.
How should they engage? How should they find and get started in the projects to develop? What's a good onboarding path for those developers?
Probably the best onboarding path is, like, if you have a chance that inside your local community, there's some people that do what we call first time contribution to open source or like coding sprints. Go speak to those people because, I mean, they will help you to get on board. But then, like, if you are behind your computer and then you don't know where to start is where we have documentations that describe what do we call contribution, because contribution is not only coding.
It could be speaking, debugging, documenting, organizing sprints, and those type of things. And we have what we consider as contribution, and how you can help, basically, and where you can help. So of course, the natural thing is to come and code.
And then we explain you how to start with that. So this is on the, like, the documentation channels, the documentation web page. And and afterwards, everything is online and public.
So there's nothing private. So if we have different channel of communications, the main one is GitHub, and it's going through the issue tracker or the pull request, like depending of which side you are. And the core developer will be, I would say, twenty four hours over twenty four because we are around the world.
So that's like if I'm sitting there's somebody else in Australia or in The US that can just answer to you. And then we'll just give you feedback. And this is where your journey starts.
You should not be shy, and you should not be scared of making a mistake, because we are not judgmental. We all started by that stage of saying, I don't know what I'm doing, and I need to ask people what should I do, and that's a normal step. And afterwards, you just grow with the community, and then the the community bring you over.
I mean, but the most difficult thing is, yeah, is the first step, like engaging and saying like, so I'm the impostor syndrome as well. But that people say, I don't want. Mean, like this, like those very skilled people, they will never want to speak to me.
And that's not the case. So just come and just try your best and then people will just communicate with you for sure. Great guidance there.
As we wind up, I'd like to get, for each of y'all for both Probable and for scikit learn, kind of what you think about for the future. And I'll let you define what time span the future is, whether it's a few months or years out. But I'd really like to wind up, paint us a picture of when when the duties of the day have finished and you're just relaxing and you're thinking about what's possible going forward.
What do you think about? I'll go with the mission.
The mission is bigger than me, bigger than us. And so that's why the governance creates a self sustaining model. So, of course, you know, it's not trivial.
So there's a lot of work to achieve the mission long term, but that mission ends up with an IPO. In other words, this company is not meant to be sold or wrapped up. The goal is to do an IPO so that this company can carry on with the mission and allowing people to invest and be part of that story.
And that's why earlier, Daniel had asked a question about, you know, investors and all that. So we do have 70 individual investors, including people who were contributors or are contributors to Psychic Learn who don't have the chance yet to be employees full time of the company. So the goal is to create this sort of dynamic vehicle.
And if we look at the North Star, there is no such company today that is the provider of open source machine learning technology. That company does not exist, and we aim to be that because we need that in an age where there is too much concentration within just a handful of players. That's not okay.
It's not okay for the global South. It's not okay for Europe, which is lagging behind. But it's not even okay for The US.
The US may have big tech, but that's not okay as a single model. We need people to own their data science. That's why that is our tagline.
That was good. Guillaume, what are your thoughts?
Yeah. So maybe more on on so on Probable, I'm I'm really thinking that we have missions, let's say, to help more data scientists. But I will speak more about, like, about Cyclone and and the ecosystem.
So for me, the mission is we should stay focused on what's happening out there and make sure that Cyclean is still relevant. So we have the foundational model. That's fine.
But we need as well to understand where this is deployed and how this is used, because we can make such progress that bring, for instance, make it easier to bring databases to Cyclone or to bring Cyclone models into productions and to reduce friction and everything. And as well, bring values on understanding the model. I mean, are speaking about AIX as well in Europe for now.
So I'm sure there's plenty of, let's say, areas where we can have real impact. And then there's, well, technology that's moved very fast. So for instance, before we knew pandas, now this is Polar.
So we need to move in a fraction of sequencing, how do we deliver value to the user that just makes the switch and still can use scikit learn? Like, can we do like accept those things? And then so we have to make this audit of what's happening.
So this is difficult to say where we will be in five years. Because in five years, we have all those things that can, let's say, we have the full chain of machine learning that probably will be here. So we should be aware, but we should be aware of whatever moves very fast around us to stay relevant.
That was well said too. Gentlemen, you guys have done a fantastic job of, of teaching the rest of us, about this, and thank you very much for coming on the show today. You're welcome.
Always a pleasure.
Alright. That is our show for this week. If you haven't checked out our changelog newsletter, head to changelog.
com/news. There you'll find 29 reasons. Yes.
29 reasons why you should subscribe. I'll tell you reason number 17.
You might actually start looking forward to Mondays. Sounds like somebody's got a case of the Mondays.
28 more reasons are waiting for you at changelog.com/news. Thanks again to our partners at fly.
io to Brakemaster Cylinder for the Beats and to you for listening. That is all for now, but we'll talk to you again next time.
Shared via Hopper