Open Source as a Business — Models, Paths, and Practice
Alexander CS Hendorf, Ines Montani, Sylvain Corlay, Yann Lechelle
Building a business around open source requires distinguishing the software's distribution method from the commercial business model. Open source serves as a community and marketing asset rather than a revenue stream itself. Common monetization paths include offering specialized consulting services, creating complementary proprietary products, or developing enterprise-grade SaaS layers that address needs not met by the core library. For example, the creators of spaCy developed Prodigy, a proprietary data annotation tool, to support power users of the open source library. Similarly, QuantStack grew from a single-person consulting firm into a 30-person organization by providing expert support for the Jupyter ecosystem, Apache Arrow, and CondaForge.
A more structured corporate approach involves creating mission-driven, for-profit companies. Probable utilizes the French Loi Pacte to legally mandate a duality between generating shareholder dividends and maintaining scikit-learn as a state-of-the-art open source asset. This model allows the company to raise venture capital—such as Probable's 18 million euro seed round—to build commercial products like Score while keeping the core library's governance and permissive BSD3 license intact.
The rise of AI agents is shifting the landscape by polarizing usage toward established libraries. Because LLMs are trained on massive open source corpora, they prioritize stable, well-documented libraries with high backwards compatibility. To adapt, developers are creating serializable APIs and comprehensive docstrings to make tools more discoverable by agents. Strategically, open source is viewed as a critical component of European digital sovereignty and economic resilience, providing a way to reduce dependence on US-based hyperscalers and avoid the risks of proprietary "black box" systems.
This description was generated by Open-Source AI using the transcript of the session and the original submission contents.
Submission
The proposal as submitted by the speaker before the conference.
The discussion centres on concrete experience — different starting points, different ecosystems, different business models — with the shared thread being open source as a deliberate professional and commercial choice.
Key Questions:
Different entry points: You each came to building a business on open source from a different direction. What drove that decision — and what did you not expect?
Where the business actually starts: Open source is the foundation, not the product. How do you define what you sell, and to whom?
Community and commerce: How do you maintain trust and credibility in an open source community while running a commercial operation around it?
Open source and AI: The AI landscape is consolidating fast around closed systems. What does that mean for open source projects and the businesses built on them?
European perspective: Is there something specifically European about the way you think about open source as a business — around sustainability, sovereignty, or independence?
Advice: What would you tell someone who wants to build a business on open source — or switch to doing so — and has not yet started?
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:00]
You can also come up.
Speaker 2 [00:08]
as a moderator i go here
Speaker 1 [00:12]
Get some water. Yeah. Nice diversity. Maybe.
Speaker 2 [00:16]
Maybe we can offer that.
Speaker 1 [00:16]
. Or?
Speaker 3 [00:18]
Ha ha ha.
Speaker 2 [00:21]
Distribute some water just in case.
Speaker 4 [00:27]
Yeah, we take the big bottle.
Speaker 2 [00:31]
So, yeah, you can think, yeah, that's fine.
Speaker 3 [00:34]
Let's close them.
Speaker 1 [00:40]
Yeah, I'm kind of worth it like yeah, I'm better doing it with a lighter and I like Berlin style. Yeah With my teeth, but I'm not gonna do that here
Speaker 2 [00:52]
chance to
Speaker 5 [01:06]
Video team, please change the slide. Yeah, thank you.
Speaker 2 [01:10]
Yeah.
Speaker 5 [01:17]
Hello, everybody, and welcome to our panel, Open Source as a Business, Models, Path, and Practice. Our guests today is Ines Montani, Jan Leschel, and Sylvain Corley. So this panel is moderated by Alex Hendoff, so I will hand to you.
Speaker 2 [01:36]
Thanks for the introduction. Thank you. Thanks for joining. This is also part of trying new things at PyCon and bringing other contexts and views to the conference. So I'm really amazed to see how many people are joining. So thanks for joining. In this panel, we want to discuss different backgrounds. the people who work with open source and also as a business and we want to go on backgrounds and the narratives and just like show different approaches how you can also build a business with open source or operate so um i would like to start with you young you used to be a ceo of scale way. I think you had like five, there were like 500 employees. Yeah. And it's also like data center. And, you know, it's now you're with probable, which is based on scikit. And I think this is like a very bold move. So what made you go like, I'd rather go open And so it's them just like staying on.
Speaker 4 [02:49]
Thank you, Alex. That's a very good question. So cloud providing is an infrastructure play, requires a lot of capex. So you need to spend a lot of money in actual hardware, physical facilities. But all of it actually is software. When you think about hyperscalers, the actual value lies very much in the platform as a service, right? So if you look at AWS, Microsoft Azure, GCP, it's about 400 products or more. And so this is the hard part, writing the software that actually allows these infrastructures to scale and clients to actually use it. And so in Europe, we have a number of cloud providers of small size, Scaleway being one of them, but you perhaps know of OVH in France, but also in Germany, Yonos and Hetzner, now Stackit. So we have these providers, but they're tiny. They don't scale because we have not reached that critical mass. And most importantly, we consume from the main players, usually U.S.-bound. But China is also strong in doing this. Why? Because it's all about software. But I've come to realize that as an entrepreneur in France, in Europe, for the past 25 years, I have failed to create a big tech. and so have we collectively so I came to the conclusion and I'm not from the open source movement I'm not from the open source tradition even though I do have some hacks that I can talk about as an anecdote but I came to the conclusion that we have perhaps no choice but to adopt open source as a European strategy to catch up because it's all about making sure that the software gets distributed to as many people as possible to create a level playing field so that so that we can catch up and build new value on top of it so that is the reason why i chose this project but the project shows me as well so it's that's the probable story awesome
Speaker 2 [04:55]
So, Ines, you started space, I think it's 10 years ago?
Speaker 1 [05:01]
Ten years ago?
Speaker 2 [05:01]
Yeah.
Speaker 1 [05:03]
Actually, the company we started about, yeah, exactly ten years ago.
Speaker 2 [05:05]
Ten years ago. Ten years ago.
Speaker 1 [05:06]
Ten years ago.
Speaker 2 [05:06]
So tell me, what was your story like, Spacey? I know you and Matthew met.
Speaker 1 [05:10]
In Berlin, socially.
Speaker 2 [05:11]
In Berlin, socially. In Berlin, socially. And then Matthew, you don't know, was like the natural, the linguist at the time. Or like...
Speaker 1 [05:20]
Also the original author of Spacey.
Speaker 2 [05:20]
Also the... Yes, and you just met and then you started this. What was your story? And also, tell us a little bit.
Speaker 1 [05:32]
Yeah, so yeah, my background is more in linguistics as well, and also I've always developed. I started making websites as a teenager, when I was 11 we got our first computer and I realised I can put things on the internet, but I actually didn't go into programming, so I don't have a classic computer science degree, because when I was at the point where I had to decide what am I going to do with my life, I didn't really identify with the boys from the computer a club. I didn't feel like a developer. When I met Matt, we realized, oh, there's a lot of synergy. There's a lot of places where our skills fit together. I'm doing a lot of product. I can build apps.
Speaker 2 [06:10]
was it just like let's do this out of idealism because it's cool and it ails people or was it already we should also do it we could also make a living from that
Speaker 1 [06:20]
Yeah, so he left academia when he saw that companies wanted to use his research code. And his research code was really just, it was intended to print a number and exit. It wasn't meant for production use. So companies started wanting that and wanting to license that. So he realized, oh, at the time, there wasn't really any library for NLP that was designed for production use and for practical use. So he left academia because he got a small inheritance, so he was able to take a few months off. And so he started writing Spacey. And that's when we met. And the plan was always like, hey, if we can, as you said, like get it in the hands of a lot of people, there will be an opportunity to make a living off that at least. So that was always the plan. We started a company together. And yeah, and then very early we saw that like, okay, we needed kind of a more product strategy to monetize it because at the time a lot of people thought, well, how should we monetize Spacey? Well, one way would be put it in the cloud. And we didn't see like NLP APIs at the time. that was not the type of use case. People need to develop with developer tools. You need to write code. So we can't just put it behind an API and ask for money. And the other option is, well, just do premium support, have people pay for that. But that also kind of conflicted with how we wanted to present the library because we wanted to have good docs. We want to make it easy for people to get started. But if our business model is charging people money to help them get started, you're kind of in this conflict. If our docs are too good, nobody's going to pay us money. if our docs are too shit, nobody's going to start using the tools.
Speaker 2 [07:46]
using the tools. And your first strategy was let's do Prodigy as an add-on product?
Speaker 1 [07:50]
Yeah, so it gets to a product and ask people to pay for that separately and especially something that appeals to power users of spaCy. If you're using spaCy, well, you need to create data at some point and you need an annotation tool that didn't exist. And also we offered it in a way that played to our strengths. So we were completely self-financed at the time and we offered it for an upfront cost, lifetime license, no SARS, because that really made it easy for us to just get something out.
Speaker 2 [08:16]
And… Then you said, let's do something with… Yeah.
Speaker 1 [08:19]
Yeah, let's put it in the cloud.
Speaker 2 [08:20]
Let's put…
Speaker 1 [08:21]
So that was, you know, that was the next step. Okay, naturally it needs to be in the cloud and it needs to have the data privacy that makes our products successful. Like, you know, we can't just have people upload their data to the cloud to annotate. So it's a very infrastructure heavy project. In the depth of COVID, we actually ended up meeting a US investor who really understood our project and our open source roads, which is also not, you know, not the norm. and we wanted to raise the money we need, not come up with plans that justify the maximum of money we could get. So we found a very good deal and we started hiring a team and really tried to scale up what we were doing, but we didn't actually manage to finish the product. So then after a few years, we were like, okay, let's go back to our previous way of running things, running from revenues and work on a new product. So it was a very tumultuous time. But it's also like a...
Speaker 2 [09:20]
It's also like a startup. Yeah, it's a... It's still like startup and, okay, you fail, you learn, you do something different.
Speaker 1 [09:26]
Yeah, and also, we were always, it was difficult because we were always kind of unusual.
Speaker 2 [09:26]
Yeah, and also...
Speaker 1 [09:29]
Even the open source way of doing things was always, you know, we were never like a typical startup. And that meant we didn't really fit neatly into any categories. But yeah, it's also, it was definitely, you know, very difficult to scale, traditionally scale what initially made us successful. And also the timing, you know, it was just before coding agents and coding assistants. So now we're like productive in a very different way.
Speaker 2 [09:55]
So what's your story with Jupiter and founding QuantStack as your company, and what was the reasoning?
Speaker 3 [10:04]
There wasn't any deliberate reasoning. It was very accidental, I would say, in the beginning. So when the Jupyter project was founded, I was already a co-maintainer of the iPython project, which was re-baptized into Jupyter at the time. And I was an employee of an American big tech that is Bloomberg, and moving away from being a quant researcher, quant finance researcher, to being more of an open source maintainer. And then for family reasons I had to relocate to France and I went to see my boss at Bloomberg and I told him that I was not quitting, I would just have to come showing up to work and be elsewhere. And so I started this one person company and the only goal was to support my own activity as a maintainer of the project and other related projects. And very quickly business came to me and the main reason is that most of the other maintainers of the project were either employees of large corporations or in academia, and no one else that was listed as one of the main authors was available to do consulting. So very quickly, I had more business than I could handle. So I hired one by one, and then we expanded the scope beyond Jupyter. And now it's a team of about 30 people, a bit over 30 people actually, not just in France but also in Germany, Austria, Spain, and in the UK. So a very European team. And the scope covers the Jupyter ecosystem still today, but also Apache Arrow and CondaForge and some other things that may not be as known, but like the Extensive Stack and XMD, which now underlie many accelerations in Firefox and other packages. So initially, this was really not a deliberate thing. It's a completely accidental startup. I think it was a setup.
Speaker 2 [12:02]
It was a step to leave Bloomberg, right? It was a step. And it was also, I think, your first client, right?
Speaker 3 [12:07]
So, yeah, Bloomberg was actually one of our first clients, but actually technically, even though we sort of already had an agreement to get started with them because they were my former employer, it took some time to materialize and we had other clients already paying checks before the one that sort of bootstrapped the whole thing. So, yeah, the bootstrap, the initial bootstrap was really accidental. Now we're in a bit more deliberate strategy on how we want to grow. So we started really as a service company, and the reasoning was that we were going to solve boringly utilitary problems that people were facing with the software. And by being merely useful to our clients, we were going to grow that way. It's only recently that we started getting into building products. than raising funds and trying to go fast. The product we're building now, I think, is really, again, something that is in a way boring and utilitarian. We are solving a problem that many of our clients in the consulting side of things were facing. Yeah, I hope it's going to be successful.
Speaker 2 [13:21]
But also, like, part of your offering is, like, parts of your team are also, like, deeply involved in the software and the open-source software. You also offer for consulting. So you have, like, the best people, basically, who work on it because they actually are creators or co-creators of the software to bring real value.
Speaker 3 [13:40]
Exactly. Yeah. The team is really singular in that many of the people on the team are award-winning authors of really important software, or C-Python core developers, Apache Arrow developers, Numba maintainers, and Jupyter, and so on. So, yeah, really people who've built tools that everybody uses already.
Speaker 2 [14:11]
No, so so open source is a foundation like how do you define what to sell and to who? Or is it just like who calls and hello sure it's like an open question
Speaker 3 [14:26]
Is this for me?
Speaker 2 [14:27]
I mean, for everyone, like, whoever wants to take first, like...
Speaker 3 [14:27]
I mean, whatever.
Speaker 4 [14:31]
Or perhaps the story of Probable. So perhaps a show of hands here. Who knows Scikit-Learn? Impressive. Everyone. So thank you, the Python community. Thank you for Cython underlying. Thank you, Cython, whom I met today. So 4 billion downloads, more than PyTorch and TensorFlow combined, actually. 200 million downloads per month. This is foundational, and it's open source as good as it gets. BSD3 license. So it's the reason why it was adopted so broadly. I have no merit for scikit-learn. I'm just the entrepreneur. But when I came to see the project, the project came also with its context. And very early on I said, there are as many business models as there are open source projects. Because open source is not a business model, number one. Some people confuse that. Open source is a distribution. It's a community. It's a federation. It's a governance concept. It's a marketing asset in many ways. And so very early when the team came to me, because the goal was to actually extract scikit-learn from the INRIA research lab in France, which was where most of the core maintainers were being employed, the goal was to extract it. So, from the get-go, to turn this into a business, the question was, what is the equation? So, it was deliberate in our case, because the French government said, okay, scikit-learn seems to be important, the numbers show, we're going to finance this, but...
Speaker 2 [16:14]
It's really big. I'm not sure if my math is accurate, but if you basically U present Psi, U present Psi, U present Jupiter here, I think we could statistically argue probably every human on the planet has at least downloaded once. In the scientific...
Speaker 1 [16:30]
Well, and then also, I mean, as you mentioned, if you go deeper down, you
Speaker 2 [16:35]
This is also the scale we have, it's like three titan open source libraries here represented.
Speaker 4 [16:43]
And so, I mean, nobody knows this, but open source software is just about everywhere, full stop. And it represents $8.8 trillion worth of value that if we were to recreate it, it would cost that much. So it is a societal capital. It belongs to humanity as a whole. But the question is, what is the arbitrage? what is the business model that actually binds to the underlying scikit-learn community in our case so scikit-learn is mature in a sense that has been around for 15 years adoption is global fine the license is permissive therefore there is no point in touching that that's the number one decision and this is exactly what I recognized this is also what the team wanted so perfect no conflict and then the idea is to say okay what is it that we build and how do we protect the asset because there's this tension between the business requirements especially as we go to venture capital that expects massive returns so there's a temptation to you know disrupt the core asset if it is to be protected. So the way we did this is to create a for-profit company in France, but we used a recent law called loi pacte, so the 2019 pact law in France that borrows from other nations in Europe, I think, where we have a special statute in the bylaws of company. And so we are a mission-driven company, legally speaking, so for-profit, mission-driven, and the mission is to develop, maintain, at a state-of-the-art machine learning software in open source. So this actually, perhaps you're familiar with B Corps in the US. The B Corp is thinking about the environment. So there is an explicit desire by the company to do good by the the environment in our case we want to do good by the open source and it's by law so we are now responsible for this duality and I think a modern company nowadays should be dual it should be focusing on distributing dividends to its shareholders as a company does but also focusing on non-financial dividends to the rest of society and open source is one way to do this so that's that's the narrative change at a corporate level and then of course the difficult thing is to actually find the product. So finding the product that will be commercial is the most difficult thing in the market of software, where everything is competitive. And you have to please investors, because they expect 10x, 10 times return on their investment. So we decided to focus on what Scikit-Learn did not provide to the enterprise market. And we conceived a new product specifically to address that. In a way, we're a traditional startup. That's what we pitched, and it comes with a cost to maintain, but it comes with a great marketing asset. We cherish the fact that we are the trusted people behind Psyche Turn, which by the way is a governance of its own, so the governance has not changed. We have people inside the company. We started with 14 co-founders, so we're quite atypical, but we have also maintainers outside of probable. The governance has not changed. We just moved the center of gravity from the research lab to the company. But then, of course, and I'm happy to say that we launched a new product. We found the product. We took care of the well. That was part of the business modeling. We knew it was going to be costly, so we raised 18 million euros between pre-seed and seed, which is probably the biggest round of seed funding in Europe for strictly open source, but now we are VC compatible as a machine driven company. So it's a very unique model that came with a lot of willingness.
Speaker 1 [21:01]
Yeah, I mean, for us, it was actually kind of similar. Like, when we started out, we actually did what we call, we raised the client round. So we did some consulting because there was a lot of demand and also to check, okay, is what we think we should build actually the right thing? And for us, it was also immediately clear people needed to work with data. I think that's where machine learning is kind of unique. It's like software 2.0. You have code and data, and a lot of the value is in the data. So we can provide the code basically for free and open source and then focus our product offering on the other part that's more custom.
Speaker 2 [21:34]
Yeah, do good things and good things will follow, right?
Speaker 1 [21:36]
follow. Yeah, so we basically thought, okay, an annotation tool makes a lot of sense and also I think when transitioning into more of a startup what was always very important to us is even in a VC context to not go down that route where you bet your entire open source work on the VC outcome because I think that's how we've seen a lot of companies die and a lot of the classical what's called unreliable startup you adopt a project and then the startup goes behind the project goes and raises a lot of money and then, I don't know, they pivot here, they pivot there and in the end the open source is kind of collateral and
Speaker 2 [22:13]
You know. Yeah, man, here we go.
Speaker 1 [22:15]
That happens a lot. I think we've all had open source libraries we're using and they just sort of disappeared when the startup went bust. And we're like, that's definitely not something we want to do. And also for us, we want the open source to be kind of...
Speaker 2 [22:28]
kind of you still feel very responsible for yeah
Speaker 1 [22:30]
Yeah, and what we raised money for was like a product, and that's kind of the other arm of things, and they are definitely things that need upfront capital. That's kind of normal in business, but we definitely see the open source as special.
Speaker 2 [22:43]
special how do you maintain trust and credibility in the community like sometimes people like or is is it is it just just a misperception also maybe in germany the background or many people if you talk to people who are not in in open source they very have very often have like a perception of open source this is like a single person somewhere very idealistic driven and we We had like the new movement earlier in Germany. Germany was also like a thought leader back in the days on open source or freely available software as well. But it was always like, yeah, you can use it, but you must not make money on that. And that has changed eventually. And the mindset changed. So I don't want to say this was a legit claim as well. I mean, people spent free time on creating it. And of course, you are the creator and you control what you do. And then we have the shift to MIT license, I think, was the biggest shift then, driven from the States. So is there even like a trust issue to say, yeah, I'm doing open source and business? Or is it just like times have changed and you see that there is no contradiction?
Speaker 1 [23:57]
I think you do have to have like integrity and that's something you kind of can't fake and I think like working with the developer Community means like developers have a very good bullshit detector And so, you know You can't like go into open source and act like a typical like corporate asshole like people people with the suit, right?
Speaker 4 [24:12]
the suits, right?
Speaker 1 [24:14]
It's not about that, but, you know.
Speaker 2 [24:16]
I love that, that's really good, yes, protector, I love that.
Speaker 1 [24:17]
Bullshit. People can see through that, and I think that's also, you know, in the startup context, again, why a lot of...
Speaker 2 [24:26]
Basically, I could rephrase it. You have basically your social network or like the community network, the people who develop with you are also like your smoke detectors, if something goes wrong.
Speaker 1 [24:37]
Yeah, yeah, and you kind of always more like you have to be legit like, you know There's a lot of other areas of business. You don't have that you can you know, you can put up a nice corporate front But I think open source similar
Speaker 2 [24:48]
Do you have a similar experience like smoke detectors at Jupiter, co-developers?
Speaker 3 [24:55]
Wester, your first question earlier, actually, why open source? Yeah, no, both, both.
Speaker 2 [24:57]
I know both, both, both, just like that.
Speaker 3 [24:58]
In a way, I don't think we have a choice, especially if we were doing, I don't know, a content management system or an operating system or an accounting software, we could do something closed source, probably, right? That would be accepted, that would be honorable, people would probably buy it if it was good. If you're doing anything scientific, then that doesn't work. and the deep reason is that anyone who's done science at some point in their life understand that there would be a huge contradiction in trying to understand the world like physics or biology or anything with a tool that you don't have the right to understand or look into. So we can't deal with black boxes in science. So even more than any other area. And so in a way, if you want to sell software, any sort of SaaS model or whatever that is meant for that public, the core of the logic and of the intelligence will have to be open source. Otherwise, nobody is going to buy it. And that's a big part of the trust that people have in the software is that it's been audited by experts. And actually, there is a lot of the AI, you know, sort of resistance that we're seeing is that people feel that it's taking that away from them as well. Like, because it's becoming, in some ways, sometimes a black box again, right? And so in terms of trust within the community, we've had conversations at times where people were worried that we were pushing an agenda into a project that is openly governed and in which we have weight. And so there are always conversations and sometimes our opinions were met with suspicion.
Speaker 2 [26:58]
Maybe, what was the reason, I think sometimes the experience is maybe people just don't, maybe lack of transparency, not also unintentional, because like everyone's busy, or maybe people just having different views on perspectives, or being overcritical as well, I mean not every criticism is...
Speaker 3 [27:18]
Well, just being a company in your software stack that is very much developed by people in academia, there will be some, you know, scrutiny into what we do in that stack. And I think it's expected and it's probably a good thing, sort of keeps us honest. So I like it.
Speaker 2 [27:36]
I think the interesting thing here,
Speaker 3 [27:37]
I think so.
Speaker 2 [27:39]
all your free libraries basically originate in a way from academia. So it's a transfer from academia into business application.
Speaker 4 [27:48]
And we should also say that, and to complete what you said, I think managing expectation is an effort as we create a corporate around the underlying asset is an extra effort. Because if you have a leader that says, okay, we take the open source for granted and we do business, then sooner or later you're going to upset a wider group of people who have expectations. so we need to manage expectations be very transparent, very explicit we're going to do this this way and by the way, governance doesn't change or governance has to change but we're going to tell you how and this is a process it's actually a lot of work but you cannot not do that as a corporation around open source
Speaker 1 [28:34]
Especially with open source, you do start out with a lot of goodwill from the community. You've provided a lot of stuff for free and people want to vote for you. So there is a good foundation there as a company. You just can't go and fuck it up.
Speaker 2 [28:47]
Fuck it up. It was a driver. Like, Spacey always had the most excellent documentation ever. Like, was it a driver? So you basically also say, okay, this is, I make people's life easier. Yeah, I mean, it's like... Taking this step because, like, Spacey was almost, like, insanely good.
Speaker 1 [29:02]
I mean, I've always liked documenting things and also I think especially in the beginning it was like we did have like this good matching skill set that I was doing, you know, all the docs and I was doing the front end and that wasn't really, you know, in a very academic environment that wasn't really the norm and we were kind of able to do that with like our skills together. So it's also, you know, it's fun. It was something I genuinely enjoyed. And I don't know, it's kind of funny nowadays like we, because we put so much work into like all the writing and all the boring backwards compatibility, not breaking people's APIs, consistencies. This really paid off now with agents and Gen AI and coding assistance, which is something we did not anticipate.
Speaker 4 [29:43]
This is a topic for another panel altogether. To me, the product is the documentation when you talk about tooling. That's it.
Speaker 2 [29:50]
Have you ever had the feeling like, okay, we build space, we do something new, but we also have migration problems, and we see backwards compatibility. It's increasing complexity in a fast-moving time. Yeah, actually, we're just like head AI agents. This is like leading us to the next question. So actually, like, AI landscape is consolidating around closed systems currently, and you can have, like, software. They can write software. probably they can recreate parts of software you created just smaller bits do you think this has any impact or mean for your businesses or is it just like you have like of course majority documentation, community trademarks, following expertise I mean like this is like you feel still like AI agents coding For smaller libraries, I don't think an agent can just recreate something like Scikit. There's so much experts and optimization there. But for smaller libraries, people argue, why use the smaller library? The agent can just write that part for me.
Speaker 1 [31:05]
I mean, it depends. I'm sure it's a very different environment and very different now to start an open source library. But I think for us, it's interesting because it also unlocked this whole new user group who are not even actively seeking out the library. And probably Scikit-learn has something similar or even a Jupyter ecosystem where you could now type into an agent like, hey, I want to extract company names from my news reports or I want to do this type of prediction. And you don't even have to say, use Scikit-learn or write this in spaCy. And then the agent will pick the library. And so you have, like, this whole user base that kind of comes into the field like that, which is, like, very different.
Speaker 2 [31:40]
We could also argue you do community blog posts, a few put online, makes the agent work. Oh, yeah, this is the way to go.
Speaker 1 [31:47]
the way to go. Yeah, and I think also, especially with the backwards compatibility and stability, it's just my theory that a lot of the post-training corpora that the coding agents and the companies have with code examples, a lot of this basic code is very likely to work because it's so backwards compatible, so a lot less of it is kicked out, whereas if you have a library that changes its API for every release, the corpora end up with a lot of code examples that do not in fact work, and then this code gets excluded, so it's at least my theory, I don't know, from just what I've heard from people working on coding models.
Speaker 4 [32:18]
coding models. But I'll confirm because I mentioned the 200 million downloads per month. Clearly there are not 200 million data scientists on the planet. There's only 3 or 4, maybe 10 million. So these are your numbers. So a lot of the downloads are CICD triggered by cron jobs or maybe lambda functions with a ripple effect. But what we're seeing recently is mostly agents. Most of the users are agents. Precisely.
Speaker 1 [32:47]
Check, like, the stats again.
Speaker 4 [32:49]
It's crazy. I mean, it's impossible otherwise.
Speaker 1 [32:49]
It's crazy.
Speaker 4 [32:52]
And so what's happening, I think, is that libraries that are established, especially when the licensing is quite permissive, are going to be polarized. In other words, they will pick that library because an NLM, an agent, is an energy function. It will use the least amount of energy to get to the conclusion. So therefore, if there's been self-reinforcement across the training corpus, and that is the case because of the documentation and all that. Therefore, it polarizes the usage of core libraries. There is no point to recreate a library that is fully open source because that actually will fight against a library that is defined by name as the reference. So I think this will get even, not worse, actually better for established libraries. It will be difficult for incumbents, but it's fascinating to see that agents are now picking up the load.
Speaker 2 [33:53]
Yeah, it could even like, it could even end up that what's already in the corpus basically manifests. There's a bias now to use scikit and even like, I'm just like open question I have in my mind.
Speaker 1 [34:04]
But I think there's still a lot of potential for...
Speaker 2 [34:06]
Would polos happen again because pandas would be like super present now, right?
Speaker 1 [34:10]
I think there's still a lot of potential for innovation because that is, you know, the things that also, you know, the models can't do. Like, we need, you know, if you're a library developer, you're thinking about, okay, what is the API that people should be using in the future? And what is something that's new that, like, a coding model does not produce? And then also, I think it's still, you know, contrary to what some people think nowadays, how do we even need software anymore? Do we even need to care about library APIs? I actually think having programmable software and programmable tools is more important now than ever because of the coding assistance. And I think that's also why there's so much value in thinking of future APIs and future ways that both humans and now agents as well can use.
Speaker 3 [34:52]
There's probably going to be a market of trust, as well, in that, as you said earlier, there are reference implementations in core libraries and that all of these agents rely upon. And the expectation is that these won't break. And so people who are responsible for making these things not break will be even more in a position of being the references in these areas. And open source especially, because there is so much corpus out there that these models are trained more on open source software than they are on things that they can't see. So in terms of whether this impacts our business positively or negatively, at the moment, I can't say. And I won't make any predictions, because we live in science fiction now in many ways. But what we're seeing right now is that there is a high demand for, for example, we are seeing a lot of contributions in core Jupyter projects from people basically using agents. And a lot of AI flop coming. So we removed the easy fix tags on many repositories so that we don't get these anymore. So the very things we had done and bought new contributors, we are removing now because of this.
Speaker 2 [36:25]
Yeah, also having a structure like a company can also be protect against AI slot now. So before we use through European perspective, just a quick question and one quick answer. Do you change anything in APIs, how you build your software or product because of agents and different processes?
Speaker 3 [36:47]
Oh, yes. Yeah. Yes. Yeah. So we've been working on a tool called JupyterAI and a pendant for JupyterLite called JupyterLiteAI, which allow using the Jupyter user interface through what we call commands. So most of the actions you can do in the Jupyter user interface are backed by something called commands, which are basically APIs, but not operating on live objects. They are serializable APIs. They're very amenable to tool calling. So we directly made them discoverable by LLMs. And it was performing very badly at the beginning. And the thing that fixed it was improving the documentation, basically changing the doc strings and making them very comprehensive and very explicit. And then Jupyter AI was working very well. So yeah.
Speaker 2 [37:41]
Yeah.
Speaker 3 [37:41]
That's what's
Speaker 2 [37:42]
Oh, it's basically.
Speaker 1 [37:42]
Yeah, yeah, so I think definitely the stability like, you know, we're much would be much more hesitant to introduce like big breaking Changes that previously, you know, we could have maybe done with the community and then I think more generally like we've definitely seen Okay, what are the parts of the library that like the agents are good at? What are parts that they're less good at? How can we improve that and of course in our case we are so we're currently working with Coding agent or we're working on our newest product. That is essentially, you know, kind of I call it vibe NLP So it's basically imagine cloud code but for NLP projects and as part of that we also developed a bunch of skills So we knew okay, the agents are a bit worse at like writing code for commercial tools Just because there's less code available freely on the internet So let's write some skills for that write some skills for certain best practices So we have been like I feel like you know, definitely targeting our new product even more specifically around those kinds of workflows
Speaker 4 [38:38]
And the way we are dealing with this is to actually flip the equation based, again, on scikit-learn, which is a platform in a way already. So scikit-learn has a lot of libraries that gravitate around scikit that are compatible with its API. But all of it is kind of dumped in a fixed file or on GitHub in some way. So it's kind of organic. So we decided to reorganize around something that we call scikit-learn central. So it's a website that is a sub-website managed by Probable. And so we start to identify the libraries that really, really comply with the scikit-learn way. And we now have an MCP server so that we can feed people who want to do real machine learning around the scikit-learn environment. And they get access dynamically to use cases and combinations of libraries. Because scikit-learn, in a way, is a very mature, very massive object that needs to work everywhere. Version 1.8 is going to be curated for six months to a year, because we need to make sure that the 1.3 million projects on GitHub are still happy with it. So we have a responsibility. So scikit-learn doesn't move fast by design. But we are building new libraries around it, including one that is the commercial product called Score. So it comes with a library and a SAS solution. But these things need to be somehow well-presented to AI agents. So we orchestrated it around this Sakit Learn Central community.
Speaker 2 [40:15]
And now we center about Europe. So anything specifically European about the way we would think about open source as a business, sustainability, sovereignty, independence, you have also written a book about it.
Speaker 4 [40:38]
So I've written a book about, and this was my first statement, which is, as Europeans, we are so far behind in terms of geo-economics that we have to lean on something that allows us to catch up. And China actually made open source their core strategy since 2022. So as Europeans, we probably need to help our governments understand the value of open source. But it cannot be dogmatic. It has to be bound to our ability as Europeans to recreate value capture. So it sounds counterintuitive, right? You distribute to recreate value capture because the numbers are terrible. I'm sorry to say. I hope you had coffee. Europe ships 280 billions of euros to the US per year. One third is cloud. One quarter is SaaS. one-fifth is hardware, but once you buy the hardware you have it, but the point is most of the value is as a service, software as servitude, if you wish so it means that the money is going to the US and it's not a criticism of them they're doing well, their products are great, the thing is we are not creating the jobs we're not creating the taxation we're not managing to recreate the bigger companies, you're lucky in Germany you have SAP. That's probably the only big company in software in Europe. Terrible. Only one of that size. So we have to fix this. And open source is a way to just, again, reset. So this is the topic of my book, which you might help translating into.
Speaker 2 [42:19]
I will catch up after the conference.
Speaker 4 [42:22]
He's been busy. But the book will be called, if you allow me, Ouvertarismus, German with a French root. But anyways, so to me, that is a change of paradigm, a change of narrative where open source is here to stay. Open source always wins. And open source is part of the fabric, including commercially. We need to find a way to do this tactically, not dogmatically. So, everyone needs to be able to work with these instruments at the macro level. But again, it's counterintuitive. When you do an open source, as a European, you can be perfectly proud to say, my code is good for the world, but we need to create a business model that actually somehow captures new value that allows us to create jobs and create great companies that are powerful and and are meaningfully participating also in maintaining all of that, right? Because the cost of maintenance, you know it. It's extremely costly to maintain open source. So we need a driver of value capture against it. So it's this duality again that we need to fix. It's really hard, but it's super exciting as well, because I think it fits as well as Europeans. I mean, this is the core of our values, right? more openness and less dogma, and dogma is being imposed on us by closed platforms. So let's go back to the Enlightenment.
Speaker 1 [43:53]
Yeah, I also think that just viewing open source software as like critical infrastructure what it is I think is important and there is a lot of There are a lot of libraries and things that are very developed by small teams that like a lot is built on and including You know a lot of public infrastructure
Speaker 2 [44:08]
I mean, also, like, open source is not just plug and play, and you have it and everything. You need the experts. I mean, why should people call Sylvain if it was just, like, yeah, you installed it in Magic, and you don't need people to help with Jupyter or customize or Ecosys. Same for you, Spacey. I imagine, although it's not one of your course, I can imagine, like, every ten minutes somebody calls you, can you consult and help us in Spacey, because you're the experts.
Speaker 1 [44:35]
Which is now why we are automating ourselves out of it with a new...
Speaker 2 [44:35]
And this is, I think... of it with a new product oh there's a spot coming i didn't know that enos and matthew bought
Speaker 1 [44:45]
I assisted an NLP engineer that is sort of, you know, automating ourselves out of the...
Speaker 2 [44:49]
I think the value you refer to, we can capture. I mean, you still need advice from the creators or like people close and know how to use because it's not just, I mean, it's not just like an app and you say, click here, click here, and you buy something. It's just like way more complexity and there's a lot of like domain knowledge. So it's design.
Speaker 1 [45:07]
Also, it's designed to program with it. Like, it's not necessarily a product in that sense. It's like, oh, that's also why it's popular. Like, people use open source not because it's free. I think that's one of these misconceptions. But because it's flexible, it's extensible, you can program with it. Composable. Yeah, composable and all of these things. That means necessarily that people do compose things with it. So it is a component of a larger system.
Speaker 3 [45:33]
Regarding your questions earlier, I think the time when Iran and Venezuela were the only potential targets of economic sanctions is over, right? And this enormous power that the U.S. has over our digital infrastructure could really harm us. We are one tweet away, as we heard this morning, of the president of the U.S. saying, oh, Airbus is doing unfair competition against Boeing, so we're going to shut down, you know, AWS and Azure to them, or a company, or even maybe an entire country. Even though this may hurt them as well, they could totally do this for companies. And so the way we've been dealing with this because first we're a European company in that even though we are incorporated in France we have people in for the countries of the Union and it's that we we've started engaging into some kind of almost economic survivalism and that we have migrated all of our services to European providers and I think open source if people start thinking in these terms that we moved email our hosting of the website. We even have a contingency forge with pretty much everything we touch on GitHub mirrored so that we have something to do the next day if we are shut down, basically. I think open source will play a big role if people start thinking in these terms, in terms of what do we do the next day? How do we start rebuilding? Do we even own our data? Would our data even be accessible Should we be shut down from the internet, basically? And so I think Yann said once that the open source was the tool of the underdog, right?
Speaker 4 [47:32]
Right? Is this the term?
Speaker 3 [47:33]
The challenger.
Speaker 4 [47:33]
The challenger.
Speaker 3 [47:34]
The challenger, right. Sorry, I was looking for the term. I think it is. And so we have to use these tools as the challengers now.
Speaker 4 [47:44]
And actually to counter that, when you have too big of a company with proprietary software, it's the bully. The underdog needs resilience because the bully will bully at some point, right? So if we get into a situation where there's a crisis and we are in a difficult situation right now globally, the world is different nowadays than it was 10 years ago, then open source is certainly a component of resiliency.
Speaker 3 [48:13]
so and so Europe in this I think we are clearly in it together and So we need to start thinking of this economic survivalism together because even the biggest countries in Europe are rather small and the global scale and So we don't really have a choice. So ask yourself What if Azure? pipelines or just you know, even teams was shut down like or the the any north american collaboration suite that you use in your company what would you do the next day could you even operate anyhow and if it's not the case maybe there is something you can do about it start doing and we've done this at quantstack through baby steps migrating services by services to alternative providers so that we can continue operating without it actually
Speaker 2 [49:04]
Actually, Pioneers has a lot of experience being teams shut down because we had a non-profit license and they just stopped the program, which is not a problem. You don't complain if you said something for free, but it was not even like a polite mail, we're going to shut this down. I had to basically pick this up from the newspaper. Just a side note, we have less than one minute left, I think. So advice, short advice. If you tell somebody who wants to build a business on open source or switch to doing it and hasn't started, what would be the advice for this person you want to give them?
Speaker 3 [49:47]
I think the path of least resistance is to pick a very used project that is lacking attention and love and be useful in it and start building legitimacy in it and then you will be the legitimate person to help people who need that project to continue operating in the future. That's a really good way to get business.
Speaker 2 [50:10]
reinvent the wheel, rather build on shoulders of giants from others? Like, rather build on something that is already there and not try to Yeah, it's really
Speaker 3 [50:21]
Really, adoption is a tricky thing. And thinking, I'm going to start a completely new project, make it open source, and build a product upon it, you need to be really, really, really talented first, but also to be very lucky. There are many components to this equation, while just actually being a follower initially and gaining weight in the community is probably a better way and a more useful way to help.
Speaker 1 [50:48]
Yeah, especially nowadays, I think, because also, yeah, it's hard to give advice, because the times when we kind of started out, it was very different, or it was very different, you know, to provide value or get attention. And so I think the landscape has changed a bit. But I do think it's important to have a clear idea of what the business should be, because I think it's nice to be idealistic. But ultimately, open source is a very thankless job. And you can't expect like, oh, you're going to invest all of your free time for no pay. And then suddenly, you're going to make it big. Like, that's not how it works. and I think a classic solid business and also I do think maybe it ties back into the European question but I do think it'd be nice to have more critical infrastructure funding and options that go beyond just VC funding because I do think for purely open source businesses I do think VC funding is very unideal and it leads to the type of thing what we discussed previously like the unreliable startup or just weird corporate bullshittery weaselly stuff that developers don't like And so I think having a different mode of funding would be very useful.
Speaker 4 [51:51]
And to that point, but you've both experienced this difficulty, is that if you choose the scalable model in terms of business, therefore maybe leveraged with VC funding, then you cannot start with consulting. In other words, consulting is antinomic to super scalability expected by VCs. And I find that it's really, really hard. I have a lot of people in my network who've created a consulting entity on proprietary software or open source, doesn't matter. It's really, really hard to actually switch and vice versa.
Speaker 1 [52:28]
It's also, I mean, something that just VCs don't like. I mean, it's like you can totally, you can do consulting as we see it on a startup. It's just like all the VCs tell you, ooh, but like the playbook says here, you do not do consulting.
Speaker 4 [52:36]
It's actually the valuation. The valuation of a consulting firm is one time the revenue. Valuation of a SaaS company is 10 times the revenue. That's it. They are looking for 10 times the revenue in valuation. Therefore, consulting is not what they like.
Speaker 1 [52:50]
But I think also in general, just don't bet like your open source project on like, you know, don't put yourself in a position where you're dependent on someone else and have to like, you know, fuck over your entire open source community just, you know, to save your ass or to save your company. Like it's not.
Speaker 4 [53:03]
We're fine with it because we built a new product.
Speaker 1 [53:05]
Yeah, no, and I think also that it's something we were very conscious of and I think what you see that in most companies So yeah, so I think we're over time. Thanks very much
Speaker 2 [53:13]
Thanks very much.
Speaker 3 [53:13]
Thank you very much. Thank you.
Speaker 2 [53:15]
Let's move to, no, we still have Q&A from?
Speaker 3 [53:15]
Let's move to-
Speaker 5 [53:17]
yeah thank you very much for that amazing panel we have a bunch of questions so
Speaker 2 [53:17]
Yes.
Speaker 5 [53:21]
please vote the questions you feel interested so i will start with the first question so how can you build an open source business around a around a bit of software that is used by very few businesses, okay, it moved, I'm sorry, ah, here, very few businesses where the demand of consulting service is likely to be low, so who want to answer this?
Speaker 1 [53:56]
Well, I mean I think as a to start like said previously I think consulting isn't even necessarily the best business model because again you are kind of limited And you have that problem of like oh how if if you make your software too easy to use the demand kind of goes down Depending on the industry, so I do think it depends like I have definitely seen companies and projects Do really well even in very niche markets where they add like a lot of value like even if it's I don't know something really specific to a certain type of medical imaging where like in that field people definitely need that and it's really niche and you have someone who has the expertise there I think that can be really successful because like just if you if there is a lot of demand there's potentially a lot of money even if it's you know not a broad project that like the average person uses like if you have like a few good customers that are like really good and have a really high demand I think you have you can potentially have a business.
Speaker 4 [54:50]
Yeah, possibly. I mean the main question in that situation to me is about lifestyle Do you want to work as on this as a passion project where you're driven intrinsically by your? personal drive with regards to the topic in which case consulting alone or Keeping it as a side project in a weekend if you can do that It depends if you have kids and a spouse and all that so there's a huge question about lifestyle if however you feel that there is a huge opportunity it's all about ambition speed velocity then you probably need to raise capital pitch the idea and say this is gonna be big but then you're no longer doing it alone you're doing it with VC partners who will be taking part of the company and that's a choice it's a it's a legitimate choice but it's a different story altogether
Speaker 5 [55:44]
Thank you. And what is your thoughts on OpenAI's acquisition of Astral, e.g. big corporate buying up a small open source focused company?
Speaker 1 [55:55]
I mean, that's sort of part of the corporate bullshittery stuff that I was talking about earlier. At least that is the risk that you end up with or if you put yourself in a position where you're also, for example, dependent, very dependent on future funding or can't just, you know, we were very lucky we were able to make the switch and say, okay, we go back to being fully self-funded. Depending on what position in a VC-funded ecosystem you've set yourself up for, you cannot do this. So you are entirely or you're dependent on, you know, the investor following you to the next round and if they don't then you can shut down and you lose everything and you close down your open source project or then you might end up in a position where you have to sell or maybe you want to sell and you're like you know it's a great offer so I think it you know it depends but I do think there is you know
Speaker 2 [56:40]
The sprawl was packed from the very beginning.
Speaker 1 [56:42]
Yeah, so I think if that's the route...
Speaker 2 [56:45]
We're also open about it. We have a VC back.
Speaker 1 [56:46]
Yeah, and if that's your goal, or if that's the classic exit, people can do that. But I do think, again, that's where we go to the sort of stuff. If people treat the open source project as sort of really just a vehicle that can leave a very bitter taste in people's mouths, that can really hurt the credibility in general. And then I'm also... I don't know if open source really is the right route for things. It's not, you know, yes, it's a marketing tool, but, like, you know, you should be providing value.
Speaker 3 [57:18]
For Astral, I think it's really an amazing illustration of something we said earlier about building something that is boringly useful and utilitary. They built this, you know, LinkedIn tool and a package manager, right? If you go to a startup party and people ask you what you do and you say you're in package management, people don't really get excited about it. And because package management is really what gets into in the way, right? You don't want to be doing this, but you kind of have to, right? And it's like living in a city. You want a sewer system, but you don't want to be the one building it, right? And so they went ahead and built a great product to do something that everybody needed to do and nobody wanted to, right? By the way, there is another company here in Germany that does something like this, very similar to Astral. That's PrefixDev, right? In PrefixDev, they built Pixe, which is amazingly fast, just like Astral, but it's just like UV. but also applies to other non-Python ecosystems. But then in the case of this acquisition, there is obviously the risk of concentration of power into a few big players. And it's like the same playbook that we've seen so many times. And this is worrying, obviously.
Speaker 1 [58:32]
So I think Fundafact like Prefix, as far as I know, they are funded by the sovereign tech fund, among other things, who do fund critical infrastructure in Germany. So I think that is kind of a successful example of how this could work.
Speaker 4 [58:44]
What we hope for is that European companies start acquiring these companies as well because we need that critical mass at some point. So they hopefully manage expectations and do things right when they buy those companies. But hopefully at some point some European company steps in, steps up and starts consolidating as well.
Speaker 1 [59:03]
Maybe they can start by just publishing some software open source themselves first.
Speaker 3 [59:09]
Okay. It's really building something that's meant to be used.
Speaker 2 [59:12]
Startups and investors panel, we really lack funding beyond series A or B. If you really want to go big, there's just not these investments available to startups.
Speaker 4 [59:21]
We have that now in Europe. It's not the problem. The problem is the exit. It's really the exit. Typically, we don't have that consolidation. We don't have an IPO market. So that is what is hurting us the most because we create amazing companies, either open source or proprietary. Europe is a prime producer of software no doubt but more often than not and I've had this experience myself at some point you the buyer the US buyer comes in and say okay I'm writing the check and there's nobody else in Europe that actually writes that sort of check so again the VC route is tricky at the end of the road so we need to fix that but that's a European market scale problem
Speaker 5 [60:08]
Thank you very much. I would say we have one last question for you. This question is for Jan, but feel free if you have anything to add for it. And it's really hard to answer this short, but please. Yeah. You mentioned the importance of the legal structures of your company that forced you to focus on your open source mission. Open AI used to make similar claims, but then they changed their minds and the legal structure did little to prevent that. Is lethal structure really useful to enforce future behavior?
Speaker 4 [60:42]
It's no guarantee, but OpenAI set itself up for a difficult path. They started up as a non-profit foundation that required a huge amount of capital. What can go right? This is a foundation with Elon Musk money that was used to buy a lot of GPUs. So at some point, people want their money back.
Speaker 1 [61:10]
That's it so led by some Altman out of all people yet
Speaker 4 [61:12]
people, yeah. Out of all people. I don't know him, but yes, I think I wrote an article, actually, if you hit me on LinkedIn, maybe I'll send you the link, but I wrote an article about the governance issue of OpenAI with two faculty professors at INSEAD Business School. So we actually unpacked the governance issue. So there is no guarantee, but we found a model that seems pretty resilient in managing expectation, in managing the transparency and in caring for something which, by the way, I'm happy because for scikit-learn it's too late, it's free.
Speaker 5 [61:48]
Thank you very much. There are much more questions, I know, but we're over time. So thank you.
Speaker 2 [61:52]
Thank you. Thank you.
Speaker 3 [61:54]
Thank you.
Speaker 2 [61:54]
Thank you.
Speaker 3 [61:55]
Thank you. Thank you.