Cooking up a ML Platform: Growing pains and lessons learned

What is an ML platform and do you even need one? When should you consider investing in your own ML platform? What challenges can you expect building and maintaining one? Tune in and discover (some) answers to these questions and more! I will share a first-hand account of our ongoing journey towards becoming an ML platform team within Delivery Hero's Logistics department, including how we got here, how we structure our work, and what challenges and tools we are focusing on next.

This session took place in track Sponsor and was classified suitable for novice domain / novice python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:03]

All right. Hello, everyone. My name is Cole, and this is the talk, Cooking Up a Machine Learning Platform, Growing Pains and Lessons Learned. Really excited to be here to talk to you all today. I guess it's a topic that's interesting for many people to define what exactly we mean with machine learning platforms. And that's what I want to share with you today, is more hands-on direct experience working with machine learning platforms in an organization like Delivery Hero. But first, what is in it for you today in this talk? What do I want to cover? Largely speaking, I want to organize the talk around three main questions, which I will do my best and try and answer for you. The first question, quite important, what is a machine learning platform? I think we should all be able to agree on something there before we can go into any more detail. Also maybe dig into the motivation behind why we need one in the first place. And that's kind of the second question as well, is when do you actually need a machine learning platform, or how much machine learning platform do you need? Which also is very important. I think we should always ask ourselves this before we put too much effort in. And finally, how do you actually build one? So really here, I can't give perfect answers. I'm not an absolute expert here, but what I can share are my experiences, my perspectives, the challenges I've faced and the things I've learned along the way as we try to build our own machine learning platform and tools within Delivery Hero. So first of all, maybe a bit of background and context about myself so you know where I'm coming from. My name is Cole, I'm working at Delivery Hero in the logistics department and I'm the manager and tech lead for our machine learning platform team. Our team is still pretty small and scrappy, we're myself, a product manager, and five other great machine learning engineers who are working together to solve these problems and figure out where we're supposed to be going with all of this. In total we're supporting around 25 different data scientists across the logistics department who are spread around five different teams who work on various different kind of domains, cases and models. If you don't know much about Delivery Hero, I can give a quick introduction. We're here in Berlin. We also have a booth at the conference, so you can come by later, ambush me and ask me any questions you might have. But basically, we're a food delivery and grocery delivery app. We operate all around the world. And so in total, within logistics, we have 30-something models running in production across all of our markets around the world in over 50 countries. Generally speaking, our models, I mean, we're working on logistics, so we work with a lot of structured data, classical machine learning, a lot of boosted trees, not so much neural networks and deep learning. And generally I would classify our use cases in one of two groups. One is the kind of batch or offline scheduled prediction models, and then we also have the real-time models that are deployed as, like, API endpoints as well. And just for a sense of scale, one of those low-latency user-facing endpoints and models can serve something like 130 million predictions in a day with P95 latencies of less than ten milliseconds to ensure a good user experience as well. So that's a bit of context, just so you know where I'm coming from. I'm sure many of you work in very different contexts, maybe larger scale, smaller scale, different types of problems. But yeah, this is where I'm coming from at least. So the first question I want to get into, what is a machine learning platform? And I want to just check what ChatGPT has to say about this. And in general, it gives a pretty solid response, as we might expect. But I'd also like to nitpick this a bit. I think this is a very good answer because it kind of represents the common understanding of machine learning platforms. But for someone who wants to actually build and operate one within a company, I think there's some misconceptions here as well. The first thing I want to highlight is this statement here. A typical machine learning platform will offer a range of features such as data preparation, feature engineering, model training, model evaluation, deployment, and monitoring. To me, it kind of sounds like the platform does everything, so why do we even need data scientists, right? Obviously, that's not the reality, but this is maybe what people expect from a platform. You'll do everything for me. The second point, of course, some popular machine learning platforms, all the big cloud providers. Yes, this is true. These are popular machine learning platforms, but often what I find is these aren't enough on their own, right? For most data scientists, you kind of need some glue to fit these things together and to work with the components that you need to solve your problems. And the final part isn't a nitpick. Actually, I think this is a really great definition. We'll see this actually multiple times in other slides. It goes into the motivation of why we need machine learning platforms, and in their words, they can be used by data scientists, developers, and businesses to accelerate the development and deployment of machine learning applications. So to me this is 100% correct. The whole point of a platform, right, is to make developers' lives easier. So some good things, some bad things. So how would I define it better, right? Again, I'm not an expert here, so I think what you should do is just go read up from these authors of team topologies. They have a lot of really great work defining what it means to be a platform team in an organization, how you should structure them, how you should collaborate. And they also go into quite a bit of detail about platforms and platform engineering in general, which I take a lot of inspiration from. So first, what I think Matthew Skelton here describes is kind of what we want to avoid if we're building platforms within a company, right? And they describe these kind of legacy platforms of the past, which were very big, massive, monolithic, hard to use, black box systems, kind of enforced top down and mandatory. And I think all of us have some experience working with this type of platform. And generally, we don't like it, right? We have to do a lot of hacks and workarounds to make it actually work for us. But sometimes this is just what happens, right? And bad platforms are made with good intentions, nothing wrong with this. But if you are a platform team or you're looking to build a platform, I think it's something to be aware of. On the other hand, when we talk about what does a good platform look like, I also really like how they describe things. In essence, a good platform is also a product, right? a compelling internal product, again, to accelerate delivery by stream-aligned or product teams. And I think this really emphasizes the key thing to take away from this talk, which is, if you're building a platform, realize that your platform is a product, and realize that your developers, your data scientists, are your users. And you should treat them with the same respect that you would your paying customers, right? So focus on the developer experience. How do you make things easier and faster? How do you make things so that they're not slowing your developers down, they're actually speeding them up? And there's a nice concept they've also coined called thinnest viable platform. So it's kind of a play on MVP, of course, which I think is also a really good way to think about platforms in general. So the concept of thinnest viable platform is a contrast between a thick platform, which is something kind of monolithic, black box, hard to extend and hard to be flexible with, versus a thin platform which is just enough to solve your problem. And to really drive this to an extreme, in Matthew Skelton's own words, in some cases, maybe all you need for a platform is a wiki page, right? And if that's the case, then write your wiki page and be done with it, right? Thin is viable platform and iterate from there. So I think just bringing all these ideas about good product management, good software development in general, bringing that into your platform teams is really the key that we want to strive for. All right. So that's my rant about platform teams and blah, blah, blah. What about machine learning platforms, right? So I'll give a quick definition, and to be fair, the definition is quite loose. So for the literal answer of what a machine learning platform is, or what any platform is, it's just a collection of shared infrastructure, libraries, tools, documentation, processes, et cetera, et cetera. So I think what I want to emphasize here is it's not just tech, actually. It's the onboarding, the user experience, talking with your users is also important, with the overall goal to enable and accelerate development, deployment, operation of machine learning use cases at scale. So this combines some of the things ChatGPT said, also with team topologies in general, and from my point of view is a really good way to look at what we're trying to do. And to give a bit more context of what we are working on within my team at Logistics and Delivery Hero, we roughly divide up our machine learning platform into three main products, visualized here at a very abstract level. So on the top you have kind of the core of most machine learning platforms, which is your model orchestration, development, and training tools. That's where you need to run your training jobs at scale, construct your DAGs, et cetera. Once you have a trained model, you need to actually release that somehow. That's the second kind of product on the bottom right, which is all about how do we deploy and serve the models at scale? How do we also monitor them? How do we A-B test and experiment them, evaluate how good they are? And finally, the third product line is about feature store, and in particular we're focusing on what we call a live feature store, so essentially capturing data in real time from the rest of our back-end systems and providing tools to help data scientists aggregate and serve that in their models to provide better quality data for better quality predictions. So that's just a little bit of insight, but that's not really the core of my talk. So if you have more questions about this, feel free to find me later. And I want to double down a bit on the motivation behind this, right? So why do we actually need a separate platform team? Why do we actually need to worry so much about platforms? we already have DevOps platforms and stuff in our companies, isn't that enough? And to answer this, I would also refer to this really nice Google research article. I think the title kind of says it all. If you haven't read it, look into it. It's not too long and has a lot of great information. It's from 2014, but I find it's still extremely relevant today. So yeah, the title is Machine Learning, the High-Interest Credit Card of Technical Debt. And I think for anyone who's worked on machine learning or data science at production scale, I hope you agree with the statement, technical debt is just hard to avoid, almost impossible to mitigate, and it just can escalate out of control, right? And so this is a bit of why we think about platforms, right? We want to handle this. And the article talks about a lot of reasons why this is true. A lot of it, you know, talking about data and a lot of other things, but one aspect that I want to zero in on is just the tooling landscape. So if I look around at what tools are out there to do machine learning, it's ridiculous, right? I mean, it's almost comical, and the cognitive load that the average data scientist has to deal with is enormously high, and this is what results in really things dragging out and slowing down in a lot of organizations. So if we think about these three main streams that I focused on before for our machine learning platform, within each of these, you have a choice to make, right? You have all of these great tools out there, some open source are managed, you can build or you can buy, and you need experts to decide which of these tools actually makes sense for us, you need to learn about... Oh, one second. All right, sorry about that. So tech debt, tooling landscape, cognitive load, right? So if you think about all these tools, you have to make a decision which tool are you going to invest in, right? Also you need to deploy and operate this somehow, and sometimes, you know, data scientists can handle one or two of these individually, it's fine, but when things go wrong, who's responsible for it, and do you really have the expertise that you would like to work with these tools at scale? And this is like a curation of tools that I personally have either used or researched quite heavily to see if they make sense for us. The reality is much more grim. This is a graphic I just found online, and it's kind of comical, right? like, okay, if you want to break into machine learning, what do you do? Learn all of this. Good luck. So if you know anyone who can do all this, please reach out. I think we'd love to hire them. But, yeah, this is why we need platforms, right? We want to make this easier and we want to make sure that data scientists can deliver value to the business in the end. So when do you actually need a machine learning platform? Given this kind of loose, thinnest viable platform definition, I'll cheat here and say you probably needed one yesterday. But I would also say this is kind of the wrong question, because if a Wiki page is a sufficient platform for you, then you already have a platform. So rather I would ask, when should you invest in a better platform, right? When should you dedicate more time and resources to actually figure out how to build a good platform? And this has, of course, a lot more nuance to it, and I think it depends on the context and your judgment and your organization's. But from my perspective, the big reasons are one of these. So whenever you recognize that you're having problems in your organization around onboarding new data scientists, maintaining your production flows and projects and systems, keeping things online, firefighting, dealing with incidents, doing things at scale over time, right? If you continuously face problems with this in multiple different projects and teams, then it's probably a good sign that you could do with having a better platform, right? And so if you get this far and you're saying, yeah, I need a better platform, how should you actually go about it? I think this is really the hardest part to answer, and I can't really give you all the answers but I can share at least my experiences and what I've learned along the way. And I want to really show some real hands-on journey that we've gone through at Delivery Hero, but before I do that, I can give my kind of snappy platitudes first. First, start small and iterate, you know, like any good product. You shouldn't lock yourself in a room and build something over a year and then expect it to work in the end. It usually doesn't pan out so well. Second thing, listen to your data scientists, which implicitly means talk to your data scientists. I've heard too many stories about platforms that were developed in a vacuum over the course of one or more years and then kind of flopped, right? And it's not a surprise. I mean, if you build something and you think you know exactly what people need but you didn't actually ask them and you didn't actually prototype with them, you didn't actually collaborate with them, then you can't really be surprised when it doesn't work exactly as you expected it. And in summary, treat it like a product. It is a product. It's internal to your company, maybe. Your users are not paying you dollars, but they're paying you in their time. And your data scientists will be a lot happier if you treat them with the same respect that we do with all our normal customers. So yeah, moving on from here, I wanted share a bit about how this has happened for me. My previous company also at Delivery Hero and my colleagues at Delivery Hero. I don't want to hide anything here. I want to be very honest, share with you our mistakes as well, what we've learned along the way and the challenges we've faced. And to do that, I want to zero in on this model orchestration as a specific component of a machine learning platform and how that developed over the years at Delivery Hero. And for many of this, I wasn't there, so this is kind of secondhand information from my colleagues, but it's pretty accurate as far as I can tell. At the beginning, you're kind of like, I think this is pretty common, I've had this experience in multiple places, you're kind of like a start-up within a larger company, right? This is how a lot of data science comes to be. You've hired the first two data scientists or ML engineers, and you just kind of say, hey, we heard data's good, do something, right? And at this point, obviously, your tech debt, you actually have no tech, so you have no tech debt. You also have no no platform to speak of, and so at this point, you just have to kind of put your head down and figure out how to get things done, right? And at this point, it's definitely too early to think about ML platforms. Like you have two headcount, they should be spending all of their time to provide value to the business, right? I think that's very clear. And you know, we're smart people, we're data professionals, engineers and data scientists, et cetera. Eventually we manage to do it, right? And eventually we also figure out, oh, we have to talk to the business, we have to convince people to use our model, integrate with other systems, et cetera, et cetera, et cetera, and eventually, maybe a few iterations down the line, you do it, right? You're able to provide value to the business, the business is making more money, everyone's happy. And at this stage, it's quite interesting because based on my definition of ML platform, there is kind of the seed of a platform even at this stage, and this is what it looked like in the first maybe couple years at Delivery Hero. It's a little bit embarrassing, but essentially our entire compute infrastructure was the two personal laptops of our first two data scientists who joined. One of them would set an alarm every day to wake up and press a button to, like, run the forecast that day and publish it to production. So you can imagine, you know, if you spill a beer on your laptop, then production's down for a week until you get a new one. Not the best platform, but it is kind of what was working and what was needed to get to that first stage. Right? There's nothing wrong with that. But we'll see. Eventually that reaches its limit. Also, in terms of documentation, basically none. You probably have, like, a notepad somewhere in your computer where you've just copied these lengthy commands to make sure you can actually run things when things break. And this is what it really looked like for us at Delivery Hero for some time. But, of course, once you get past this first hurdle and you've kind of deployed your first data science model into production, you've provided some business value, then things start to accelerate very naturally, right? And that's where things start to accelerate. You start hiring new people. You start exploring new use cases on the side. You also start to iterate on your existing use cases, probably scale it up quite a bit in terms of the data features and complexity of your models. And as you do that, this high-interest credit card of technical debt that you've been spending so freely in the early stages starts to kind of surface its head a little bit. And you discover that, or at least we discovered that yeah, this old way of doing things, things need to change. But that being said, at this stage, there really is no true platform to speak of, and things emerge quite naturally, and it's quite an exciting phase to be in. Things are moving quite quickly. For us, what it looked like was we figured out, OK, running without Docker and on local machines is probably not the best way to do things. We need some kind of cloud-based infrastructure, so at some point we set up Airflow, right? And we figured out how to run our DAGs on Airflow, how to set up projects. We even wrote a Python framework to generate Airflow DAGs from a a YAML file so we could onboard new projects easily. We also started creating some documentation, onboarding guides, right, how do you get access to all the tools, which tools to start using when you first join the company, and also just some shared Python libraries and modules that we could put some like common code in so that all the projects could share those, we could maintain it a bit easier, right? And already, looking back, it's kind of impressive. Like it was a fully functioning platform in a sense, but there was no like proper maintainer or owner of it. And I think that's actually really good, because at this point, it's like a user-driven platform, right? Users creating the platform for themselves, and so it works quite well without even thinking about it too deeply. But problems do start to grow, right? I mean, without paying too much attention to the tech debt, eventually, for us at least, things did eventually, I would say, get out of control. In particular, we hired more and more data scientists, we tried to tackle more and more different types of problems and use cases, and I think at some point it became quite clear that things were just not working like they used to, right? You talk to a data scientist and you hear them say to you, you know what, I don't really know when am I doing data science anymore, right? Like I'm spending all my time, I don't know, debugging and fixing things, this one country failed, the data's invalid here, these data types are broken there, right? All these little problems that pop up. The infrastructure's not working, okay, who knows how to restart Airflow, right? And these are the issues that we faced over time. And at some point, we had to really figure out how to improve the situation. And I think as you grow, it becomes quite obvious that the boundaries, the responsibilities aren't clear. It doesn't really make sense for the same data scientist to maintain all of your infrastructure and be working full time on some specific data use cases. And so for us, the first step we took, which didn't actually solve anything at first, was was to repartition our work a little bit. And this was really the start of a new development, a new direction of things, because we had dedicated people that were focusing on, like, Airflow, the infrastructure, the Python tooling that we had built. And at this point, we didn't call ourselves a platform team, and we really weren't treating our platform like a product yet, right? We were just saying, hey, you've been at the company a long time, you know how all of our infrastructure works, you're pretty good at it, I know you're a data scientist, but how How about we call you an ML engineer now, we shift you one disk to the right, give you a new hat, and that'll solve our problems, right? And eventually it did, of course, clarifying responsibilities is very important, also the data science team started to segment themselves along specific domains and use cases. But it also took a lot of time, to be clear, to actually get things under control and pay down this technical debt and get things stable again, and then also deal with all the constant changes happening in other parts of the organization and keep things running. That does also take quite a bit of effort, to be honest. But eventually we were able to achieve that, and this is kind of the era or the phase where I actually joined the company. I joined directly into this ML engineering functional team at the bottom. And so there was this subtle transition as we hired more people and things evolved. This team at the bottom became much more engineering focused. We weren't data scientists who got renamed to engineers, we were actually engineers hired into the company. And of course, we had a lot of things to resolve. We had to support a lot of new use cases. We had to do some ad hoc firefighting. But of course you prioritize things and eventually you figure out how to not just patch it, but to get things in a pretty healthy and stable state. And once we got to this point, I think this is where really we started to have a series of small existential crises where we thought to ourselves, what is our job actually? Why does our team exist? What is our role in the company? And this is what led me to discover all the things I'm sharing with you today about how we should think about platforms, right? And now we're officially rebranded to a platform team, and we're also redefining how we collaborate with data scientists, et cetera, so it's an ongoing journey. But this is, I think, one of the big turning points. And when we achieved this, where we figured out how to manage the tech debt, we had some space and some time to pause and think and prioritize, some interesting things actually happened. The elephant in the room for us really was Airflow. I don't know, I guess many of you have experience with Airflow, I see some chuckling, but no one really liked Airflow, right? But we just kind of accepted that Airflow was our master and we were just going to serve it forever, and data scientists figured out how to make it work, and we built layers and abstractions on top of it to try to make it better. But it wasn't really working in many areas, and we start to think, okay, like, what's going wrong here, or how can we fix this, and how can we make this a more long-term sustainable platform or product. To give you a bit more background about how we did things, and how we still do things to be honest, we were using Airflow, we had developed this framework, we called it One DAG internally, where you could essentially spin up these DAGs, which would partition across our different countries, and partition your DAGs into concrete steps that are kind of standardized across all the projects. And so there were some really good things about it, to be fair. For production stable runs that weren't changing often, it was perfect. You have a thing, you have a dashboard, you have a UI, it runs, and when it breaks you go read the logs, you re-trigger it, whatever. The YAML interface was also pretty nice for onboarding new people. You didn't have to learn all of this Airflow magic and operators. We could have data scientists just kind of use Airflow rather than having to operate it. But I think the biggest thing that it didn't solve for us was the development and prototyping, which is kind of the most important thing that a data scientist should be doing. Also it was quite rigid and restrictive, so whenever the projects would scale or the use cases would change, people really had challenges to say, or they would come to us with requests like oh, can we do this in Airflow, and we're like ah, maybe, but we'd have to really figure that out and keep it backwards compatible, and it became very difficult to kind of keep up with all the changing requirements. And so what did we do? We kind of took a step back, and we thought about, okay, if Airflow isn't it for us, what else is there out there, right? What can we look at instead? And we did a survey of the landscape. We looked at a lot of tools in some detail, trying to figure out what could work for us. We did look at Kubeflow. I think it's the most obvious choice if you look around the internet. It's like the reigning open source champion. And at first, it was like, wow, this is amazing. We've been used to Airflow. had our own abstractions to move data between tasks, and now, like, Kubeflow, you can just do it, and it works. But somehow it didn't work for us. We just tried to reimplement our Airflow use case in Kubeflow, and then we found some GitHub issues where they said, yeah, sorry, the DSL doesn't support it. Better luck next time, right? So we were a bit worried that we would get locked into this rigid DSL that really didn't solve our use case and would again lead to the same problem we had with Airflow so far. We also thought to ourselves, maybe Airflow is not the problem, maybe we're the problem. Maybe Airflow actually is a good tool, we just need to learn how to use it. And we did a little bit of research here, I mean, to be honest, I think we all lacked a bit of motivation to do this properly, but there were cool features that Airflow had that did seem like they would solve real problems we were facing, but all the time it felt like it was very difficult to adapt it and make it usable for our data scientists, right? It would really take a lot of work to package all of this in a way and make it usable, whereas we look at these other open source frameworks and it seems like they've solved these problems from the jump, right? Daxter was a pretty cool one. I think what I learned from Daxter is how important local development is. They have a really nice way to just spin everything up locally, run purely on local, and then you can later port that to the cloud. But it didn't really work with containers and Docker images the way we might expect, and we had a lot of legacy stuff flying around, so we were a bit afraid to kind of completely scrap everything we had and start from scratch. Argo workflows was our kind of reigning champion for a while. We actually are still using it to some degree. If you don't know, Argo workflows is the backbone to Kubeflow, and surprise, surprise, it's a lot more powerful, but it requires a lot more work to actually use. But we could port over our Airflow stuff into Argo workflows pretty transparently and actually have it kind of backwards compatible for a while. And the main limitation we found with Argo workflows was that you have to do everything in this Kubernetes native YAML way. And you know, you have templated YAML, you have for loops and if blocks and this and that within YAML blocks, and I mean, I had some people on my team who could manage that. I sat down for a day and I just had so many bugs and weird error messages I couldn't figure out and I was pretty confident we can't give this to data scientists, right? Like we would have to maintain the YAML templates for them if they had a question, again, they would have to come to us to extend and adapt them to their needs. And so how did we get out of this kind of confusing situation, how did we decide where we go next? I think there was three main factors that came together for us. The first thing, I want to emphasize again, we talked to our data scientists. And what did we learn from talking to them? I think at one point I interviewed 10 different data scientists within the span of a week or two, and I got 10 different answers about how they actually do their development. Not only process-wise, but also which tools they use, what infrastructure they use. We had people launching custom kube jobs because the airflow was too rigid. Different types of frameworks to visualize and understand the results. It was just a chaotic mess, and it all made sense at the local scale. Everyone was finding workarounds to their problems, and no one was doing anything wrong, but it just really proved that our tooling was not solving their needs, otherwise they would be using it. In summary, prototyping development needed to be a lot easier. Onboarding and understanding the system needed to be a lot easier, and also just in terms of scaling and dealing with these larger and larger amounts of data and complex models needed to be able to be handled in a more flexible way. The second thing that happened was Metaflow announced Kubernetes native support for Argo workflows and after hearing about this and reading the docs a bit, it really, like, everything somehow clicked. Like, we had been fighting with tools for so long and we kind of skipped Metaflow because it was at first only in some AWS-managed fashion. And we weren't really excited to figure out all of that tech and how it works. We're pretty Kubernetes-centric in our deployment so far. And Argo workflows we also were pretty confident about. We had been using it actually at scale, and we trusted the tech, and we also trusted that we could kind of understand and invest more to become more experienced with it. But most importantly, Metaflow solved a lot of the problems that we were having when we we tried to use Argo workflows directly. So Metaflow allowed you to really do purely local development, have a really clean Python DSL to create and generate your DAGs, to easily transition with a few CLI arguments from local development to like cluster, like Kubernetes development, and also use Argo for your scheduling needs, et cetera. And so we really became convinced that if we did it properly, that Metaflow could be the right tool for us. And finally, we took a closer look at our priorities. So I think, especially the early days of exploring Argo and Kubeflow, we always kind of did things half-half. We were thinking, okay, we have this, like, air flow system that we can't really get out of, and so we tried to find, okay, can we find a development tool that can kind of complement air flow, but that necessarily puts you in a position where you're, like, dealing with the worst of both worlds. And so finally we kind of decided, you know, if we can find a better tool, then why not make that the tool of choice, right? It's not going to be easy, of course, and I'll get into that in a second, but this is is ultimately the decision we came to, that we should aim higher, that's one of our values at our company, we should also not just focus on incremental improvements or adding new things to help solve the problem, we should also figure out what's the legacy stuff that we want to deprecate, what's the stuff that's creating too many support requests and creating too much friction for our users that's actually worth spending the time to eliminate or migrate away from. And with all this in mind, so this is the decision we made in the end. We did some pilots and prototypes with Metaflow, we got great feedback from the first early adopters of it in our company, and now we're in the process of actually figuring out how to make this the tool of choice. We still have stuff running on Airflow, we still have stuff running on Argo workflows, so it's still a bit of a mess and it's still a learning process, but that's where we've ended up now. And that's, I think, kind of the closing point I want to highlight here. So I don't want you to take away that Metaflow is the answer to all your problems, right? It definitely isn't. Everyone has their own context and situation. For us, Metaflow seems really great so far, but it also opens up a whole host of new problems, and to be honest, it makes a lot more work for us in the beginning, right? But we commit to it, of course, with this long-term vision in mind. So always keep that in mind, right? I mean, legacy projects and migrating them, that's always going to be a pain. Also just convincing people that this is the right choice to make, right? Some people, they've gotten so used to Airflow in our company that moving them away from it, despite how much they might hate Airflow, is still going to be difficult, right? They have to unlearn certain patterns that they've developed over years, and they have to figure out how to adapt, and we also have to figure out how to develop new tools on top of Metaflow that actually fit their needs properly. So yeah, it's not an easy path, being a platform team, in summary, but I think it's very worthwhile and very fascinating to figure out how to make this stuff work better for your data scientists and help your company achieve much more on the whole. That's all. Thank you, everyone. Again, I have a, I think we have some time for questions, but I'll also be here all day. You can ambush me and find me if you want to talk about any stuff in more detail that I didn't cover in the talk today. I'm super happy to meet anyone else here who's working on ML platforms or ML engineering or data science at scale. Thank you, Cole, for the amazing talk. have time, five or ten minutes for questions. Here's the first one. Thank you for the talk, first of all. I'm curious, you mentioned once as a side note that you, I think, we redefined how we collaborate with the data scientists. I'm really curious about, like, the reasoning behind that and the old and the new workflow. Sorry, could you repeat? Sure. You once said, we redefined how we collaborate with the data scientists as a platform team. Redefined, you mean? Yeah. Basically, you had one old process, and then you basically scratched that and came up with a new process. Got it. That is my understanding. I'm really curious about that. Yeah. So I think at the beginning, because we have to remember, at the beginning, we were one kind of cohesive team, and then we just kind of split into two chunks. And so the boundaries weren't very clear. And so what we kind of naturally became was more of a support team, right, an infrastructure team where we would deal with ad hoc problems, also help implement some, like, Greenfield use cases from scratch. And what we're trying to currently redefine ourselves as is more of a platform team, focused more on self-service and documentation, and really empower data scientists to do things on their own rather than having them, forcing them to come to us with problems and then wait for us to solve them for them. So this is the core of how we want to change the collaboration model between the teams so that each team is as independent as possible and has its own kind of vision and goals in mind as well. Thanks for the talk. My main question is about your relation, your team related to other teams. You set up this platform. You're using some kind of infrastructure. Did you set up everything from scratch? Is there a data engineering team or an infrastructure team? Yes. Very good question. Yeah. It's a complex landscape. And it's always evolving. We have always had a very strong, we call it a foundation team. They manage all of the Kubernetes clusters and core infrastructure, Terraform, Helm, all this kind of stuff. So we build on top of what they have, but what we need to do for data science is very different from what they usually provide. So we often have to kind of build another layer on top of this. We also have a pretty strong data platforms within Delivery Hero, especially for data warehousing and data curation is what we call it. And for a while, actually, currently, these data teams actually do manage the Airflow infrastructure, but now with Metaflow, we've become the infrastructure team for Metaflow, which is also an interesting change. It's a lot more work, of course, right? You have to think about these kind of challenges as well. Okay. Thanks. A small follow-up question. So what kind of skills do you, as ML engineers, now need that you didn't need as data scientists? And what kind of skills do you not use anymore compared to data scientists? Yeah. So I think email engineering, right, you always have, like, platform teams, you also can have embedded email engineers. And I think those are quite different roles in the end. And I've, yeah, a lot of people, I think, struggle to figure out how to define these. For us, like, as a platform team, there's a lot more infrastructure work, to be very clear, right, especially at the beginning of these projects. So we need to know a lot more about Kubernetes, right, we need to know a lot more about Terraform and Helm and all of these core technologies using the cloud providers, right? And try to make that as easy as possible for data scientists to get into. And yeah, we're naturally a bit further away from the actual direct modeling and the business cases that go launch into production, right? We still collaborate and pair up very closely with data scientists when they have issues and to help them unblock them, right? But, you know, we're not every day sitting down with Scikit-learn and figuring out how to make these features and these models work for these specific business cases anymore. Thank you. Hi. Aren't you afraid of linking your software, your data science teams are writing with the infrastructure by doing these type of code annotations and things? Because how Metaflow looks to me is like now you're injecting infrastructure into the code and now you're linking it. Yeah, yeah. And don't you want to have this flexible? So aren't you afraid of having this because now you're dependent on Metaflow? Yeah, this was one of our hesitations for sure. Because with Airflow we had several layers, right? But the problem with Airflow is that the DAG logic was here and your project logic was here. And so there was never a way, if you wanted to make a change to one or the other, it just became very messy. And with Metaflow we can combine that all into one repository. Still, I think it makes sense to, you know, the Metaflow files should define the DAGs in the infrastructure, and then your actual business logic should be separated to separate Python files. This is a best practice that we would recommend. But it does mean there is room for some more messy stuff. But it's a trade-off. For us, I think we've learned that data scientists need flexibility, especially in how they structure their DAGs. So for us to scale some projects, they need to train at country level, city level, zone level, this and that. And we can't possibly think of every single use case they might have and provide that as a YAML config somewhere, right? It's just not an efficient way to work. So what we're aiming to do is give them the power over the infrastructure and provide as much support and abstraction so that they don't break things, of course. But that's the trade-off, yeah, it's not an easy choice and we thought about it a lot, we had a lot of discussions on this exact topic, so yeah, that's where we stand on it. We think it's useful and from our experience so far, but there are some downsides, definitely. Great talk, thanks. How do you handle deployments to production, or in other words, who will be called on the weekend when there is a problem on prod? Is it your team or the data scientists? Yes, this has been like a running joke almost. Every couple of months we talk about on-call teams. So luckily in Delivery Hero we have this definition of tier 1, 2, 3, and 4 services and we've designed and generally all of our data science production models to be in the tier 3 category. So we always have other systems in place that are proper software engineering maintained systems with on-call engineers which can fall back gracefully if the data science models are broken. So for that reason we don't have dedicated on-call for like weekends and stuff like this for data science or for our platform team. So we can handle some downtime. But this was the problem in the old setup, where we kind of felt responsible for all of the Helm deployments and Kubernetes deployments. So when things did break, we kind of got looped in somehow. And it really felt like we should be an on-call team, but we weren't, which was a bit confusing. Now the direction is, we're actually moving also the data science into more cross-functional teams, so they will have dedicated software engineers as well, which would be much better suited to be on-call. Because for us, I mean, we have maybe three engineers maintaining 10 different model services. It's just not feasible, right? So our direction is more that the data science teams should own their systems end-to-end, also ideally be cross-functional and have their own on-call if needed. Last question. What is the reason why you're not considering cloud-native tools such as Amazon SageMaker? It's a good question. I think one reason, big reason, is actually cost. So we run our own Kubernetes clusters on spot and it's just orders of magnitude cheaper than paying for like even Vertex AI or SageMaker. Another reason is just like all of these cloud solutions, at least from my experience, you usually need to kind of build another abstraction on top of it to make it fit with the rest of your systems and tools. And so for us there's really no strong benefit to use it. Like we might pick and choose a few things here and there, but in general we're much happier with running our own code and our own Docker images on Kubernetes. It's much more, I guess, flexible, and we can adapt it to our needs as we want. Well, that's all. Thank you very much for the questions, and Cole, again, thanks for your talk. Thanks, everyone.

Cole Bailey

About — in the speaker's own words

I am leading a scrappy but growing team of ML engineers at Delivery Hero who aim to bridge the gap between software engineering, DevOps, data engineering, and data science. I hope to make data science easier without restricting the creativity and flexibility that data scientists need to make an impact in their role.

Social card for talk: Cooking up a ML Platform: Growing pains and lessons learned