Public Money, Public Experiment - open source processes in the public administration

As one of many data labs in the public administration, sharing code and software increases the speed with which technical problems can be solved and reduces overall costs. In the previous months, we started collaborating with other public units to share a python prototype between labs. Now it's time for the next step: as we approach PyCon DE & PyData Berlin 2024, we aim to make code publicly available.

The presentation will address the following questions:

  1. How can the process of publishing code look like in a public administration and where can you get access to code already published? (Spoiler: Check out OpenCoDE)
  2. How does open source align with public administration principles?
  3. What legal and political and security requirements shape the process and possibly the code base?

Whether we succeed or encounter challenges, this talk serves as an attempt to transparently share our journey and contribute to the broader discourse on the intersection of public administration and open source initiatives. Join us at PyCon DE & PyData Berlin 2024 and stay tuned for a glimpse into the evolving landscape of our code publication.

This session took place in track Others.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:05]

Okay, hello everyone here and remotely joining us. Thank you for coming and thanks for the introduction. Yeah, I'm very excited to have this platform here today to talk about open source processes in the public administration. So, can I ask quickly who here today is actually working right now in public administration? Okay, cool. You can also raise your hands online if you want to. So, we have some in the room. And the next question, who is working with public administration? Okay, cool. My name is Lisa, and I do work in public administration. I'm a data engineer at the data lab of the Federal Ministry for Family Affairs, Senior Citizens, Women and Youth, or short, Family Ministry. And I wanted to introduce you to the world of bureaucracy with a story about my journey as a data engineer, starting a job in public administration and as it usually goes it started really well until there was a problem but first things first last July I started working at the ministry and in case you are wondering what a data lab is they are in-house tech units that are being established right now in all the German federal ministries to improve how data is handled and in the ministry and to increase data expertise within the government body. And one of the first big projects I did when I joined was to do a data inventory for a data catalogue with all the relevant data sets of the ministry, which was really fun for me to be able to explore what kind of data sets there are. But, of course, it's also a bit cliche data engineering project to build a data catalogue. But for me it was a great project because I got to know everyone since we are collaborating with all the departments in the ministry. And while I do love talking to the colleagues, I was also glad that next to collecting all the information, we are also working in parallel on the software for the catalogue. And we did some market research and started to build a prototype with C-CAN. And CCAN is an open source Python-based application. It's commonly used for data management, it's a data management system for data portals or catalogs, and what we like about it is that it uses, it's composed of different modules which can be adopted to your needs. It has a free and open source license so you can adopt it to your needs with extensions. And for me, as an open source enthusiast, it was a great start to be able to work with an open source tool, because in our environment there's usually also a lot of closed source applications. So we took this default, this is the default catalogue you see here, and we took this and build a prototype, we started to do this, so change the front end basically to our in-house CI, maybe you can recognise the little blob from my first slide, this is our CI, and we were able to do this thanks to the licence and could start prototyping with this to see if it fits our technical and non-technical requirements. This was a good start, being able to work on an open source application. Now, we are not just one data lab, but actually as I mentioned, there are data labs in all the ministries now, and we are in a network. It's a peer-to-peer network where we communicate about what we do and try to exchange insights with each other and we showcased our prototype in one of the network meetings and as it happened some of the other data labs were also building catalogues in order to establish like a data governance in the long run in their houses and the question was well can we share our prototype with them and yeah so this was the next good thing like we were working on something and there was already demand for it from other projects who are working on similar things and now the question is like can we share what we are working on with them and actually there is something called the EVA principle. EVA stands for einer für alle or one for all and that's It's like a principle in the public administration when it comes to IT projects, which means that if you are building something, you should actually try to also make it available to other units in the public administration in order to save costs and not do duplicate work and be more efficient. So actually, with this principle, we were even encouraged to share our work with the network. So far we have a pretty use case for open source, someone builds something, others want to use it, you can share it, but when we wanted to actually do it, we ran into some roadblocks, and the roadblock is the infrastructure. Now, our ministry is not a software company, but a bureaucratic organisation, and there are no established processes to share code, so we didn't have a GitHub account. I'm not going to start with fax machines. Because of this, we didn't have infrastructure to share the code in an efficient way. There are ways, but yeah. And since there's no infrastructure in place, and there are no processes established for publishing code. We need to actually establish them. Now, if you want, if you think about public administration and you want to establish a new process, you need to keep in mind that the administration is optimized for stability. And this is good because states want to be able to function in times of crisis or chaos. But the flip side is that new processes take a lot of energy to establish. But the good thing is we are an innovation unit and we have the energy to do this. And since we have the word lab in our name, we decided we have a use case for open source in our hands and so we asked how much progress can our data lab make in publishing code within the three months leading up to this conference. So now I would like a quick show of hands, also online, if you want, who thinks that in about ten minutes I will show you our CCAN extension on the public website? Okay. There are a few believers. Cool. Let's find out. I will walk you through our process from the past three months, and we will see where we ended up. So first I will go a bit more into the detail of the situation we were in in February. Then I will go into how our endeavour was planned and executed, and along the way I will highlight some learnings. And then once we arrive in the present, I will summarize the results of our experiment, and we will also look ahead at what we are planning in the future. This is our project plan for the experiment, so on the left we have the four stages, and then some steps that we will take and some keystones we want to accomplish, like access to the first access to a code publishing platform, setting up a repository, and establishing a code publishing process, and underneath you can see we are also in parallel working on moving our prototype from the prototype into production, but this will not be part of the talk. That's probably two more talks. So, yeah, let's dive into the initial phase of what exactly is our problem. And let's go a bit back to the specifics of our data lab and what our situation is right now. So as I said, we are a new unit with staff that has more technical and data-focused skill sets. So in our team, we have data engineers, data scientists, and analysts. And basically, we try to make the German public administration fit for the digital age. data labs and their tasks are conceptually rooted in the data strategy from 2021 and 2023. And we are complete as a team since November last year. This is when we started, like, we finished the build-up phase basically. And one thing I want to mention is that how we work is like in established processes. There are established processes that define and govern how we can work, and one very important factor is IT security. We work in high security environment which means that internal products we develop must run on-premise and without outside internet connection, so cloud solutions are not our first go-to. So basically IT security is always on our mind, and as a technical unit, or like as a more technical unit than your normal public administration work, we also have new requirements, right? So for the main tasks in the ministry, like the office suit is fine to work with, but have you ever tried connecting to a database with Word? Probably not. So with our new skill set that allows us to develop data products from within the ministry, there also come new requirements like access to a code publishing platform. So let's come back to our use case for sharing our prototype. This is the problem we face. First, we need to establish a process for publishing code, and second, we need infrastructure, so we want to publish our prototype or even the production-ready CCAN extension. So we look at how can the process look like and where can we publish code. Let's start with the process. So when researching, we came across the Berlin Open Data Handbook, which I can really recommend with anything related like how can public administrations publish data. But I think they propose like a publishing check and a data protection check before something is published. And even though the guideline is for publishing data, I found it also helpful when thinking about publishing code. And yeah, But there are a few things to consider before publishing data or code. So for example, a loss of confidentiality, like in a publication check should consider whether the publication can lead to a loss of confidentiality. If secrets are disclosed, I mean we had a few security talks where the talk was about access tokens being published, like that's the kind of thing you don't want. And then there's questions about compliance, are you publishing something you don't have the right to publish, are there copyright issues, and then in the end you always need to think about the effects this can have, whether there are detrimental effects on public safety, information security, or the conduct of legal proceedings. And then in the process of thinking about this in our team, we also came up with further considerations in this process, like the question whether we can publish unfinished software and who is responsible if there are errors or security issues in the code that we published. Now on to the infrastructure, it was pretty easy to find a product, and it seems to be pretty much what we were looking for, so Meet OpenCode. It's a code publishing platform for the public administration, and OpenCode aims to simplify and promote the use of open source software in the public administration. It's run by the Center for Digital Sovereignty, CENDIS, and the project was initiated from the Federal Ministry of the Interior and two German federal states. It's geared towards public administration, which we will hear later on. Some examples, for example, the German data catalog GovData has their code published there. In general, when it comes to code publishing, Schleswig-Holstein, I can really recommend to check out their open code. They're really great with these initiatives. Okay, now we have specified our problem and looked at ways how we can go forward so we can continue with the planning phase. And what I learned in my short time in the ministry is that it's always wise to first think about your stakeholders. Like, oh, okay, who are the experts you need to talk to and who has to be involved in the process? and who can give you the final go. So first, this is what we did. We started with mapping the stakeholders, and first I talked to my unit, and we decided that this is a new process, and we will have to get the approval of our Secretary of State in order to open an account. And then we also talked to our IT unit and some other specialists in our organisation. And this process took a bit longer, so let's move our timeline a little bit. And now we are in March, and we needed to get the stakeholders' approval. In this case, our Secretary of State's approval. And how this process works in the public administration is with a decision template, and for those of you who work in public administration, it's always a lot of fun to work on Leitungsvorlagen. And for those of you who don't know, there's basically three things to consider. You need to list all the stakeholders who need to give their goal for you to go forward. You need to state the facts of the case, of the decision that you want your Secretary of State to make, and then give a recommendation of what you think they should decide. What we did is first then this Leitungsvorlage goes on the horizontal, so all the stakeholders sign off on it. In this process we found out that in order to sign up for the code platform we actually need anonymous emails, that's something we didn't know beforehand, and here it shows that OpenCode is one of the few platforms where you can actually have anonymous email address, they provide it for you when you sign up, so this just shows how they really adapted towards the needs of public administration when you sign up to a code platform. They also have extensive guidelines, for example for licenses, which I can also recommend for private people to check out where they have established licenses that are good to use when you are in public administration. And then it goes, our Vorlage goes all the way up to the Secretary of State, then it comes back, and we were successful, we got the goal to go ahead on the 27th of March, So this was time to celebrate. And now let's go to the next step of setting up the infrastructure and setting up the repository. And here I brought a little screenshot from a first hello PyCon repo we set up on the website. And if you scan the QR code, you can see it, first we did it private, but now during PyCon we launched it live. So you can go check it out, hopefully it works. And yeah, now we have set up our repository. But as you can see, we are now in the now time. And so, yeah, this is basically where we are right now, and the results basically are that in the last three months, we managed to set up the infrastructure, get access to the code platform, and when it comes to the publishing process, we are still in the process, but We have published a first repository, not the extension, and the template for the process now exists. We have uploaded a template for how you could write a Leitungsvorlage for any other public administrations that might need to go through this process, and it's on our repo, so please share it if you know anyone who's in our situation. I will skip through this, I'll go a bit quicker. So looking ahead, we are still now in the process for defining like a real publishing process in our house in cooperation with the information security officers and other units to think about like which steps this publication check and the data protection check entail. And when it comes to this, like who is responsibility, When it comes to the considerations from before, can we publish unfinished software? Is software ever finished? That's the real question. And on open code, you can mark your project according to the stage it is in so that others also know, like, okay, this is still an early-stage prototype. When it comes to the responsibility for security issues, we still haven't fully answered this question, but our current understanding is that the unit who is actually operating their system and their software are responsible for the security but OpenCode in the future is going to implement security screening in cooperation with the Federal Office for Information Security so that the code that is on the platform will have a screening result and it can already give you a first information about the state of the code that you might want to use yeah we will move forward with our prototype to production process and we hope to finish this in the by the end of the year and who knows maybe next PyCon you can see the published extension but we are still on the yeah on the way there and if you take anything away from this talk public administration is not only fax machines data labs are technical innovation units within the German federal government and please spread the word about OpenCode it's a GitLab for public administration thank you so much and I brought some more readings if you are interested

Speaker 2 [21:50]

Thank you for showing us we're not just about faxing life. It's very important. I have a few questions, so I'll skip the ones I'm particular about. How are you cooperating with the kind of recently launched open-source competence center? The what? The open-source competence center for Berlin.

Speaker 1 [22:11]

We are not cooperating with them right now.

Speaker 2 [22:15]

And how is the OpenCodePay part of the Europe interoperability or is it just really focused on German government? Can you say that again? Is it cooperating with the Europe interoperability or is it just focusing on the German government?

Speaker 1 [22:32]

For now, we are, like, focusing on our internal administration, but, and, yeah. So, oh, yeah, one thing, it's also linked here, the family, like, our ministry is also funding a civic data lab, which is, for example, then more geared towards, like, NGOs and the civil society to help increase data literacy there and do data-based and data-driven projects. So this is something... Our data lab is for the internal processes to make systems speak to each other in an automated way, hopefully. And the Civic Data Lab, for example, is geared towards the public.

Speaker 2 [23:20]

I have a bunch of questions. How competitive are salaries for data jobs in public administration? Is it hard to find talent?

Speaker 1 [23:30]

I think right now there are two positions open in one of the other data labs, BMWK. I think they are looking for two data engineers right now, and you can see the TVÖD. It's a set salary, basically. Yes, tariffvertrag, exactly. So you can publicly kind of view what the salary ranges are. I would say they are competitive because we have something called the IT-Zulage, which can be like a bonus for technical jobs in the public administration. So this is one instrument where the government or the public sector also tries to be competitive in this field.

Speaker 2 [24:22]

And is the code, the open code, really open source, or can every private person participate, create PRs?

Speaker 1 [24:32]

So the access, like everyone can read what's on there. When it comes to publishing, I'm not 100% sure, but it's more geared towards public administration units or the cooperating partners that they work with. So in our case, we have a technical unit in our ministry, so we can at some point push code. But, of course, not all the public administrations, also at the communal level or like on the lower federal levels have these data labs so they can also work with contractors and then these contractors can upload code for them onto the platform this is like something that works and in the wiki there's like a faq which is also very much geared towards non also non-tech people who want to find out like how can i talk to The contractors are like what what does the code need to look like to be uploaded to the platform?

Speaker 2 [25:35]

So it's like a transparency to see, but there's a more structured, very strong governance model for you to actually contribute.

Speaker 1 [25:48]

You can read it up on the on the website how actually you can contribute but I think the use case is more that public administration can can publish it and Yeah, take take it from each other based so it's a bit like the safe if our principle that the the public public administration units can share code with each other and Others can also see it

Speaker 2 [26:19]

I consume public data. These projects often end up unmaintained and unsupported. What's the plan for when the public money moves onto the next big idea?

Speaker 1 [26:38]

I don't really know, like, how to answer this question. I don't know what the next big idea is or if something will move on. What I know is that this is, from my perspective, like, working in the public administration, having a platform like this gives us the ability to actually, like, even start open-source projects and work together. So for me, I see it more like an opportunity to create an environment where these projects can be maintained and move forward.

Speaker 2 [27:20]

Yeah, I believe for the one who asked the question, it's more about, or the question is more about, correct me if I'm wrong, about putting governance models in place to reassure the memory, the documentation, and the transparency, and I think particularly the questions that I've been asked, how is this maintained, who fixed the bugs, are important questions for this question.

Speaker 1 [27:45]

and

Speaker 2 [27:46]

And we have like one or two more. Are the anonymous email addresses also used for all the git commits? How do you track code on or review with an anonymous email?

Speaker 1 [28:01]

Yeah, so the anonymous email is optional you can also have like your clear name or like your username Some some do but there is the option to have the anonymous

Speaker 2 [28:15]

Did you also discuss not only how to share code but also to share the data over the data labs?

Speaker 1 [28:24]

Oh, yeah, of course. So what we do in our house is also like a building block for something called the Darton Atlas, which is also why we like the CCAN, because it has an API, so you can also harvest. Like, CCAN instances can harvest each other. And the Darton Atlas is basically like the overarching project where it's a data catalogue for all data sets in the government. And, yeah, so basically these will flow into each other at some point.

Speaker 2 [29:00]

What would you do differently if you would have faced a similar change again?

Speaker 1 [29:16]

Start my talk preparations earlier.

Speaker 2 [29:24]

Maybe one last one to push the time. What kind of technical projects are you working on that you are excited about and is the open code a full feature in GitLab including the CI CD and so on?

Speaker 1 [29:41]

What am I excited about? I'm excited about this and being able to move this prototype into a production environment. For me, it's a very interesting aspect, the IT security aspects of it and how can you actually make sure that something runs internally. I think it's very nice to build flashy prototypes and be very quick and show something, but then the real work starts when you actually try to implement it under these conditions. Conditions and that's something I'm really excited about and the second part

Speaker 2 [30:17]

If it's all the structure, not just the code, like the pipelines and the CICD structures also, but also at this, because I think it's quite important, is that an accessibility goes for the project.

Speaker 1 [30:30]

So there are accessibility, there are guidelines, I also linked them here, so there's basically like a set of, like a compliance thing that you have to acquire to, like it's a requirement for projects that we develop.

Speaker 2 [30:51]

Perfect. Thank you so much for sharing that.

Lisa Reiber

Lisa works as a data engineer at the in-house data lab of the German Federal Ministry for Family Affairs, Senior Citizens, Women and Youth.

Social card for talk: Public Money, Public Experiment - open source processes in the public administration