Is digital sovereignty a new buzzword in AI development?
Digital sovereignty in AI development addresses the risk of dependency on foreign technology providers, political instability, and restrictive legal frameworks. A primary concern is the use of "black box" software, such as Palantir Gotham, which lacks transparency regarding training data and bias, making it difficult to maintain or audit. Furthermore, reliance on hyperscalers like AWS or Azure introduces legal conflicts; for example, the US Cloud Act allows American authorities to access data stored on servers located within Germany, potentially contradicting the General Data Protection Regulation (GDPR).
To achieve true sovereignty, developers must evaluate five critical clusters: hardware, network, data, platform, and software. Reducing dependency involves moving away from API-based models toward open-weight or open-source models, such as Toykin or Apertus, and deploying them on-premises or through local European service providers. This modular approach prevents vendor lock-in and ensures that individual components can be updated without compromising the entire system.
Key takeaways emphasize that digital sovereignty does not block innovation but rather secures it by ensuring long-term maintenance and predictable costs. While large-scale models are often marketed as necessary, many professional use cases—such as information extraction and classification in the legal sector—are effectively handled by smaller, open-source models. Adhering to GDPR principles of data minimization and informed consent is essential for avoiding the high-risk classifications defined by the AI Act. Ultimately, sovereignty requires a strategic shift from rapid scaling via global cloud providers to ethical, transparent infrastructure that protects fundamental human rights.
This description was generated by Open-Source AI using the transcript of the session and the original submission contents.
This session took place in track Ethics & Privacy.
Submission
The proposal as submitted by the speaker before the conference.
AI development typically prioritises feasibility and implementation. While solutions should be efficient, high-performing and scalable, sovereignty and data security are often overlooked. These issues tend to be overlooked when solutions are being found, even though we don't use AI as an end in itself, but rather to benefit or support our customers. Customers operate within a regulatory framework and rely on responsible technology. Rather than seeing regulation as a hindrance, we should view it as an opportunity to drive innovation and create sustainable, trustworthy solutions. However, this is only possible if we understand the full meaning of sovereignty. This presentation will explore the various aspects of the term 'sovereignty' and its potential impact on AI projects. We will discuss current examples from politics and development to identify best practices for secure data processing.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:00]
physicist by training and many, many other things, including an ambassador for women in AI and with a long track record of being an AI expert with a focus on responsible and ethical AI. Her talk today is digital sovereignty, sorry, is digital sovereignty a new buzzword in AI development? So we will have 25 minutes talk and then afterwards, any questions, please ask them over the talks URL. We don't take the questions in the room just over the website. So with that, I'd like you to welcome our speaker and enjoy.
Speaker 2 [00:39]
Thank you so much. A few people are still left. I know it's more or less one of the last talks of today and of the conference, so very happy that you're all here and find the energy to still listen to something. So today I would like to talk about digital sovereignty and why is this like a topic for me or why do I want to get into the topic. It's more or less two different reasons. One reason is that Westernacher Solutions, the company I work for, create software products for courts, lawyers, notaries, also churches, and they have a very high security standard. And if we create like a sovereign AI solution, for example, we are on a safe side with the security level here. On the other hand, what I have seen within the last, I think, more or less one year, is that many people use sovereign AI system as a buzzword. They use it to have like a marketing thing going on, And also a lot of politicians say, okay, we need to support digital sovereignty. But in the same sentence, they say we want to use hyperscalers even more to have more innovation and everything. And I was like, okay, this doesn't go well together. And we already have heard in the keynote in the morning that this is not a very good idea to claim AWS setup as like a sovereign AI setup. So let's get into the topic a little bit more. but before telling you more about sovereignty and what it really means and what you need to take care about, I want to ask you who is aware of the Palantir software Gotham? Who has heard about it? Everyone. Okay, perfect. Then I don't have to explain so much. You maybe know that it's a very controversial software. The founder of Palantir is Peter Thiel, who is very close to Trump. And the software takes different sources of data, combines this makes like an analysis and can also predict some crime and so one state in Germany like the Hessen state uses this software already to check or to get into the crime scene of Reisberger who knows who what Reisberger are okay good then you don't need to explain it these as well and they always show the success cases here with Palantir and Gotham but they never show what can go wrong and what kind of bias you have here inside. And we have already seen from America that things can go wrong, especially if you use this kind of software. And if you ask our Minister for Digital Affairs, Carsten Wildberger, if this goes together like digital sovereignty together with Gotham software that we already use in different states, then he claims, yes, of course, it is okay. Because the data is on a local location and the data doesn't go out of this local location, yeah? So he only focuses on the data security, which is not really fair because digital sovereignty is not only data sovereignty, it is even more. So let's go a little bit more what it actually means. I mean, in the morning we have heard already that it's more or less about dependency, yeah? Who you depend on. And if you ask the Google AI, it will already tell you that it is also avoiding the vendor login here. So we try to avoid to depend on like a company or a different state outside of Europe or you would like to avoid depending on different political issues or changes that we already see quite often. So this is everything you would like to avoid. And something we really need to take care if we want to set up like a sovereign AI solution for example is that you have to know about your technology and your data. So how you depend in each of these parts on certain companies or different aspects, political issues, also data dependency, where do you get your data from, for example, is it like an open source area, is it the internet, is it like the customer data or everything? And both of these aspects, they go together with a legal framework. And this legal framework, I think it is very often in AI systems or AI projects, everyone forgets about these legal frameworks because these are the most important parts, especially in Germany. We all know about it. But in the morning, there was one question as well about how secure I think an AWS setup is. and it's not only about the data if it is like located here in Germany for example it is also more who has access to the data potentially and I will go into detail a bit later in the talk but first let's check what kind of topics we have within the technology part of the sovereignty here so first of all if we talk about like being independent of something we have like different aspects we need to really check within our whole AI system. The hardware cluster we already had the keynote in the morning so it was very great explained that you have to buy your hardware somewhere from some company you have a dependency on the hardware side. As well as the price. The price depends very much on the political issue the price depends very much on the availability of the hardware and everything so here you really try to reduce the dependency as much as possible you cannot like totally avoid it but you try to reduce the dependency here another part is the network cluster the network cluster is that more or less the communication channels between front end and back end and you try to be as secure as possible so that none of the data you put from the front and to the back end can be taken by hackers or somewhere. So you really try to cover everything in a secure way. Most of you, I think, they are already aware of the data cluster and that this is the most critical element in most of the AI projects. But here you also should check where do you take your data from. Is it like the internet? Is this kind of data quality still available in the future? We see or we already have heard about the collapse of AI models, so the data becomes worse because more and more generated data is available. So can you guarantee also getting new data with high quality in the future or not? So here you depend very much on the data source itself. Or do you use data from a customer, then you're more or less safe in the site, but still you should think what kind of data do you use at the moment to get your project running, but also in the future. Another part is the platform cluster. Here we have, of course, the on-premise solution that is more or less a very secure setup. You just need to buy all the hardware, set up everything. But, of course, we also have the issue here of using different cloud systems. Do you use a cloud system provided by a German provider or an American provider? And is the cloud cluster available here in Germany or outside of Germany. I mean, sometimes people say, okay, but I can install everything in Frankfurt in these cloud centers, but this is not really true because if I go, we pretty often use the Azure AI system and here some of the models are just available, for example, in Sweden or something. They are not all available here in Germany. So this might be a big issue for our AI system and we really need to be careful by choosing the right location and also model here. Last but not least, the software cluster, and here we really need to check what kind of AI model do we use. Is it API-based, is it open-weight, or is it an open-source model, and how do we deploy it at the end? So people also ask me pretty often, if I talk about the previous stuff, okay, it seems to be super difficult to get like a sovereign AI project at the end, and it might be super costly and expensive to set up everything, to buy everything, to have it secure and everything. So they tell me, okay, this is maybe an innovation blocker then. Why should we really follow up on digital sovereignty? But my point is it is not really an innovation blocker. It could improve innovation even because if you set up your project in a way that you have like different modules, like you have a module for extracting information another module for generating like summaries or something, or another for a classification task or something, then you can always make an update within each module. And this guarantees you that even in the future, you're up to date with your innovation. So I think it's not really a blocker. It more secures your innovation in the future as well. I mean, you always have to see it from a different point of view from my perspective. And if you have these modules, you can also adapt to other customers and also to other use cases you may want to have. So it could also lead to more innovation. Another point I would like to point out is maintenance. I think this is something many people underestimate in AI systems or AI projects because a big focus is on delivering the project, like just setting up everything and send it to the customer. But then what's happening after? And for the Palantir software, Gotham, for example, it is such a black box that only people from Palantir can maintain it. They can add new features, they can add updates or anything further, but this could also lead to a security issue. I mean, if you always have people from another country or another company who do the maintenance, then you can always have other people checking all the data, checking all your system, and they can take this knowledge away to America. And at the moment, I think this is not the very best case. Yeah, and another part is also the costs. So for the costs, we have the following. Many customers want to predict the costs even in the future. If I tell them, okay, the cheapest thing is to set up everything on an Azure AI because you pay by tokens. It's pretty cheap, but you don't know if the costs are still stable in the future. You don't know. Maybe the costs depend on some political issues. We already have heard maybe the Azure AI will be closed next year. I have no idea. So you also have to take a look on the cost part and how it is predictable for our customers in the future. Bless you. Another part is the data so here you already know most of the things that data is like very sensitive, you need to be very, it needs to be secured and that you need to be careful to handle it but what I want to point out is here transparency and access control. The point about transparency is we have heard already that the Gotham software is like a a black box. It is not transparent at all. We don't know on what kind of data it was trained on. We don't know how the data is processed. If we have biases, and we always have biases, we cannot avoid them. But the point is how we can deal with a bias. And if we don't know the data for training, if we don't know the model, we cannot really account for these kind of biases that we have in the systems. So transparency is always It's very important not only for the training data but also for the model and how it behaves. Another part is access control. I already told you that it might not be very secure to have your AI solution on AWS. Why is this? You always or people always tell me, yeah, but it is secure because it's my environment, you know, whatever it means. But at the end, if you use, like, the AWS server in Frankfurt, you still have the Cloud Act. Who knows about the Cloud Act? A few already. Perfect. So the Cloud Act allows the US authorities to get access to servers here in Germany that are located here in Germany. Which at the end means that the people from America, just say it like this, can read our data that is located here in Germany, yeah? So here you already see that we have a legal framework and we are not only living GDPR, of course GDPR is important and we need to take care of GDPR, but we also have the Cloud Act that is against GDPR at the end. But it's still there. And at the end you have to do some risk management to check if this is really worth it. What kind of data do you put on the cloud server on AWS? Is it like high sensitive data or not? Or do you need to put like an additional security layer around? Or is it really needed to use an AWS server? Maybe you can also use a server from a German provider because then we really have full data protection here. And maybe also what I want to mention and just forgot is about the access control for the maintenance, what I already mentioned as well. So also make sure that the maintenance can be done by you or by someone who's really reliable. I mean, also the point is that at the end, if people are just finding a new job from Palantir, who's then maintaining the whole system if it is just the black box, if it is really needed to have expertise here? It's not only a security issue, but also having the real expertise to maintain the whole system at the end. So what you see is that the legal framework is also super important. And the most important principles is coming from GDPR, you all know, hopefully. So I just want to point out a few GDPR principles here. The first one is consent and agreement. So if I put my data on a platform like LinkedIn, I use it very heavily. So if I use the data on LinkedIn, I agree that the data is used for this algorithm on LinkedIn. But I have never agreed that this data is used by external parties, for example, to make analysis about my behavior and maybe use it for the Gotham software at the end. I have never agreed on it. But what is happening with Gotham, they do exactly what is available on the Internet and also on LinkedIn. So I think they can just take this kind of data, put it in their analysis and use it without asking my permission. And at the end, I can ask, okay, what are you doing with my data? But still, I mean, it doesn't matter for the Gotham software. They still take the data without asking for minimization. And at the end, these data or these kind of systems, they need a lot of kind of data. And this is also against another principle, the data minimization. So if you really want to use the Gotham software as it is planned, then it is a very high-risk AI system because it does not follow GDPR. The point is, again, here that GDPR is not really relevant for suspects. So you can use it or you can just analyze any information of people without asking of permission if you're a suspect. However, if I just call the police, then I'm as well in the analysis. And this is against GDPR then. And therefore, these kind of systems are also listed under the high-risk AI system environment of the AI Act. And you have to do some procedures for high-risk AI systems. So you have to do risk management, you have to check if everything is transparent, you have to do data governance, and many more. And at the moment, the whole authority to check it, if these things are fulfilled, are not in place. And this is why we have Gotham. Just simple as this. And, yeah, last but not least, many people say, okay, the AI Act is just really a blocker of innovation. But at the end, for me, it is not a blocker of innovation. It just secures my human rights at the end. It just tells that we are all equal and that we should all be treated equally. But if we use software like this, it's not really possible. And therefore, we have the AI Act. And hopefully, I really hope, that in the future it will go a little bit different and we all see that this kind of software is not really a good software to use and it's against our fundamental rights. I mean, I just have this source from December 2025. Switzerland rejected to use Palantir. The reason here is that they checked the whole Palantir software, Gotham, and wanted to use it in the military. And then they saw, okay, the maintenance is an issue. And also the codes are not really clear. And it is just a black box, so they don't want to use it in military. And I really wonder why Switzerland is rejecting this kind of software and we are still using it. And it's still under discussion. And now I think people said, okay, maybe we will not use Palantir in the future, but maybe something similar. Then you're more like sovereign, but it is still against GDPR if you use this kind of software. So that's why I always tell my customers, please use one of the open source models. If you want to use like a large language model, just use one of the open source models like Toykin or Apertus. It's good for many of the use cases that we need in the legal space, for example. And we don't need something super fancy like ChatGPT, we need something for classification, information extraction, and then here these models are pretty good. And what we have built for our customers, we created such a playground where you can run Apertus against ChatGPT 4.0 Mini, for example, and you can check what the output is of both models for a certain question, and if you already get what you wanted to see. And if the models like Apertus is good enough for your own project later our use case later. You can also check or compare a pathos against Troiken in our setting and what you see that it's working pretty well in the English language however we still have issues with German language to be honest. You see here I ask in German what indigenous sovereignty means and you see here the whole answer is in German but here you have like a switch into English I think here but I think it is something we still could fine tune and still get over these kind of difficulties so but what I tell to my customers is always if you use this kind of open source software use it in your location use it like on a service provider here from Germany we are very secure and at the end we really have like digital sovereignty for your AI project and it's not just a password it is really secure and good to use Thank you everyone for listening and I hope you have a few questions.
Speaker 1 [20:48]
Thanks very much for the exciting talk. And we do indeed have a couple of questions. So I will start with the one that came through first, which is, does the CLOUD Act make it impossible for all of us, for all us companies to be GDPR compliant?
Speaker 2 [21:05]
Yes. Yes. Okay. I mean, if you don't put, like, a security layer around, I mean, you could, like, do different things to protect your data if it is on AWS, but, yeah, if you do not protect your data on AWS, then you shouldn't use it.
Speaker 1 [21:25]
Great. Very clear answer. Thank you. Do you think it is realistic for most companies to get to digital sovereignty in the next five years?
Speaker 2 [21:38]
I mean, we had it in the keynote today. I mean, it depends on you as well. If you set up everything on AWS and say it's a good project, now we can scale, we will never reach the phase of being very secure in our place and be independent at the end. I mean, I see it as my work to tell customers we can use AWS, but we shouldn't. And we can set up everything fast with AWS or like with the Azure AI, we can do everything. But maybe you should look a little bit further and the dependence we have, like with the political dependence we have and also about resources that we use. And this is also why I really like working for the company because churches are really aware of these topics and they want to make good decisions. And it's not about scaling. It's about making the right decision, the most ethical decision. So it's up to you.
Speaker 1 [22:38]
Great, so that's a message to everyone. Start working.
Speaker 2 [22:42]
I
Speaker 1 [22:44]
we remain economically relevant if we only use the smaller, weaker free models that are GDPR compliant?
Speaker 2 [22:52]
But this is what I mean. I mean, for most of the use cases that we have, we don't need like the big models. I mean, they are just taking too many resources and they cost a lot of money. And I don't see a point really using a big model if I can make like a categorization also with a smaller model. I mean, if a big model is really required, go for it. Yeah, I mean, use it. But most of the cases for customers, They don't even need big models. And it's really like they don't care if you have like, I don't know, GPT-5 or GPT-6 or 7. And they don't care about it. They want to have a running use case. That's it. If you do it with a small or bigger model, if you can make it run.
Speaker 1 [23:42]
And I think we've got a question that follows on from that very nicely, which is how to best convince decision makers that sovereignty is not a blocker for innovation.
Speaker 2 [23:52]
That's a very good question. I have the same. I wanted to ask the same in the morning because I really have no idea at the end how to convince them. It is super difficult because everyone wants to have it like yesterday, the project. And this is super easy to do with AWS settings or like with the Azure AI stuff. But I mean, I just try to repeat myself. And I was also wondering if I just like put a stamp on all of my slides. we are sovereign, even though it is not really true, but I don't know. I think for me, I try to be honest and tell them, you know, it might be like a big investment today, but you invest in the right thing for the future.
Speaker 1 [24:39]
Yeah, I think that's probably true. We have a final question, sorry. I've dropped my phone, apologies. So this actually moving away from the companies and the institutions that you mentioned, it's about us individual users and what do you think of digital sovereignty for the individual, self-hosting, smaller online communities, what are your thoughts on that?
Speaker 2 [25:03]
That's a very good question. So my phone is like six years old. I still use a very old phone. And I don't have a ChatGPT account. I don't have an Amazon device. I'm not using Amazon anymore. And I think it is really up to you what you want to support in the future as well. And maybe I'm like the slowest person in writing abstracts abstracts because I don't know because I'm just aware of all the security incidents we had in the past and I'm also not using like generative AI about in this in the way that just write an abstract and then I send it out I really use it carefully and I think it's up to us who are making the decision what we use in the future and who we support and I really try I mean I use instead of Amazon I use Otto for example it's a German company and it works the same seriously it works completely the same it also has a marketplace and yeah I hope this answers
Speaker 1 [26:10]
great talk, great answers yes, I'd like everyone to just thank our speaker and thank you
Speaker 2 [26:16]
And thank you very much.