Chasing the Dark Universe with Euclid and Python: Unveiling the Secrets of the Cosmos Keynote

The Euclid mission, a European Space Agency-led mission launched in July 2023, is set to transform our understanding of the Universe by exploring its most elusive constituents: dark energy and dark matter. Together, they account for 95% of the cosmos, dictating its structure, evolution, and eventual fate. Euclid is currently surveying one-third of the sky to construct the most extensive 3D map of the Universe ever created. By using deep imaging and spectroscopic data, it traces the distribution of galaxies and the subtle distortions caused by gravitational lensing with unparalleled precision.

By connecting theory with observations, Euclid aims to uncover the properties of dark energy driving cosmic acceleration and the distribution of dark matter shaping large-scale cosmic structures. At the heart of this endeavor lies the challenge of cosmological statistical inference: extracting robust conclusions about the nature of dark energy and dark matter from vast, complex datasets. This talk will explore how cutting-edge statistical techniques and powerful computational tools, including Python-based analysis pipelines, are being used to compare theoretical models against Euclid's observations. We will discuss the role of Bayesian inference, machine learning, and advanced simulations in constraining cosmological parameters and testing extensions to the standard model of cosmology.

This session took place in track Keynote and was classified suitable for novice domain / novice python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:07]

Thank you so much for the lovely introduction and the invitation to be here. I'm extraordinarily delighted. This is the first time that I'm speaking to a particular type of public, that it is you, like moving towards the software. Usually I'm either on fully scientific cosmological community or the lay citizen, the general public, to tell them what we do at the European Space Agency. So let me know by the end of the talk if you actually like what we do or the challenges that we are facing at the European Space Agency, and hopefully you will understand a little bit more how we are using Python to make this alive. I know that it's really early in the morning, so I just would like to introduce you to what has been my life so far, and that is it, the European Space Agency Euclid mission, and I have a video prepared for you.

Speaker 2 [01:21]

In 1915, Albert Einstein astonished the world with his general theory of relativity. It described the behaviour of the entire universe based on the matter and energy contained within it. The theory sparked the modern discipline of cosmology and the hope that we would finally understand how the universe came to be. But in recent times, the effort to define what the universe is made of has given us a very big surprise. Visible stars and galaxies make up less than 5% of the universe's total matter and energy. Beneath this visible layer is a mysterious celestial realm, consisting of shadowy particles and unknown energy fields. For decades, astronomers have puzzled at their nature, calling these elusive substances dark matter and dark energy. ESA's Euclid mission will go in search of the answer to the fundamental question, what is the universe made of? A European designed mission, Euclid is built and operated by ESA, with contributions from the International Euclid Consortium and NASA. ESA selected Thales Alenia Space to lead on building Euclid, with Airbus Defence and Space providing the telescope and payload module. The telescope and scientific instruments form the heart of the mission. Together, they will observe billions of galaxies over more than one third of the sky. Producing record quantities of data, Euclid will enable scientists to draw a precise map of the Universe across space and time. This will allow researchers to investigate the effects of dark matter and dark energy on the apparent shape of galaxies and on their motion and distribution over immense distances. In turn, this will help reveal the true nature of dark matter and dark energy. The spacecraft and data communications will be controlled from ESA's European Space Operations Centre in Darmstadt. To cope with the vast amounts of data Euclid will acquire, ESA's ESTRAC network of deep space antennas has been upgraded. These data will be analysed by the Euclid Consortium, a group of more than 2,000 scientists from more than 300 institutes across Europe, US, Canada and Japan. Understanding the elusive nature of the universe has drawn astronomers throughout history. It remains one of the most challenging investigations in modern science. But Euclid is up to the task. The Euclid mission is a quest into the unknown. A mission to shine a light on the dark side of the universe.

Speaker 1 [04:15]

I actually don't know how many times I've seen this video. This was the first time that I actually was actively helping the ESA communications teams right after I joined the European Space Agency, which is where I work on. Although I'm a full-time researcher, I dedicate part of my time in communicating what we do. And the reason behind it is that we actually need to tell you what we are doing. The European Space Agency is an organization based on different member states, based on public money from these member states. We have different missions that are exploring the solar system, different planets, but Euclid is about cosmology. It's about actually understanding what is the history of our universe, how did it begin, how it began, evolved, and what its eventual fate could be. Actually, we do not put a millionaire mission such as Euclid and a spacecraft into space without understanding and having some clear idea of what we would like to do. And for that, we are building up. Cosmology is actually really recent and really modern in its essence. So we have been building our knowledge about the universe over the last 30 years. And there were some observations that were crucial, that were the cosmic microwave background. So for you to understand, we're the first light that actually we could detect of the beginning of the universe, what we meant, the Big Bang. We believe that our universe started in a really hot and dense state, and right after this Big Bang, that we speculate how it could actually be produced, our universe started evolving. The first nuclei were formed, the first neutral atoms were formed, up to some point, light could travel all over the universe. which is what we have actually measured from space with the European Space Agency PLAC spacecraft. Our universe continues expansion and up to some point, gravity has started to play a role. And the thing is that our universe is quite simple. Assuming the theory of general relativity of Father Einstein, we could actually see or we could predict how the first stars were formed and then how these stars are starting clustering together to form the first galaxies. These days, we know that our universe, when you look at it from very large distances, it has the form of what we call a cosmic web. So these galaxies tend to form together different filaments that resemble a web, therefore the name. And this is what we are trying to measure with Euclid. Why are we interested in detecting and measuring and modeling and mapping how our universe look like a very large scale? Well, because we know that our universe is expanding, but recently, since the last decade, we know that this expansion is actually accelerating. And we don't know what it is behind this acceleration, and because cosmologists are really poor in marketing, we call it dark energy. Some kind of a nonsense that we don't know what it is that potentially is fueling this accelerated expansion of the universe. Because we know that this accelerated expansion of our universe started quite recently in the history of our universe, and it has might have affected how a structure have formed at this stage. This is why we are aiming to map this part of the universe with Euclid. Ultimate, what do we want to do? This is all about statistics and statistical fitting. In cosmology, we have a model, a really simple model that we call the standard cosmological model. with really few parameters. Take into account those models that they are used, for instance, for weather forecast or in vital medicine. We usually have many parameters, whereas in the standard cosmological model, even though we believe that our universe is really complicated, with really few parameters, that basically explain how much matter we have, how much ordinary matter we have, that I will tell you about in a minute. Some initial conditions, this expansion of the universe, the geometry of our universe, and how much of this dark energy we have, we could actually predict what could potentially happen in the future or what it could have actually happened in the past. However, this model is what I like to call a cosmic embarrassment. Because if you think about it, we have a really robust model, statistically speaking. However, we do know nothing about basically what this model tells us about. Because when you fit this model against the data, what you see is that this ordinary matter, so the matter that stars, galaxies, everything that emits light, us, ourselves, are made of. This is what we call baryonic matter, and only 5% of our universe is in this form. We know that there exists some other type of matter, that this time, we were a little bit better with the marketing, we call it dark matter, because we cannot see directly, it doesn't interact with light, but I will tell you later why we know that it must be there, compose around 25% of this composition. And the rest, up to 70%, is this dark energy. So we have a fantastic model that actually, we pretty much, physically speaking, we only know 5% of. This is really, really bad. Now, if we would actually like to say something about the nature of dark energy and dark matter, you need to feed different models. So as I said, this is all about the statistics. I would love to tell you that actually, in cosmology, we are doing very fancy things. And yes, you will hear some of the fancy things by the end of the talk. However, most of the knowledge that we have built up on cosmology these days are based on Bayes' theorem. So I have already seen this formula yesterday in some of the talks. But the key here to do the statistical fitting is to obtain the statistical probability distributions of these parameters that I was showing you before. So the amount of matter, the expansion rate of the universe, how much dark energy we have against some kind of data. And for that, as I said, we rely on our beautiful Beast Theorem. And there is a key here that makes the connection between our theoretical predictions against the data. And that is the likelihood. So this is statistical probability distribution that, as I said, in cosmology we do not do many fancy things yet. We just model it to be a Gaussian probability distribution. So that's it. We basically predict that the logarithm of the likelihood is proportional to the difference between that data that we are measuring with respect to the theoretical predictions that we have, and, of course, we weight with the uncertainties that we have in our measurements. The better the measurements, in principle, the easiest is going to be to have a robust calculation of the likelihood and, therefore, sample that probability distribution called posterior to obtain the best fit of the model. Okay. So, this yet is not rocket science. Data, Euclid, theory, everything that goes beyond this standard cosmological model, that was equations, you guys are not going to see in this talk, but I will basically tell you a little bit about that. So now the question, what do I do? Well, this is my role within Euclid. What I'm attached to is to calculate this likelihood and to make the software that actually calculates this likelihood. So let me reintroduce myself, who I am. So basically, hi, I'm Guada. Apart from working at the European Space Agency, I'm being Euclid, one of the the spokespersons for the agency, so being today here is an honor, but I also had the opportunity as I asked ChatGPT yesterday to tell me a little bit about myself, yeah, well, I have, like, some more media appearances than Netflix actors, that was slightly scary, but anyway, also something that I do for this mission called Euclid is that I'm a manager, so I basically coordinate the effort of which models beyond this standard cosmological model that we have are worth testing against the data. Overall, in reality, I spend my days mostly coding. So as I said, I'm also responsible of the software that calculates the likelihood and not only the likelihood probability distribution, but also how we calculate these theoretical predictions that are based on different cosmological models. And because I don't have tons of patience, I've been working a lot, as you will see by by the end of this talk on how we could use machine learning and artificial intelligence to speed up all this process so we could have results hopefully in this decade and not in centuries from now. So let me see if I could tell you or you could actually grasp the theory of general relativity of Father Einstein with halcyon equations. This is a challenge, but let's see how well I do it because at the end of the day we want to study dark matter that we cannot see. We want to study dark energy, fantastic, really fancy, but the only thing that we know is that potentially it's there fully in this accelerated expansion of the universe, however, we cannot grasp what it is. We only see its effects. So in reality, this is some kind of literally dark magic, right, because we must learn how How we could measure something that we cannot see based on the effects, on how you could actually predict these effects on the different structures that we do take pictures of on this photograph. Okay, so challenge accepted. Now, you need to understand that I come from the academic world, the European Space Agency. Even though, yes, we do have a fantastic set of different contributions from the member states at the end to basically explain really basic concepts that we use on a very daily basis. We use really basic stuff. So now, I would just like you to see this video. This is just a penny that one of my colleagues, this was from Dr. Jason Rhodes, actually working for NASA, created a few weeks ago to be able to explain this phenomenon that we call gravitational lensing. This penny is put on the floor of a bath. And because you actually see the image of this penny distorted, you know that there is water in the bath. Now you cannot tell me how much water you have because you only have one data point, which is this penny. Now imagine that you have, like, many pennies on the path. You could tell me a little bit more based on how the different shapes are formed, basically if the water is standing or maybe there is a tilt. Now imagine that this path is actually the largest swimming pool that you could ever imagine. This is what we are aiming to do with galaxies in the universe, because there is an effect that is really similar to what you are actually seeing with the shape of this penny. The light, that it is coming from very far away galaxies, its path gets distorted by the fact that there is something in between, a galaxy or this dark matter that we cannot And as a result, we actually see the shape of these distance galaxies completely distorted. This is just a diagram. But guys, with Euclid, like eight weeks ago, we actually saw one of these. So this is what we call an Einstein ring, and this is one of the biggest effects that are predicted by Einstein's theory of general relativity. On reality, the shapes of all galaxies get up to some point distorted. So by averaging over all the different distortions of the galaxies, we could actually infer how matter is distributed in the universe. However, we cannot say anything of it is the matter that we know something about, this ordinary matter, this 5%, or that matter. We need something else. But we do have that something else. Now I want you to put yourself in the situation that, yeah, somehow you are in a spacecraft coming from an alien civilization, and you would like to colonize Earth. And this is where you arrived, okay? So you visualize, you do a strategy, and you're basically visualizing the Earth during the night and you will say, hey, I'm going to start colonizing and I'm going to attack first those areas that are emitting more light. Because I assume that basically, more light means that more people are living in it. Well, you will be missing quite a substantial amount of the world population, either in South America, China, or in Africa. This is what we actually see with galaxies. Galaxies emit light, so we do see where that ordinary matter is, but we don't get to see where all that matter is, we cannot detect that matter. However, combining this effect of gravitational lensing, so the distortions on shapes of galaxies, plus actually knowing the two-dimensional position of those galaxies, we could actually discern what is dark matter and what is ordinary matter. Fantastic. Like two parameters of my model, something that I can measure. Now, we know that our universe is expanding. How do we measure an expansion? In a very similar way that when we are driving our car, and unfortunately we get a speed fine because we overrun the speed that it was allowed to be used in that road. This is a real image from Euclid. Here you can actually see many, many galaxies, some of them are obvious, but some others are not. Like for instance, this little one here is also a galaxy. Because nothing can go faster than the speed of light, this is something that Einstein tells us in the theory of general relativity. We could classify those galaxies according to their distance, and very far away distance galaxies means that that light took more time to travel towards us. Therefore, if we classify, it's like taking different snapshots of how the universe look like at different times. So we do something like this. We classify and then we have like observations at different times. If we study how this evolution has changed over time, voila, you have the expansion rate of the universe. Okay, so in principle, what we need is quite simple. To do statistics, we basically need positions, shapes, and distances to galaxies. Now we are in the era of big data, and that also applies to astronomy. How many galaxies we actually need to be able to discern statistically between this standard cosmological model or other models? Well, like many. Many means many millions. So actually this is the reason why we need Euclid. And this is what you're going to be seeing now. Euclid is so weird to be talking about a mission that I dedicated basically the last decade of my life preparing to start talking about it in present time. We launched on the first of July of 2023 from Cape Canaveral on a Falcon 9 from SpaceX. It was the first scientific partnership that the European Space Agency had with SpaceX. It was a really smooth launch. I was not in Cape Canaveral in Florida. I was actually here in Darmstadt in the operation center because there was something that I I wanted to see, and it was the moment that you have this launch happening, and as I said the launch was fantastic, it was extraordinarily smooth, but one of the keys is when you launch that you actually set free your spacecraft, and you switch it on, and you see if it is see if it is actually alive or something was broken right after you launch it, right? So to get signal from the spacecraft, this is what it is doing really close from here at ESOC here at Darmstadt. And I was here because I really wanted to see that graph, like showing the first peaks that my little Euclid was alive because for many people at the European Space Agency, this was the end of a project. For me and many scientists, this was actually the beginning. It undertook a travel, well, a trip to the most expensive neighborhood in our solar system. And I'm not going to be arrogant because, honestly, I don't know if there are other civilizations out there. Hopefully, yes. But, yeah, probably the most expensive neighborhood that we have in our universe right now. It's a really crucial point that we call Lagrangian Point 2. What does this point or area in our universe has a special on? Okay, so it is a point that basically is between the sand, the Earth, and this area. The sand is basically behind Earth, so it is pulling solar energy towards the spacecraft so we can actually operate. But it protects, somehow, the solar panel protects from further solar contamination, so light contamination, so we can look at the deep universe. Why it is expensive? Because most of the missions that need to look at the very deep universe are operating from there. So, James Webb Space Telescope, Gaia recently, Planck in the past. Euclid Spacecraft is actually quite rudimentary. It's a telescope. So, a very common telescope, 1.2 diameter, what it has as a joy is what happens in what we call the payload. Basically, the cavity within the spacecraft where we have the instruments. I'll tell you a little bit about it right now. What I can tell you is that I had the opportunity to see the spacecraft in person. You do not realize on my face, but I was extraordinarily nervous. I couldn't sleep. So this is actually in France in 2023. And I wanted to compare because, well, you see here that, I don't know, I'm quite small. I'm 150. But it was really big. It was really shocking to see how big actually the spacecraft is. So as I said, the key of Euclid is what it is hosting, the spacecraft. So one thing is the telescope. This telescope allows us to maximize these images that we need to take from the universe. So remember, shapes, positions, and distances to billions of galaxies. But for that, you need instruments. Which instruments? Well, you need a camera to take those pictures, so you can actually measure shapes. And Euclid has right now the largest ever camera sent to space, that it is operating from space. And you need a really important device that measures the distances to galaxies. So this is, we do it with photometry and spectroscopy, for those who have that background. And it currently takes measurements of distances to galaxies of really billions. Every night. It's absolutely crazy. This is how they look like. So if you actually look at the camera that you have on your iPhone or in your phone overall, This is just the largest version, six times six, and in the case of the device that measures distances, it's four times four. So this is Adelaide, provided by the UK and France, provided by France, Italy with German contributions, actually from very close from here, Heidelberg. This is the kind of images that you take from the ground of the Earth. So you might wonder, we know that we have telescopes. I mean, I'm Spanish. We have a beautiful telescope operating in the Canary Islands. We have a beautiful telescope operating in Chile. What is the point of sending a telescope to space? It's not just to create more beautiful pictures. It's actually to create sharper pictures. This is the same field observed with Euclid. Where do you prefer to measure shapes of galaxies? Here or here? This is the reason why we went to space. Because we don't want to be contaminated by the atmosphere. And by the way, this is just a raw image. So you see here like some straight line, this is contaminations. Here you have some stars. These guys are galaxies. And you are going to see right now how much we can actually submit on that. You already saw in the video, I'm here representing the European Space Agency, but I'm also actively working within the scientific collaboration behind the scientific explotations and the the scientific collaboration that it was in charge of providing those instruments to Euclid. More than 2,700 scientists spread all over the world, and it is a pleasure to work with all of them. This is a picture that we took like three weeks ago, actually in Leiden, in the Netherlands. I was one of the organizers. I was right here. So this is just a snapshot of some people that could join the meeting in person, 600 people. And it was really funny to see that our Euclid model didn't actually fit in the picture. We were so many that it was completely covered by everyone. Guys, we are already operating. But it's not that we are actually operating. It's that we released already data for you to play, even if you would like to. Even if you don't have, like, astronomical background, you will see later how. But I would like to take you over some of the fantastic pieces of science that we are doing with Euclid in this first quick data release that we have on 19th of March. This year So this is only 63 square degrees You will see later how much Euclid Is going to measure during his life So what we are doing in this video is zooming in and zooming in, so it's not only about the field of view of Euclid, but also the Shermans. This is an area where stars are forming. An area that we used for calibrations of our instruments. It's crazy to see that you have, like, some galaxies, like, they are so well. This is one of the patches that we have in the Northern Hemisphere. We also release patches in the Southern Hemisphere. The reason why we are working in the Southern Hemisphere right now is because we also are incorporating the data that we have from these telescopes that are operating from the Earth. Even though the data might be less quality, Everything that we've learned on that data is really, really values to basically minimize the statistics as well. So see, some lenses. The universe is basically clouded with this kind of objects, more than a million galaxies in this beautiful cluster of galaxies. Some stars, you will recognize if a picture is taken by Euclid because you will see like a really particular deflection pattern of the estuaries of six different peaks. This data is already public for you to play with. And even if you might say, hey, I don't have an Astronomy Cup background, doesn't matter. You will see now how you can even actually open the files. So first rule, just make me a favor. What you're seeing here is in really low quality. Compare what we can actually offer to you. Go to ESA Sky, which is our platform when we are putting the 24 times 24K images and just zooming and zooming and zooming and get wonder by the universe that we live in. There is another option. If you would like to play with data, the European Space Agency have made an extraordinary effort to finally have what we call our data labs. So basically servers where you can do your sciencing. You can basically do your statistics on. You don't need to download the data anymore. And if you would like to actually see or say thank you to some of the people who are actively working on that, this is my colleague, Sandor Kruk, who has provided like extraordinary support to do all the science that you were seeing in this video before over data lapse. Which kind of science? I was telling you before, it's about galaxies. But we care about some kind of galaxies. And therefore, for that, we need to classify. Classify, guys. We cannot do it by eye. We do rely on citizen science. So some of the classifications of the shapes of these galaxies have done with what we call the Galaxy Zoo, where we have volunteers like classifying those galaxies that we're feeding some training data to actually a neural network that was helping to classify. We did this exercise of actually classifying galaxies as well for these strong lenses, these things, these objects that we actually need to understand how matter is distributed in the universe. And the good thing is that, well, the good thing, the fantastic thing about this catalogue is that it discovers so many objects that we didn't see before. So, Euclid is not only valuable for cosmology, as it is taking so much data. You guys know how much data it's taking. Like, at the end of operations, we are talking about petabytes of data. This is why we needed to improve the capacity that we had at the European Space Agency, not only to host, but also to move all that among us amount of data. And this exercise of actually setting up our archives and making everything connected of data labs have managed to help scientists like me to be able to do this exercise of constructing catalogs so easily. Why I'm telling you how much data Euclid is taking? Well, some telescope that is really well-known in the pop culture is the Hubble Space Telescope. This is every time the Hubble takes a picture, this is how much it can photograph. This is the field of view. Every time that Euclid takes a picture, this is what we measure. So what does it mean? Well, that basically in the amount of lifetime that Euclid has been operating, that it has been like around one year and a half, we have already collected as much data as Hubble in, Well this year it is their 35th anniversary of operations. So this could give you like an indication of how much data we are talking about. And you can actually see online how much data we are retrieving. Like you can go into this web page, there is a really simplistic Python script that it is running behind the scenes, that it is retrieving how much data we are taking overall. So you have seen what we can offer with 63 square degrees. That is nothing. To do cosmology, or to start thinking about doing cosmology at these big data statistics, we need at least 2,000 square degrees, or the order of 1,000 square degrees. This is what we are going to have in public next year on the 21st of October, 2026. And it is 2,000 square degrees. And what I would like you to see is those first 500 square degrees. So if you got amazed by 63, let me show you what we can actually do with 500. Now 500 is one-fourth of the data that we are going to be using for cosmology. So here what you can see is just a projection of the sky that is seen by Euclid. This is contamination in the infrared. We are looking for galaxies that can be seen both in the visible and in the infrared, and We're basically looking for galaxies that are not contaminated by the light of our own Milky Way. So all of that. Now let me show you this 500 square degrees. This data was taken by Euclid only in two weeks. Like it's amazing. Like to build this with HAVL or with James Webb, it would have taken decades. So this is real data. Actually some of these areas that you are seeing here are unfortunately some data that have to discard because we have contaminations from solar bursts. We are actually taking data in the cycle of maximum activity of the sun. You have seen it, northern lights. Let me zoom in. So it is not about all the galaxies that we are aiming to detect, but also about the amount of data that it is extremely valuable for legacy astronomers that would like to understand things like galaxy evolution, as we were seeing before, classification of galaxies, how galaxies are forming, how stars are forming. Let me just zoom in a little bit more. These are two galaxies that are interacting with each other. Those are spiral galaxies that are interacting with each other. And the good thing is that you can actually resolve the stars on the arms, and you can actually see galaxies which are far in the background. So we would like to measure shapes of these galaxies. We are measuring shapes of these galaxies. So this is 500 square degrees. And this is what we do data with. Hopefully all these products will be available for everyone to play with October next year. But for that, we need to make an exercise of understanding all the products that you need. It's about taking the pictures, making sure that the pictures are properly saved with all the metadata. Also, we need to start building those catalogs. What is a start? What is a galaxy? Its shapes, its distance, its position in 2D, how it was calibrated, which all the data you used to calibrate, all of that. Well, we are making an exercise, well, actually, my colleague, Sam Farrance, a fantastic Python mentor that I have, was working really, really hard to make sure that we have everything properly classified and properly stored in what we call our data product description document, actually running in Python. Hopefully the software that was used to create all these products, all these catalogs, will be stored if we get the seal of approval that it is really likely that it will happen on GitHub in the Euclid Consortium organization. I wish that I could do cosmology just by looking at this picture. The reality is that I can't. I need to compress all this petabyte of data on something that I can compare these equations that you haven't seen against. So what we do is that we build something called summary statistics, or as we like to call it in astronomy, power spectra. So what does it mean? That we start taking a galaxy that we have classified and we see its correlation with another galaxy. And then we repeat the exercise with a galaxy and another galaxy at a farther distance. And then another one to a farther distance. And this is where you construct power spectra. So you see what is the correlation of different galaxies at different distances. The code that does, that calculates this power spectra on the data for Euclid is public. It's called Herakles, and it has been developed by one of the best colleagues that I have in the Euclid Consortium, Nicolas Tesori. And it basically computes, it compresses all that amount of data into a few kilobytes of data that I can use to compare different cosmological models with. Now, you might think that it was obvious that we would need to read that data to do patient statistics. It would have been lovely to make sure that we have something prepared for that. The reality is that no. When I started working at the Euclid Consortium, I quickly realized that we didn't have anything prepared, anything homogenized to read all these data products. So this is something that I partnered, and this is how I got to meet Dr. Nicholas Tesore about basically making sure that we could read all this amount of data. The package is public. It's called Euclid Lib. We are actively looking for contributors from the community to make sure that we could incorporate like all the reading routines that we need for this. Now in the last five minutes of my talk, I would like to bring you to what has been my work over the last five years. As I said, this has been calculating the likelihood distribution, given some theory and Euclid data. This is done with the cosmology likelihood for observables in Euclid or Chloé, fully written in Python, and what it does is that it calculates, it samples this probability distribution for the cosmological parameters, and in particular, as I said, we are interested in dark energy. So basically, this is the plot that potentially might tell you nothing, but this is the most important plot that we will ever produce in the Euclid Consortium or in cosmology talks, or in cosmology experiment. So basically, how well, how tight the uncertainty on the dark energy parameters might be looking like. So I think that it is important as well to talk a little bit about failures. I worked on this code for five years, and it was only used by the people who wrote it. That's me and a few other collaborators. The reality is that we quickly realized that we have to reinvent it, because it became a black box. We never took into account the overall user, and it was a really good lesson learned, because good science really needs good engineering and really good software. So what do you do? You lame your injuries and you restart. But this time, what we have restarted was analyzing what is the current state of the art and what is happening overall in data science and in the Python community so that it incorporating machine learning and artificial intelligence algorithms to make sure that we could tackle an issue that we have these days. There is a reason why I'm not showing you the equations. First of all, because they are really ugly and really hard to understand. I hate them. because they are really pricey and computationally expensive to compute. And we need to compute them and solve them every time that we would like to sample a point in the posterior distribution. So it is extraordinarily annoying. This is a bottleneck in our analysis. Like obtaining the plot that I was telling you before that it was one of the most important plots that you will ever have in cosmology takes approximately 20 days. We don't have 20 days to analyze the data that it is going to be public next year in October. Therefore, we needed to reinvent ourselves. This afternoon, I'm going to be part of the bi-lady's panel on artificial intelligence. The reason why? Because I've been basically building different neural networks that learned the results of these equations of general relativity so that we could make computations and predictions extremely quickly. So I will tell you a little bit about, within this panel, I will share my concerns of how well we need train these algorithms to be able to use it for this high-tech science. For the first time, an experiment studying large-scale structure of the universe has been written completely in auto-differentiable framework using JAX, and this has granted us to be able to use gradient-based samplers. So it's not only about making the predictions quicker, but making the sampler even quicker. The community is public, the repositories will be public really soon, actually next month, and this is a fantastic moment. All the people that I have basically have been working really hard with me on developing this, and the good thing is that, well, since we went open science within the U-Click Consortium, finally we have been receiving feedback from our own people. So in the time that we were working on the previous code, zero bugs reported. Since we actually open it and we follow open science, which is not only about making it public but changing the mentality of how people should be thinking about visibility and accountability, we have feedback every week. This is something that I'm extraordinarily happy from. So in total, how much data we are going to analyze? 14,000 square degrees in six years from now. This is when we will be able to say something about that energy. Take-home message. Well, it has been ten years since I've been playing with Python. It has been an honor for me to be here with you today because I feel it like a full cycle moment. Hopefully, you will get to see these codes all in public and maybe get to inspire you this kind of Python algorithms that we are using to solve one of the most or biggest mysteries in the universe right now. And it has been a pleasure to do this with friends, colleagues, and mentors. So thank you so much for listening.

Speaker 3 [42:31]

Thank you so much. So who can relate to the refactoring part?

Speaker 1 [42:44]

I don't know.

Speaker 3 [42:45]

So, I think it's really great to see we all struggle with similar problems and we do conferences like this to talk about it and to find solutions. So, we have time for two or three questions. So, what's the most malleable lesson you learned when embracing the role of science communication?

Speaker 1 [43:04]

Well, I think that, yeah, the biggest lesson is that scientists or cosmologists, in this case, that it is my field, we tend to specialize so much that we get to talk only with people that are doing what we do, what I do, every single day. And we lost a little bit of the horizon that at the end, we are doing this not only for ourselves. I mean, we would like to contribute to understanding more how the universe works, but also to tell and to share it with the society. So for me, it has been like putting a reality in front of me of basically being sure that you understand that the effort that we do in basic science, that it is this case of cosmology, it has some impact in society as well.

Speaker 3 [43:52]

Very technical question. What is the resolution? E means the max zoom. How many square degrees does this single pixel cover? And how does it compare to Hubble and James Webb?

Speaker 1 [44:05]

Okay, so in terms of picture sharpness, Hubble and Euclid are really, really similar. Now, for the field of view, it's around this size. So every snapshot that we take is about the size of the moon on a very light and clear night. So whoever has that question, I could tell you, I could point you later exactly to the resolutions of the numbers.

Speaker 3 [44:33]

Last question. Are people manually classifying the galaxy images consistent classification noise? I'm just reading it. You compared images from Hubble. How does it... Okay, we covered it. So are people manually classifying the galaxy images?

Speaker 1 [44:49]

Yeah, so first you need to teach people what it means classifying a galaxy. So the algorithm is trained in such a way that first when you log in, you basically have like some set of galaxies that are really obvious to be classified. That's it. If it looks like a spiral or it doesn't. If the user has classified correctly, then it shows some other galaxies, say the pictures that are slightly more complicated but that still you could catch by eye. Now if the user continues training themselves, there will be a point in which we will take that data to be properly training data later for the neural network, but there is always some uncertainty, so we do ask the user, hey, if you are not really sure, please just skip it and don't give me that data. Just be honest with yourself.

Speaker 3 [45:37]

Thank you. So Guada is around all day. Feel free to approach her, discuss. I was also happy I didn't know you work with Sam because he gave a keynote at Euro Sci-Pi a few years ago. It was also like great So you see everything's connected. We share the same problems Let's stick our heads together and solve them together and thank you very much

Speaker 1 [45:56]

Yeah, thank you so much for listening, guys.

Guadalupe Canas Herrera

Guadalupe is a Theoretical Cosmologist working in understanding how the Universe began, how it evolved and what its ultimate fate could be. In particular, she is interested in studying alternative cosmological models with state-of-the-art astrophysical data using advanced statistical techniques and data science algorithms. Furthermore, she is interested in forecasting the performance of new experiments or new observables, for instance, Gravitational Waves.

She holds a Bachelor's in Physics from the University of Cantabria, and Master's and PhD degrees in Cosmology from Leiden University. Currently, she is a Research Fellow in Space Science at the European Space Agency. Moreover, she is an active member of the Euclid Consortium: the scientific group behind the data explotaition of the ESA Euclid mission. In particular, she is the maintainer of the code "Cosmology Likelihood for Observables in Euclid" or simply, CLOE. This software is part of the official data anlysics pipeline that will be eventually used to extract cosmological constraints of the Euclid data. Within the consortium, she is also co-leading the responsible group in charge of testing models beyond-Standard Cosmological Models to discernish the nature of Dark Matter or Dark Energy, or to test alternative inflationary models.

Social card for talk: Chasing the Dark Universe with Euclid and Python: Unveiling the Secrets of the Cosmos