When Space Weather Breaks Your GPS: Building an Explainable Early Warning System
Space weather, driven by solar activity such as flares and coronal mass ejections, disrupts the ionosphere and causes plasma density fluctuations known as Large-Scale Traveling Ionospheric Disturbances (LSTIDs). These disturbances bend and delay radio signals, leading to significant positioning and timing errors in Global Navigation Satellite Systems (GNSS) like GPS and Galileo. Such events pose systemic risks to critical infrastructure, with potential socio-economic damages estimated at 15 billion euros for a single extreme event in Europe.
To mitigate these risks, an early warning system was developed to predict the onset of LSTIDs over the European sector within a three-hour window. The problem is framed as a multivariate time series binary classification task using a dataset of 1,600 manually labeled events spanning nine years. The model utilizes physical drivers as inputs, including ionospheric status, geomagnetic currents, solar wind proxies, and solar cycle data such as sunspot counts. Feature engineering involves moving averages, exponential moving averages, and lagged features covering up to six hours of historical data.
The technical stack employs CatBoost for gradient boosting on symmetric decision trees, Optuna for hyperparameter optimization, and MLflow for experiment tracking. To ensure the system is trustworthy and explainable, SHAP (SHapley Additive exPlanations) is used to decompose predictions into feature contributions, allowing domain experts to validate the model's physical reasoning. To handle uncertainty, the system implements conformal prediction, a post-hoc statistical framework that transforms point predictions into mathematically guaranteed prediction intervals. This allows users to choose between three operating modes—high precision, high sensitivity, or balanced—depending on whether the cost of false positives or false negatives is higher for their specific application.
This description was generated by Open-Source AI using the transcript of the session and the original submission contents.
This session took place in track Machine Learning & Deep Learning & Statistics and was classified suitable for intermediate domain / intermediate python by the speaker.
Submission
The proposal as submitted by the speaker before the conference.
Space Weather doesn’t just produce beautiful auroras: it can silently disrupt navigation systems, radio links, and satellite-based technologies we rely on every day.
Travelling Ionospheric Disturbances (TIDs) are wave-like structures in the ionosphere that affect GNSS accuracy and HF communications. From an ML perspective, forecasting TIDs is a challenging rare-event prediction problem involving imbalanced data and heterogeneous physical inputs.
In this talk, I will present an operational machine learning approach developed within the T-FORS project to forecast TID occurrence over Europe. The model is built using CatBoost and integrates data from space- and ground-based observations.
The talk focuses on model design and evaluation choices. In particular, I will show how SHAP can be used to debug model behaviour, validate feature relevance, and build trust in predictions in a high-risk operational context.
Along the way, I’ll share practical engineering lessons on:
- handling class imbalance,
- incorporating domain knowledge into ML pipelines,
- producing uncertainty-aware outputs via Conformal Prediction, and
- running interpretable models in real-time forecasting systems.
The talk is aimed at data scientists and ML practitioners interested in applied forecasting, interpretable models, uncertainty quantification and ML at the boundary between data and physics.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:05]
So, I'm going to show you how it works.
Speaker 2 [00:09]
a funny experience.
Speaker 1 [00:10]
And now we give a warm welcome to Vincenzo Vetriglia, that is presenting When Space Weather Breaks Your GPS, Building an Explanable Early Warning System at the Palladium Room. The stage is yours.
Speaker 2 [00:31]
Thank you for being here. I'm very happy to be back at PICOM Germany. And today I will tell you about space weather and how we applied AI to build an explainable early warning system. Before starting, a few words about me. I have a background in theoretical physics and I work as a machine learning engineer and data scientist in a research institute in Italy, which is the National Institute of Geophysics and Volcanology, where we study volcanoes, earthquakes, but also the upper atmosphere, the environment in general, and space weather. I also love community, and actually that's why we are here today. I'm one of the organizers of the local chapter of Paidat in Rome, So if you happen to be in Rome, feel free to join us. We are always very happy to meet new friends. And this year, I'm going to be on the road to deliver some talks around Python conferences around Europe. So that's the agenda for today. We will start with space weather. And I don't know if everyone is familiar with this term. If everyone ever heard about that, possibly not. So, auroras are possibly the most striking consequences of space weather. But that's definitely not all the story. And indeed, before diving into the machine learning architecture, I would like to set the stage. So, first, space weather. Just like we have weather on Earth, we have weather in space. And this is the physical state of the near-Earth environment as it is driven by the Sun's activity. You know, the Sun is a variable star. It has a cycle which on average lasts 11 years. So when the sun throws a tantrum, like a solar flare or solar wind, this interacts with the planetary magnetosphere. And these interactions can have significant impacts on satellite operations, communication, power systems, GNSS accuracy. And that brings us to GNSS, which is a collective name for different constellations, like the GPS from USA, Galileo for us Europeans, Elonus in Russia, between China. There are also some regional navigational systems, but we are focused on the GNSS side. We use GNSS every day for everything from maps to also timestamping financial transactions. And it works by calculating incredibly precise travel times of radio signals from satellites down to Earth. But those radio signals do not travel through a vacuum. they have to pass through the ionosphere which is a layer of the upper atmosphere extending roughly from 60 kilometers up to a thousand kilometers and this is filled with partially ionized plasma. This region plays a crucial role in radio wave propagation as it bends and delays the radio signals acting like a a giant mirror up in the sky allowing for beyond the horizon communications so when this density fluctuates you can get positioning errors and that's why we should care about space weather this infographic actually perfectly maps out the domino effect of a solar storm So this is not just about 3D auroras, but a direct threat to everything from satellites in orbit down to power grids and navigation systems. And when insurance giants like Lloyd's or even the European Commission publish risk reports on this matter, you know that this is a systematic vulnerability. A study from ESA, the European Space Agency, estimated that a single extreme space weather event could cause around 15 billion euros in socio-economic damages in Europe alone. And that's because our mother infrastructures rely a lot on space-based systems. And since obviously we can't turn off the sun, our only defense is anticipation. And a few words about large-scale traveling and atmospheric disturbances, which are the main topic of this talk. This is a space weather effect of the upper atmosphere. Those are plasma density fluctuations that are rippling through the ionosphere. And they are usually associated with auroral and geomagnetic activity. And they have real work consequences, since they can disrupt high-frequency communications and also genesis positioning and timing. So the physical chain of the mechanisms involved in the formation of LSDIDs is actually clear from a phenomenological point of view. You have something starting at the Sun in the form of coronal mass ejection, which propagates through the solar wind, then you have injection of energy at the higher latitudes which then propagates equatorial as waves and then you detect LSTADs at ground. So how did we build an explainable machine learning model for that? Everything as I said starts at the Sun. So we had to figure out how to devise this task from a machine learning perspective. And despite the clarity of the mechanisms that I told you which are responsible for the formation of TIDs, the real-time monitoring and prediction of this kind of phenomena remains highly complex. So So we frame this problem as a multivariate time series binary classification, which is taking different classes of inputs as physical drivers. So we tell the model something about the status of the ionosphere, something about the geomagnetic state around the Earth, so how currents are flowing around the Earth, some proxies for the forcing from the Sun from above and also we tell the model something about the state of the Sun so where is it as a stage in the in its solar cycle the number of sunspots and so on and so forth and we trade this model based on a catalog of instances which were manually labeled by ionospheric scientists this This catalogue consists of roughly 1,600 events spanning 9 years, covering almost the entire solar cycle. The output of the model is trying to predict if LSTAD is starting or not over the European sector in the next 3 hours. Let's take a step back now. device machine learning model in general you have to pick one either a simpler model or a more complex one and in general when you choose simpler models they tend to be more interpretable but usually they may lack some accuracy on the other hand if you choose more complex models like complex deep neural networks, they can be for sure more accurate, but they for sure lose some interpretability. So in order to achieve both, we worked like that. Of course we wrote everything in Python, we used CatBoost as a framework for gradient boosting over trees. I don't know if everyone is familiar with CatBoost, I'll tell you a bit more in a moment. Then we used Apple Flow for tracking the experiments and OCTUNA for hyperparameter optimization and then on top of that we used SHARP as a layer for explainability in order to peek into the decision-making process of the model and also debug it. So a few words about CATPOST. I'm sorry there are no cats involved but it stands for category and boosting and as I told you this is a gradient boosting framework on decision trees which handles efficiently and in a very smart way categorical variables missing values as well and it also has a peculiar architecture which is the symmetric trees or balanced trees architecture which has some nice pros like an efficient implementation on CPU, reduced inference times, but also a natural form of regularization. It also integrates seamlessly with SHARP as a method for explainable AI and it's also very easy to use with Optuna for automatic optimization of the hyperparameters. In general, the trade-off between precision and regular sensitivity is a function of the end user which has to adopt and use the model. Why so? Because the cost of false positives is generally very different from that of false negatives. So instead of just delivering one model we devised what we call three operating modes. So you might go for the high precision mode when the false positives are more costly so for example you don't want to issue frequent unnecessary alerts that could lead to other fatigue or costly countermeasures on the other hand I recall or high sensitivity mode might be preferable for you when false negatives are more costly so for example you want to detect as many real events as possible even at the cost of some false alarms because missing an event for you could disrupt some critical communications for example so for early warning systems or safety critical applications you will prioritize high record or high sensitivity but for operational purposes operational systems with also possibly costly countermeasures you might aim for the high precision mode just to avoid overreacting to benign space weather fluctuations. And when false positives and false negatives matter equally to you, you can just go for the balance mode which is the one that maximizes the F1 score. A few words about SHAP. We have given centrality to trustworthy AI matters and this in order to go beyond the black box approaches, placing some emphasis on interpretability and explainability of the model. And this framework comes from the cooperative game theory, where the inputs of the model are conceived as players taking part into a cooperative game, which is the machine learning model, and the output of this model is essentially a price. So the output of SHAP is how can we fairly distribute the price, the model output, among the players taking part into the game. And as you can see the final output is decomposed into a reference value, this one, plus a bunch of real numbers which are nothing else than the feature contributions for a specific sample. So this is a local explanation, okay, where the sign tells you whether a certain driver pushed the prediction up or down and the magnitude of the Shapley value tells you how strongly that feature impacted the model output. So Shap turns the prediction into an additive story, and essentially that waterfall plot is just the additive path from the baseline to the final prediction. So with this tool in hand, we can peek into the decision-making process of the model and And we can see how it reasons over time and which drivers are contributing to the decision-making process. And when you aggregate those instances over the entire set, you can also build a global explanation in the form of a feature importance, a ranking of features which matter the most on average for your model. And this is very nice because this allows you to make contact with the domain expertise. In this case, it was physics, but it can be everything that you like. So this is very useful in order to build trust in the model, to drive user adoption, but also to debug your model. because you can go with this kind of charts to your fellow colleagues or to domain experts and ask them is the model taking the right path or is it just taking the wrong shortcut and it's not learning the meaningful meaningful things but before shipping any model to production you should characterize it a bit more and we should understand that we have to move beyond point predictions and to this end we modeled risk and uncertainty for this model as well so we are currently serving near real-time forecasts for this model via the ESGUA platform at IN2E which stands for electronic space weather for upper atmosphere. But the reasoning works like that. In the real world, acting on forecasts as a price tag. Cost-sensitive decisions can be more naturally addressed within probabilistic forecasting lengths. And just for example, posing an expensive trilling operation because of a false alarm is costly, sure, but missing a severe ionospheric disturbance could mean losing a drone or compromising a critical communication and this equation essentially just formalizes that business logic so the expected risk is just the probability of an event times the cost of that events occurring so given accurate estimates of the costs that are associated with force positives and force negatives you can build a system that automatically triggers decisions according to user-defined risk levels. But before using this approach we have to understand that standard machine learning models are poorly calibrated. Why so? Because 90% probability is rarely a true calibrated probability and in order to make some safe decisions point predictions are not enough and if you look at this chart we are essentially serving not just a single probability line but also prediction intervals and so to to get some robust intervals we need to look beyond the standard outputs so it comes uncertainty uncertainty quantification can be broadly divided in two classes of approaches intrinsic methods or extrinsic or post hoc approaches based on the fact that you do require some retraining of the underlying model or not and to give you some examples of methods you can think of Bayesian approach and quantile regression as methods for intrinsic uncertainty quantification while on the other hand good candidates for extrinsic methods are calibration and conformal prediction and I will tell you something more about conformal prediction which as you might guess it is a statistical framework for any machine learning actually AI model to turn point predictions into prediction sets or prediction intervals. This is a nice approach, also a quite recent one, because it is almost distribution-free, in the sense that you, in general, just require some exchangeability, which is a milder IID assumption between test and calibration data. I'll tell you in a moment what calibration data is. So, essentially, you ask that the distribution stays the same between test and calibration sets under some permutations of the data points. And that's the recipe for conformal prediction for one experiment. You can use it for classification and regression, and it's quite easy to achieve. Those are the ingredients. you need a model that has to be already trained after all this is a postdoc method so you do not require an explicit retraining you need a heuristic notion of uncertainty which is a metric a distance between the x and y's a never rate or an empirical risk that you have to set say five percent if you want to aim at the 95 percent confidence level and you need a pinch of fresh unseen data for calibration so you need a separate set for your model training testing validation and calibration data and those are the instructions you define the nonconformity score as a a disagreement between your inputs and the outputs. You then evaluate those S scores, which are numbers, on the calibration set. So you learn this quantile Q hat, which is related to the error rate that you set at the beginning of the process. With this Q hat you move from the calibration set to the prediction phase so you can form prediction sets and this has the nice property in the sense that the probability of a new unseen point belonging to the conformal set is related to the error level that you set at the beginning and you have mathematical guarantees about that and And that's the key difference. I'm sure I confused you a lot because it's a complex topic. So if you want to learn more about conformal prediction, you can listen to my talk from PyCon Germany 2025. And wrapping up, trustworthiness is not just extra polish for your model. It's very important. And if you want to shape models in production and have reliable operations, you have to quantify your uncertainty, it's very important, and possibly also try to reason in a cost-sensitive perspective. And the thing is also that explainable AI and uncertainty quantification answer different questions for your model. So the first one belongs to the uncertainty side. conformal prediction for this case gives you intervals or sets and so the model can express the ambiguity instead of just pretending to be certain and the other question goes in the direction of explainability so SHAP for this case helps you attributing prediction to the relevant physical drivers and also opens the door to the scientific interpretation. So what changes is that the output becomes something that you can inspect, that you can challenge, and it's also something that you can then adapt to the user context. And that is why uncertainty quantification and explainable AI have to be part of the workflow, and not just decoration at the end. And that's pretty much what I have for today so I'm open for questions
Speaker 1 [23:43]
Here are some questions. You mentioned around 1600 events over 8 years and seems like too little data to confidently predict an event within the next 3 hours. How did you solve this issue?
Speaker 2 [24:02]
It's not too little, it's how nature works, because we have this kind of events with a fixed rate over time so we can't really tell the Sun, oh please try to produce more LSTDs. The thing is that of course we try to work with colleagues to enlarge the data sets. Of course it was not possible to go backwards in time because there were not enough instruments for ionospheric measurements. We tried to control the confidence of the model by looking at the prediction intervals in output. So, when the model saw enough instances as a specific sample, it was providing estimates which were narrow enough. So in this case, we were okay with the fact that the model was confident enough. For sure, when you deploy this kind of models to production, you have to account for concept treat coverage sheets and so in that case you ship your model to production and monitor the performances and then take action
Speaker 1 [25:29]
How do you encode the time-science data to predict with the CAT Booster? How far do you look back? Also goes in line with the data points across time-questions?
Speaker 2 [25:42]
gradient boosting framework you cannot understand what time is so you have to perform some feature engineering tricks to tell the model something about the time so this was done in the form of moving averages exponential moving averages also lagged features which we were design deciding how far back in time to go by looking at correlations between features and the target of the model so if I remember correctly we go back inside to take into account like no more than the previous six hours of data
Speaker 1 [26:35]
If I understood correctly, you work on prediction of LSDRI case.
Speaker 2 [26:40]
What happens next?
Speaker 1 [26:43]
Are you sharing the results with GNS?
Speaker 2 [26:48]
a very nice question. I didn't mention that this was part of a European project, a rising Europe project, and between the stakeholders involved in this project, there was the German Federal Police, the Bundespolizei, because they have instruments like the high-frequency direction-finding systems which are affected by this kind of events so the first phase for this project was just building a model which can tell them in advance if something was going is going to happen or not we are also working towards the integration into the European Space Agency monitoring room for space weather.
Speaker 1 [27:42]
Next question. Do you know about some work that is explaining the instability of the given forecast?
Speaker 2 [27:54]
I'm not sure that I understood this question, explaining the instability, because I can understand what instability is, but I'm not sure what do you mean by explaining the stability. I don't know if you are in this room, maybe you can clarify. Maybe next.
Speaker 1 [28:12]
Next one. Can you imagine some of these methods to be used in demand forecasting in explaining the models and their either causal dependence on the future covers used?
Speaker 2 [28:30]
For sure you can use these approaches in demand forecasting as well, but as you correctly pointed out, it would be much better to also look at integration of causal inference. Because the problem with these kind of models is that they are just learning correlation and you are not sure that they are picking causal effects.
Speaker 1 [29:04]
topic and I had no chance but using MATLAB because it provides important functions. How come we are able to use Python?
Speaker 2 [29:16]
But it refers to... I'm sorry for doing that, you had to use Matlab. I don't know, what kind of functions are you referring to? Are you in this room? Yeah, for geometry and carrier function calculations, so the distance, but I think it was a bit off-topic than your topic. Yeah, at Young, to be honest, it was quite easy to work in Python these days. I don't know when you graduated, but it was definitely easy and a joy, because you have the full ecosystem for AI models. If you are curious, we can discuss it also later.
Speaker 1 [30:08]
I think that we have already answered all. Okay. Thank you very much for the talk.