On Interventional Generalisation

Interventional generalization addresses the challenge of determining whether a specific action will improve a desired outcome in novel, unseen situations. Standard machine learning validation sets are insufficient for this problem because predicting an observation differs fundamentally from intervening in a system. A primary difficulty is that counterfactual outcomes are unobservable; once an action is taken, the alternative reality where a different action was chosen cannot be seen.

The proposed approach integrates structural causal models (SCMs), Bayesian hierarchical modeling, and decision theory. SCMs use Directed Acyclic Graphs (DAGs) to distinguish between conditioning and intervention, employing "do-notation" to break links from parents to the intervened variable. To handle "known unknowns" and transferability across different groups or regions, Bayesian hierarchical models estimate between-group and between-sample variances. This allows uncertainty regarding how an effect generalizes to a new environment to be internalized within the model.

Key takeaways include the use of KL divergence to measure interventional detectability—the ability to perceive the consequences of an action. Decision-making is further refined by splitting uncertainty into epistemic (reducible through experimentation) and aleatoric (inherent randomness). By calculating the expected value of an intervention both with and without perfect information, practitioners can determine if further data collection is necessary. Finally, the methodology suggests incorporating safety margins to account for "unknown unknowns," acknowledging that unmodeled confounders typically cause interventions to perform worse than predicted.

This description was generated by Open-Source AI using the transcript of the session and the original submission contents.

This session took place in track Machine Learning & Deep Learning & Statistics and was classified suitable for intermediate domain by the speaker.

Submission

The proposal as submitted by the speaker before the conference.

If I do X instead of Y, will I get the outcome I want (in a novel situation)? Making predictions alone is pointless, one wants to act in the world. Furthermore one must act in situations that are similar but different to all past situations. The real underlying goal of all decision making is interventional generalisation: the ability to evaluate hypothetical choices in new unseen situations.

This talk covers this history and problems of null hypothesis significance testing, the benefits (and limitations) of Bayesian reasoning. Introduces the basics of Pearl-ian causality theory and its treatment of interventions and counter-factuals (things that hypothetically could have happened, but didn't), finally we discuss the next step, interventional generalisation, that is being able to compare the value of hypothetical interventions in new unseen situations. Decisively improve your modelling practically and conceptually with the mental tools in this talk.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:00]

Thanks everyone for coming. I'd like to introduce our speaker, Andy Kitchen, who describes himself as a hacker human and a lover of AI before it was cool. So he's speaking to us today on interventional generalization. So I'd like you all to welcome Andy and please kick off.

Speaker 2 [00:28]

Matt, like really before it was cool. There was a period where you'd go to a party and try to talk to someone about like AI or cryptocurrency and they'd be like, please leave me alone. And now you'll be just at a party and unprompted people will be like, hey, what do you think about crypto AI blockchain? You're like, mm-hmm, pretty cool. So yeah, I guess that the trick is to have a really, really long title that's hard to say and then people will think it's smart or something. So it is tradition for me to start with a joke. So, okay, there's this doctor, and he's on the text chat, because now everyone goes to the doctor remotely, obviously, and he's chatting to this person, and the person says, Doctor, please help me, I'm depressed, the future is uncertain, I want to make the world better, but I just feel like I'm making it worse. And the doctor says, look, you know, don't worry, there's this great new app, you can install it, you can talk to the bot, and that'll cheer you right up, it'll make you happy, and it'll cure you. And then there's just all these crying emojis, And the doctor says, you know, what's wrong? What did I say? And the response is, I am Pagliacci bot. It's an old classic. I'm sorry, I made Laura sad. Okay, so intervention, interventional generalization. What is it? So I think this is one of these talks, So the reason I did it is because this is a talk I want to see, and it seems like a problem that really exists, but nobody really talks about it, at least in this kind of holistic way. So here we go. I will claim that this is sort of the central problem of actually doing stuff, not like statistics, but just doing things in the world. If I do X instead of Y, will I improve outcome Z? And you can see that we have an intervention, We have a counterfactual, we have a utility, and we have also a question is, we wanna do this in new unseen situations. And this really simple framing seems to me like a central question that needs to be addressed, but what's very strange is that you can go through an entire data science course, an entire ML course, probably an entire degree without seeing this simple question posed, and like, well, how do we solve it, what do you do? And I would say that there's a lot of great theory that has been developed in academia, mostly, I guess, decision theory for the utility question and Perlian causality for the counterfactual and sort of interventional questions. But in terms of having some sort of coherent methodology to apply it, it seems really sparse. So you can kind of say that there are sort of like three basic things. You've got like signal and noise. And there was even like a book by Nate Silver called Signal and Noise. And this is really the focus of a lot of statistics so you're trying to say okay I have some sort of random sample I have some sort of process which can which is sort of corrupting my observations but I want to make them more accurate more clear and then you have sort of this move to will cause an effect what I actually care about is not observing and predicting but in fact acting and intervening in a way that's beneficial to me so there is I think a lot of people now who are getting into kind of causal ML, it's still not obviously as well developed as just, okay, average over noise, but it's something that people are moving towards. And then I'd say there's this last step, which is kind of linking causality to decision theory, adding utility, deciding what to do and whether it's a good idea, which I would say is like the least well developed. And also it's the least well developed in terms of having some kind of semi-coherent methodology so i'd say the big the two huge challenges when you come to intervention and why i think interventional generalization is really something you care so much about is that the normal sort of test validation sets just just don't work because if you want to observe you want to predict you can just kind of take your data and pretend you never looked at it but if you need to intervene in the world then you need to actually have some sort of live thing that you can you know You have to do it to see what the effects are. And I would say the other thing is you can't observe counterfactual outcomes. Obviously, if I have a choice between doing A and B, then the challenge is that once I've done A, I don't get to see the world where I did B, so I can only infer something about it from my observations. And I would say, why does systematisation matter? If you can make complex intuitive judgements, just do it. You can do it quickly and you can do it high quality, don't worry about it. But the real question is, when you come to novel, complex, difficult situations, what is the actual systematic kind of steps you can apply to turn your base judgments into these more complex derived judgments? So systems kind of allow us to coherently mix simpler judgments into more complicated judgments. This is a work in progress, so I'm very interested in feedback, criticism, and collaboration. Roast me. And, yeah. Okay, okay. Okay, so just a quick show of hands so I have some idea. Who has interacted or engaged with Perlian causality or like structural causal models at all? We've got like, okay, about a third. That was sort of my guess. You're all very cool. Everyone else, you're about to be very cool. So what is causality? It's essentially an extension to probability theory that allows us to deal systematically with interventions. So the kind of joke I have here is that correlation always implies causation. Why? Well, because if you have A, B are correlated, then either A is causing B directly or indirectly, B is causing A, or there's a third thing which is causing both. And this takes something that's kind of mysterious, like, ooh, spooky causality, to, like, your job as a scientist is, like, figure out which one of these three it is. That's actually kind of boring. But why do people say that? And I think it's because conditioning is not intervention. So imagine a situation where I want to compute the probability of cancer given your teeth are white. Now, here's my little Bayesian network, and I condition on stained teeth. That means I infer that it's less likely that you smoke. Therefore, it is less likely that you have cancer. So the statement T cancer given teeth equals white is actually a decrease in the... Conditioning decreases your chance of cancer. But this is obviously stupid. And the reason we're getting kind of misled here is when we use a sort of causal approach, one of the things we say is, well, OK, we have this do notation. So we're actually intervening. We're saying, I am making your teeth white. I'm whitening your teeth. And in that case, we need to cut the incoming edges of the variable we're intervening on. So if we go back to our model, we make the teeth white. We're intervening because of the do, so we break the link from smoking to stained teeth. and now it's obvious that changing stained teeth has no effect on cancer and so our intervention has the right probability. So I think when people say causation does not imply... Sorry, correlation does not imply causation, what they really mean is conditioning does not imply intervention or conditioning is not intervention. But to kind of precisely even think about and answer this question, you kind of need to engage with the concept on the concept and theory and systems of causality. And just some technical details. A structural causal model is a DAG where each node is a function of its parents, and this is actually like an equation, and then you have some kind of exogenous noise terms. So you can kind of think of this as the direct causal version of a Bayesian network. Recommended reading, Judea Bell, both Book of Why and causality. So I guess, you know, there's that classic thing, We have known unknowns and unknown unknowns, so we have to treat those first. So we're trying to intervene and we're also trying to generalize. So clearly we will have some known unknowns. And this is I think where Bayesian hierarchical modeling really, really becomes very useful because it lets us estimate between group and between sample variances. So this is where like we have a number of experiments we've done say in different places, in different countries, different states, and now we need to act in a new country, and we want to know how much does my data even transfer over. So these are also called random effects models. So imagine a generative process where the effects that we're drawing come also from some distribution. That seems kind of abstract, but you'll see a diagram and it will make a lot of sense. So what we do is we have our hyper-prior, yes, we're Bayesians, we love priors, we've got priors on priors. That's right. I heard you like priors. So here we've got our hyper-prior, and what we're going to do is we're going to draw effect distributions for, say, each of our elements in the hierarchy, each of our groups, we'll call this maybe a school, and then each outcome is going to be drawn from this distribution. And these also have parameters that are inferred, they're variation, they're inferred. So imagine if we're seeing a lot of variation between the effects, then that will mean there is a weak sort of clustering between effects in different groups. But if we have the effects are very, very, very tight and they tend to cluster together, that will be evidence that that effect will generalise to a new group, a new school, a new state, whatever it is. So what we're basically trying to do is take what used to be out-of-model uncertainty and put it into a larger model. So now the kind of uncertainty that we gain from moving between groups, moving between states, moving between countries, moving between years is now part of the model. So you can sort of imagine we have these three levels of effect. Here's some equations if you're into that, if you're like a, you know, if you're a real hardcore Bayesian, you're like, oh, yeah, look at that, sampling, sampling equations. Oof. So, and then the last thing I'll say is that if you have an observational model, which is Bayesian, you can also link it into a structural causal model. I think this is the kind of viewpoint that I have to do all or nothing structural causal modelling, but especially if you have a somewhat new situation, you can create a structural causal model from your domain information and then some parameters that you learned through observation, you can plug them in and it's the same uncertainty, it's just distributions over values. So you don't actually have to use a causal model all the way through if you don't want to do all the causal modelling. You can use some inference and some causality. So yeah, recommended reading is if for random effects models and Bayesian hierarchical models is definitely BDA3 by Andrew Gelman. So I'm going to propose a couple of things, and the caveats here is I have not used these in anger, so I want you to be guinea pigs, I want you to try it and then complain to me about how it doesn't work. So that's your job. So here we go. I really think that an important measure is the KL divergence between your outcome distribution after intervention, hypothetically, and before. And if you have a large separation, this is something you can work with, because it means counterfactually, when you've actually done the intervention, it will be something that you can see you did. The problem with being in any situation like this is that even if you intervene, when you get the outcome, when you sample the outcome variable finally, when you actually, you know, things come to pass, you won't even know if you did anything at all. So it's very, very hard to get feedback. So I'm going to claim this is a situation that's very bad for sort of like where you intervening in this new situation is going to be very, very difficult because it's hard to get any feedback. So I'm going to call that interventional detectability. And then we can sort of bring in decision theory, and so we can say, okay, we write down a value function, and I'll just add, or you can also just look at these graphs, you can just compare your outcome with the intervention, without the intervention, that's something you can do as well, just look at the distribution. So we're going to try and write down a value function on the outcome variable that we care about, so we can now start to assess costs. And I think this is really an area where all of a sudden different types of uncertainty begins to split, because when you're not intervening, epistemic and aleatoric uncertainty are really kind of the same thing, and we often treat them as almost exactly the same. But when you're intervening, they matter a lot, because epistemic uncertainty, which is uncertainty in how the world works, it's in your head, it's what you know. You can reduce that with experimentation. Aleatoric uncertainty is the uncertainty in the process itself. It's like if I intervene, it might not actually happen, I can't control the world perfectly, and so these begin to split because I can't reduce this by collecting more information. So we can compute, of course, the normal expected value of intervention, which would just be the expected value with a marginalised over our both variables, and then just subtracting the value we would have got anyway. So this is kind of doing something versus doing nothing, what's different, what's better. And then we can also compute this thing, which is a little more technical, which is the expected value of intervention given perfect information. And so decision theoretically we're saying, OK, if I had perfect information in this hypothetical world here, I have stuff I can't control, and I can compute the maximum value that I can get in expectation from my action, given all the stuff I can't control, but I in this thing have perfect information, I'm maximising over hypothetically perfect information, and then we're asking for the average over the epistemic variable. So this is the hypothetical maximum score we could get if we could run perfect experiments, if we could eliminate all epistemic uncertainty. And then we have a little kind of nice look-up thing here where we can say if the expected value of my intervention is high, and if they're both high, then I should act, but if my expected value of intervention is currently low, but with perfect information it's high, that means this system could be worth intervening on if I had more information that could reduce my epistemic uncertainty. This gives us more of a cookbook formula. We can measure these two things, we can look at the gap. So then finally, unknown unknowns. They're really hard. I would say you kind of borrow from engineering here and you add a safety margin because there's all sorts of stuff that can happen that you can't model, that's dumb. And there's this kind of, I think, general observation which is a causal regression to the mean, which is that because unknown confounders and mediators almost always reduce the effect of an intervention, Generally, interventions will just work less well than you modelled them. So we can sort of say, well, I want to know the probability that my intervention at least doesn't do any harm, and I can ask for that being over some threshold. So with my model, I can compute this quantity, and I can also say, you know, like, even though from a decision-theoretic perspective this is positive in expectation, there is just a huge amount of probability mass on this intervention actually makes things worse. So I could ask for a safety margin and say even though game theoretically it's still optimal to take this action because it's in expectation better, in practice my threshold is set to, I don't know, 0.9. So I want a fair bit of probability mass on at least this intervention doesn't actually make things worse. So this is towards. The reason I wrote towards is because you're moving towards it. Not there yet. But towards a methodology is, and it's going to be a bad one, but at least it's a starting point. We're trying to glue all these concepts together. So write down a causal model for your intervention. Write down a Bayesian hierarchical model linking your parameters to observations. And what's really nice about this is you can systematically move uncertainty from your causal model. So move uncertainty all the way from observations into action because they're now linked. So write down a causal model. We write down a Bayesian hierarchical model for our parameters to observations, and then we're going to simulate the outcome distribution and calculate the interventional detectability. If we can't even see the consequences of our own actions, this is going to be, I think, too challenging to do something useful unless you have a special reason not to. Then, if we're kind of happy, we're going to write down a cost model, we're going to identify our aleatoric and epistemic sources of uncertainty, and we're going to calculate the expected value of intervention with or without perfect information and we're going to use the gap to decide if more measurement is useful and we'll use domain knowledge to set our safety margins and then I guess oh there's a slide that I meant to put in which is a big smiley face saying act so just imagine you're looking at that right now there we go okay so thank you I hope that was useful and I'm excited about questions I'm extremely excited about questions we need like a fun slide

Speaker 1 [18:13]

Great. Thanks for that talk. Super interesting. Andy specifically asked for more time for questions, so ask away. We've got a little bit more time than usual. So we've got some great questions here. I'm going to start with one. If I get this wrong, jump in if you're in the room. So one person is asking, I was thinking A-B testing in the beginning, but is the problem instead that we want to answer the question what if I was a poet instead

Speaker 2 [18:44]

Well, life would be better as a poet. I mean, you know, should you quit data science right now and become a poet? Well, I mean, I think you should create a causal model and then I should write down your cost model and so on and so forth. Yeah. So, yeah, quit data science, become a poet. Next question.

Speaker 1 [19:04]

Alright, you could get poetic about this question, so would you say being from the future helps you make decisions?

Speaker 2 [19:04]

All right. Well, I mean, if you know the answer to everything before it's going to happen, obviously, life is a lot more fun. The parties get a little bit more crazy, especially because you know you're not going to die. So, yeah, but importantly, by applying these techniques, people will think you're from the future because you'll be so good at deciding things.

Speaker 1 [19:35]

Right. Well, we're coming to slightly less abstract questions now. Good. I like it.

Speaker 2 [19:39]

I like it.

Speaker 1 [19:41]

Could you give some sample scenarios with this theory as if I'm a high school student? So can you explain in simple terms where you might use this?

Speaker 2 [19:50]

Yeah, so we were sort of already short on time, and I think a whole talk could just be a really meaty case study on trying to apply this to a real-world situation. I would say that especially Rubin causal models have a long heritage in medical research. I mean, ultimately, there is a whole school asking the questions, you know, if someone had got this medication instead of this other one, would their individual treatment have been more effective, which is often something that's completely glossed over when you're taking sort of population averages. So starting to think about these things like individual treatment effect, counterfactual treatment effects, average individual treatment effects as opposed to population effects are all like... So that's an area where it's very important, I'd say, better developed than it is in sort of data science. And I think that the... It does seem to me that the other area it's popping up a lot is unfortunately sort of advertising modeling um advertising spend like i know some of these bigger companies wise i think someone was saying um uh uber are sort of working on like causal and larger causal ml systems and a lot of that is for sort of predicting you know if we have capacity here or value if we spend here what kind of return will we get so i mean what I think is kind of funny is that this is so incredibly applicable because most of the time you don't want to make predictions you want to act and often you get really confused if you don't have this even if you never use any of this modelling if you just kind of engage with causality you at least have a basis to say like why is this stuff so confusing because often data scientists are using conditionals to answer interventional questions and like even explaining what the role of like a mediator or a confounder is in those situations can get really dicey so um concretely you want to answer questions like especially where you can never go back on them so you're going to spend money this year and you can't it's spent you can't go back and do 2026 again a patient is going to get one treatment or the other and you can't go back and do it again i think if it's a situation where you get a lot of hits at bat like you can kind of show of people adds over and over again, then it almost is like a stochastic process. So you can kind of skip some of this stuff and just go straight to a kind of like bandits thing or something. You don't need to sort of put so much care into modeling the causal structures because if you get something wrong, you just kind of do it again.

Speaker 1 [22:27]

Great. Thanks. I've got two questions here which I'm going to combine because I think they're somewhat similar.

Speaker 2 [22:32]

Moderator superpowers.

Speaker 1 [22:35]

So, what is what tools libraries do we use in Python to tackle these intervention problems and decision problems? And the other was, do you use any packages to conduct specific experiments?

Speaker 2 [22:48]

yeah yeah bundle them so super super good questions and like i was intentionally um

Speaker 1 [22:48]

Yeah, yeah.

Speaker 2 [22:54]

like didn't mention tools at all in this talk because i think it's like a conceptual talk and there's a lot of like i would say even between like five years ago and now there's been a huge development in causal tools overall and i would actually say that i think doing it by hand is better like there's a lot of work especially on causal inference so you just have a data set and i want to throw my data set like in the black box and like a causal model comes out and maybe it's good maybe it's not but i think like there's sort of two schools of causality developing and one is the sort of causality as a supercharging inference tool so i still want to do my normal data science thing of like get a big data set close my eyes hit the button and like something comes out and then there's also this other school which is i think kind of a little bit more Bayesian flavoured, which is like no, causality is just a system for me to understand the world, for me to write down what I understand, to reason about it, to communicate it to others. And so once you have this tool and this vocabulary, the benefit is not that it's automatically coming out of the computer, the benefit is that you're writing it down and you're talking about it and you're sharing it, and the ability to calculate something at the end is like kind of a bonus. And like, you know, if you take a real simple model, like, oh, my screen turned off. If you take a really simple model, like this one, you could just write it by hand. You don't need sort of a framework. It really depends. If everything is a Gaussian and you've got a lot of linear relationships, you want to calculate some quantity out of it, you could sort of write it down. So I'll say that, like, maybe use it as a thinking tool first, and then the packages that can do inference for you are a bonus.

Speaker 1 [24:38]

Great. Now we have a question relating to agents. Do you believe today's agents could benefit from learning and exploiting this framework, maybe to become better at scientific discovery? Have you already tried to teach agents to do that?

Speaker 2 [24:51]

that that's a really really good question so um i i think the answer to me is in some ways obviously yes and i think there's sort of two principles for making that one is this sort of like um it would be kind of like the the the human the human priors argument which is simply that these uh llms have learned priors to how priors on how to think from human communication so thinking tools that benefit humans and organizational tools that benefit humans also tend to work on llms and i don't know if that's like a property of neural like the the intent attention architecture per se or a property of cognition it's just they they have social priors and cognitive priors they learned from us so if it helps you it'll help an agent um and i think i think that the the sort of second answer is like absolutely being able to systematize a judgment is extremely important if you want to work on it iteratively in groups like having great intuition and great judgment is really important and like to be a effective leader and effective thinker you have to have great personal internal judgment but the problem is that that's hard to transfer it's hard to scale and it's hard to reason about because it can when it goes wrong you just like make a bad judgment it's hard to know like why and how and so the point in some of the systematization here is that it is difficult and it is fraught but going through that process is something that you can do repeatedly in groups so i think if you have many agents that are trying to interact on some problem having like one of the things that are produced as a structural causal model that one agent can like implement in python another agent can critique in prose is something that would be very useful and yeah it's it's stuff i'd like to work on but i haven't got a chance to work on yet so if you want to teach an agent to do causal stuff talk to me

Speaker 1 [26:52]

Great. And then we have another question around simulated populations. So I think, question asked, I think there are some companies that I've seen recently that simulate an AI population to run these kind of irreversible tests. What do you think about that?

Speaker 2 [27:09]

Wow. Like, okay, okay. That's a wild question. So I think it's also a good one. And this kind of goes to the also, like what tools can you use? This is, I think causality overall is best as both a quantitative, it's an okay quantitative tool and it's also an excellent thinking conceptualization tool. So, you know, if your thing is, well, I want to do something where, like, I sample agents and I have a population of them and I test some kind of, like, product on them or something, well, you can go back and say, well, what are my epistemic sources of uncertainty and what are my aleatoric senses of uncertainty? and the epistemic stuff I want to model outside of the simulation because I sample those to get a world where the dynamics are fixed and then I go through and I have 10 agents arguing with each other and I also want to make sure my aleatoric sources of uncertainty happen inside the simulation, which in this case will just be like what the agents choose to do and how each agent will probably have its own sampling noise or whatever it is, its own temperature. So what exactly the agents say to each other in each simulation would be a classic example of, like, aleatoric uncertainty. It's exogenous to the particular model you're trying to test. So even just as an example, using these thinking tools, I think, is, like, extremely productive.

Speaker 1 [28:46]

I think we have time for one more question and then I think I'll finish off with a question that I'm reading more as a request so when it's called modelling does it mean we sample from data and then compute a statistical equation could you give an easy term what's modelling

Speaker 2 [29:04]

An easy term, what's modelling? I mean, like, what is a model is a really, really good philosophical question, right? Like, it really is. And I would say that we have, like, intuitive mental models and some people are better or worse at explaining them. And famously, sometimes people who are very good at things are extremely bad at explaining why. You know, like, you'll ask, like, a really, really good sports person why they made this particular choice or went for this particular shot and they're like, I don't know, it just seemed right. So what we're really, I think, talking about here is external systematic models. And in that case, systematism or systematisation means that it has kind of well-understood elements. The rules for combining and understanding how those elements work together are well-written. So, you know, we can write down a causal model and we can...

Andy Kitchen

About — in the speaker's own words

Born hacker. Curious human. I've started a couple of companies. I liked AI before it was cool, I swear.

Social card for talk: On Interventional Generalisation