🌳 The taller the tree, the harder the fall. Determining tree height from space using Deep Learning and very high resolution satellite imagery 🛰️
The risk that a tree poses to line infrastructure (such as power lines) is determined by several factors, chief among them the height of the particular tree. The increasing availability of very high resolution satellite imagery makes it possible to use photogrammetric techniques to extract height information from a set of stereo satellite images. By using satellite imagery we can achieve a scale not possible by manual measurement. We found that classical techniques perform poorly on vegetation, and were handily outperformed by deep learning based techniques implemented in PyTorch. This improvement was not trivial to achieve however, as creating labelled data in sufficient quantity was quite challenging. By increasing the quality of our height predictions we were able to more accurately calculate risk for our customers.
This session took place in track Machine Learning & Deep Learning & Stats and was classified suitable for intermediate domain / novice python by the speaker.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:06]
Cool. Yeah. Thank you very much for being here. So I want to talk to you, my talk today is called The Tall of the Tree, The Harder of the Fall, and I would like to talk to you today about determining tree height from space using deep learning and very high resolution satellite imagery. My name is Ferdinand. I work at Live.io, we're a start-up based here in Berlin, and yeah, just a quick overview of who Live.io are. Live.io is a start-up here in Berlin, we're an Earth observation company, as the name is live earth observation. And what we do is we take data, majorly satellite imagery, we process the data through our analytics platform, which includes a lot of machine learning. I myself am a machine learning engineer. And we translate this into some kind of actionable insights that our customers can use to solve their business problems. So in our specific case, I work on a product called Treeline. So, what is Treeline exactly? In Treeline, we do vegetation risk modeling for line infrastructure. So, line infrastructure can be railway lines, power lines, or gas pipelines. And at least in the US and the EU, vegetation interaction is the leading cause of power outages. So, this would either be trees falling into or growing into power lines. And so, this is a pretty big problem for the set of, like, customers. So in dry areas like California or Texas, for example, the interaction of vegetation with power lines, as an example, can even cause fires. And this can, of course, be a large problem for these companies as they can be held liable if they've been found to be negligent in these cases. And I hope I've said that, like, you know, by accurately determining the risk that vegetation poses to infrastructure, we can reduce the maintenance costs that these customers face and we can reduce outages. So, now that you have a bit of background, what we're trying to do here in Treeline is we're actually trying to determine what is the risk that vegetation poses to the infrastructure. There are a couple of things that go into risk. This could, for example, be the species of the trees, as some species are more prone to fall over during storms than other species due to their root systems. It might also be the health of the vegetation. Sick trees are much more likely to fall over than healthy trees, so being able to determine the vitality of a tree is quite important. But today we're going to be focusing on height, and height is quite a crucial part of this calculation. As you can imagine, if you have a power line that's, for example, 10 meters up in the air, a tree that is 5 meters tall, no matter how sick it is, is not going to pose much risk to this power line, whereas a tree that's 20 meters tall and pretty close to your power line might pose quite a bit of risk. So what we do is we segment the risk into different categories. We have to effectively tell our customers, hey, this vegetation poses low risk, this vegetation poses medium risk, or possibly high risk. And especially in the high risk cases, our customers will often go out and do something about that, like cut down the trees or trim the trees, depending on what they deem necessary. But our goal is to give them the information that they can make these decisions. So how exactly do we go about calculating the height of trees? So if you've ever worked in remote sensing before, you might know some of these terms, but I'll just introduce them to give you a bit of background. What we effectively want at the end of the day is a canopy height model, which is effectively you can think about it as an image where the X and Y location of the pixels tells you the location of every pixel, and the Z value, the actual value in the raster, tells you the height of the object pictured at that point in the image, but, yeah, we don't generally get these canopy height models directly, so we have to create them ourselves, so we start with digital surface models, which is a model of the earth with everything on top of it, like vegetation or buildings, and then we use something called the digital terrain model, which is a model of the Earth without buildings and vegetation of that. And if we then take the difference between these two, we basically simply take the digital surface model, subtract the digital terrain model, we can get the height of the object we want above just its surroundings. And so how do we actually generate these surface models? So one of the main sources of digital surface models is actually LIDAR. Some of you might be familiar with this from other contexts. LiDAR sensors get used, for example, in self-driving cars and other kind of scanning. It also gets used a lot in remote sensing, not from satellites, but from either airplanes or helicopters. And the great thing about LiDAR is you can get very, very accurate models. Some great things, like LiDAR planes or airplanes are often in contact with GPS satellites, so you can get very, very accurate location of the point clouds that these door flights create. So, like I said, LIDAR stands for Light Detection and Ranging, which is effectively a little laser scanner that scans the area around it. And you're often scanning at about 30 to 50 points per square meter, which is very dense. And what you get out of this is quite a dense point cloud. And the great thing about LIDAR as well is it can actually see through vegetation, so it can see the ground below the vegetation as well, meaning you can create nice digital surface models and digital terrain models. In this case, this is an image of a LiDAR scan, and you can see here you can get the brown points, actually the ground points that were scanned, and the LiDAR can detect this, and you can get the actual vegetation scans. So there are a couple of great things about LiDAR. It is basically the gold standard for quality. you can get up to centimeter height accuracy, and you can actually see through the vegetation, so you can actually get a model of the vegetation and the ground below it. However, there are also a bunch of downsides to LiDAR. Nothing comes for free. LiDAR often has very slow capture times. You need to, it takes quite a while to scan this. It can be quite expensive because you have to fly a plane or fly a helicopter to actually do this, and it can also be very region-dependent, Meaning, you know, in every country or every region that you work, you have to work with a different operator. So, it's not very scalable globally. So, are there any alternatives? Yes. So, there's one other way you can do this. And this is through using stereo satellite imagery. If you've ever worked in stereo vision before, you know that if you take two images of the same location from different angles, you can actually also create a depth map. in our case this will be a digital surface model that tells you how far away everything is from your camera and therefore how high it is and as you can see through this you can create with a little bit of clever processing you can create a surface model so i just want to talk a little bit about like i'll get into a little few more details in a moment but i want to talk about very high resolution satellite images which are the kind we use we need high resolution images because the The things we're picturing, trees are of the order of meters, so we can't use low-resolution imagery. So usually when we say very high-resolution imagery, we mean sub-meter resolution. And generally, the satellite sensors that we work with have resolutions of 30 to 70 centimeters per pixel. So can anyone tell me what location is pictured in this image? Yes, as a couple of you have noticed, this is actually an image that we took, or that we bought of Berlin. And if you have very sharp eyes, you will see that you are right around there. For reference, Live Viewer's offices are right down here in Kreuzberg if you ever want to come visit us. We're very friendly. And yeah, to give you a bit of a sense of scale to how big the images are that we actually work with, this particular image is about 40,000 pixels by 40,000 pixels. I think it's 37,000 by 40,000, but that's roughly the size. This is a relatively large image, but not excessively so by any means. This image alone is about 8.5 gigabytes on disk, which is quite big, and as I said, to do this kind of stereo satellite calculations, you actually need two of these images, so you are doubling up on this. This is quite a lot of data. Now, we generally don't work on these large images directly. We, of course, will zoom in a little bit, and work on a little bit of a smaller patch at a time. It's just a little bit unwieldy to process this much at one time. So if you look a little bit closer, that is also us. We are right now in the Kuppelzall, which should be right around where that arrow is pointing over there. And I also realized as I was looking at this, there's probably a pretty good description. There's an item pretty close to us here at Alexanderplatz that I can actually use to describe how we actually do this. And that is, of course, the Fahrenheit term. So if you've looked outside, you'll see the fanzotrome is there. It's a very tall building, and this nicely describes the phenomenon that we can exploit to actually calculate depth from images. So if you keep your eye on these red dots, you'll see the one red dot I placed on the top of the fanzotrome and the other one I placed at the base of the fanzotrome. So if you look at two different images of the tower, you can actually see that between the two images, the tall object, which is the top of the tower, will move quite a few pixels, whereas the, like, bottom of the tower will not have moved at all. And there's kind of a scale that, like, if you can figure out how many pixels an item moved between two images that you took, and you have an accurate understanding of where your cameras were when these images were taken, you can actually infer the depth of the objects in this image. And if you now go and do that for every pixel in this image, can create a depth map which we can then translate into a digital surface model which we can then use to infer the height of our vegetation. So, the initial approach. I'm a machine learning engineer, but my opinion on machine learning is that don't use machine learning unless you really need to use machine learning. So, our first approach was to use some techniques from classical computer vision. Stereo vision is actually a pretty well studied problem in computer vision, and this is an example from a tutorial from the OpenCV documentation, and actually in about five lines of code, you can create a disparity map. That's relatively simple. The algorithm that's generally used in these cases and that's used here in this example is called Stereo BM Create, but it uses an algorithm called semi-global matching, which is quite a nice algorithm. It works quite well. used quite often in industry. And yeah, that should be pretty easy, right? Yeah, how hard could it really be? Yeah, so we might have thought that it would be a little bit easier than it turned out to be. So this is where we kind of ran into the, you know, where we got our first reality check. If you look here, if you're not used to geospatial things, this might be a bit hard to interpret. But what you're effectively looking here is a profile view, so it's just a little profile tool that gives you a little slice that tells you what a DSM looks like, and in this case, if you look carefully, there are two rows of trees between some fields that are being imaged here. The lines in red is what we would like to see, that's our reference, that's LIDAR, because we are trying to be as good as LIDAR, and the black line is what we got when we tried these classical computer vision methods like semi-global matching, and as you can see in the one case, it completely didn't reconstruct the vegetation, and in the other case, it reconstructed the vegetation with about 50% of the height that you would expect. Of course, this is a pretty big problem. You know, if you go to our clients, you tell them, hey, the tree you're looking at is 10 meters tall, the client goes out into the field, they have a look, the tree is actually 20 meters tall, clients get pretty upset, and we do a bad job. So we realised we should do something better. So the signs were there. If you look into the literature, there's actually quite a lot of literature about this. This one says the limitations of high-resolution satellite stereo imagery for estimating canopy height in Australian tropical savannas. You don't have to read through all of these, but when you do go look through the literature, there's quite a lot of literature telling you that these classical methods do struggle quite a lot on vegetation specifically. So if you are actually trying to determine the height of buildings or other stationary objects, the classical methods might actually work well enough for you, but vegetation specifically is difficult for these algorithms. To give you a bit of an intuitive idea of why this might be the case, it's just that vegetation at the resolutions that we're working at, which is roughly 50 centimetres per pixel, is actually semi-transparent. So you're not just really imaging the tree, you're also imaging a little bit of what's behind it. So you can imagine that if you're taking an image of a tree from directly above, the pixel that comes out of that location might be very green, whereas you're also taking a picture of the top of the tree from a different angle, you might be imaging not just the top of the tree but also a little bit of what's behind it, and you might get a different color. Now the classical computer vision algorithms that are used here basically go on color matching. They're trying to see, like, hey, what is the color of this pixel? And this is a little bit simplified. What is the color of this pixel in this image, and what's the color of the pixel in that image? And it tries to match them. And so if you have this, like, change in color between the two images, it might not actually properly image it might not actually match the correct pixels, and you might just not reconstruct the object that you're reconstructing, or you might misreconstruct it as either being too tall or too short, or quite often if it just doesn't find a match, it will just tell you, hey, I wasn't able to find a match, didn't actually manage to do something. So just a little bit of the intuition. And so I'm a machine learning engineer, and, you know, I was starting to think this is starting to look like a machine learning problem. However, I wanted to be sure that, you know, just because all I have is a hammer, not everything looks like a nail, so I wanted to be sure, so we did a little bit of research, and had a look at is there any precedent for actually using machine learning, and would the machine learning algorithms actually give us better results? So I don't know if you know the Kitty Vision Benchmark Suite, but it's effectively a test case with a leaderboard that gets used quite often in academia for these kind of tests. So a lot of the work that's actually being done in stereo vision actually comes from the self-driving car community. There's a lot of reasons for them to be able to do this, and so they do a lot of the research. And so this specific case is a case where you have two cameras mounted on a car, and they have a LIDAR sensor, and the challenge is to determine the depth of everything in these images. So we had a look at this leaderboard, and we found that the method we were using, which was the OpenCV implementation of semi-global matching scores, is at number 310 on this leaderboard. So clearly, there are much better methods over there. And if you actually go and look, the first couple of methods, actually, basically the top 100 methods are basically all deep learning-based methods. So clearly there is something to be said for deep learning in this case. And just to give you an example of, like, what this actually looks like in these cases, this is one reference image from the skitty data set. And on the left-hand side, you will see the prediction that you get from semi-global matching or these more classical algorithms and on the right hand side this is the prediction from the top deep learning algorithm. As you can see there is actually a very vast difference in quality here. Semi-global matching in general reconstructs most of the big things but it really misses out on the details and it really has a lot of artifacting and especially in our case where we're working with satellite imagery where you know our trees are only a couple of pixels wide, we are actually kind of working in just the details. So it's kind of interesting to see how far the field has come. I think semi-gold matching, the paper was actually published in about 2007. So I mean, in real-world terms, it's not that old, however, in the machine learning world, it's pretty ancient. So yeah, we thought like, okay, there's definitely something out there, let's try and use deep learning. So of course, one of the big problems with using deep learning is it's very data-hungry, so now we actually have to somehow find some data that we can use to train. There are open data sets out there, for example, the kitty data set, however, it's quite a small data set, and it's a data set of cars, it's not satellite imagery. There's also some synthetic data sets that people have created just in, you know, computer graphics products, like Blender, for example, however, these don't really match with what we're working on, so there's a problem with distribution drift, so we really had to go and create our own data, training data, however, this is a little bit easier said than done, because how do we actually create the labels that we want to train our deep learning methods on? You can't go label this by hand, you can't really have a person sit there pixel by pixel and match this, this would take forever, so we needed to find a different way to do this. The requirements are that we have a set of stereo images of a location, so we can task these, we often buy the images from one of our providers, and then we also use LiDAR data to actually create our training data, so we need LiDAR of the same location at a similar time, especially because we're working in vegetation and vegetation isn't static, It's pretty important that the LiDAR scan is from at least a similar time, usually within a few months of when the image was taken. If you take a LiDAR scan that was taken, for example, a year later or two years later, the trees will have grown quite a bit, and you're just going to introduce noise into your data, and you might train, you know, your model is going to struggle to train a little bit. So, yeah, that's a pretty stringent requirement that takes quite a lot of cross-referencing to find these things at the same time. thing is you have to co-register the data with the stereo images. You may have to make sure they're at the same location. This is easier said than done, and I can go very deeply into all kinds of things to do this, but, yeah, that's one of the things about working with geospatial data. Not only do you have images, you also have to make sure that they are at the correct location. And then what we basically then do is we take the LiDAR point clouds, we project them into the reference image, into the secondary image, and by calculating where they end up in these different images, we can actually calculate a disparity map, which is the training data that we actually use to train our models. And yeah, so based on where in these images these points end up, you can create a disparity map, and you can train your model to actually then produce this on unseen data and data for which you don't have LIDAR. And yeah, I just want to shout out a couple of tools that we use to do this. We use RasterIO to open the satellite images, PDAL for processing the point clouds, and JAX, actually, to do some of the projecting of the points into the images. So I want to talk quickly about some of the deep learning models that are used here. So the models quite often have very similar architectures. This is just a little cartoon of one. But effectively, you quite often have convolutional neural networks as feature extractors. You then use these feature extractors to build up some kind of cost volume, which is effectively where you just stack the features on top of each other. And what you're trying to do then, there are different ways. Some people use 3D CNNs. You can also use some kind of attention where you try and figure out, like, okay, what pixel, and you look at the embeddings. you see what embedding of this pixel and this image matches with an embedding of an image and another of a pixel in another image, and by matching these embeddings of these images, you can actually get a depth map, and that is how you finally get out your DSM and your final prediction. Also we did this mostly in Torch, and using PyTorch Lightning for some of the nice training steps. I just want to highlight one thing, it's not just as simple as pulling off some models of GitHub and just running it, there are a couple of snags that we run into. For example, that the majority of models that are used in this space and the majority of the data sets that you can find come with the implicit assumption that your cameras are actually parallel. This makes computation a little bit simpler because then your disparity can only be one-directional. And this is quite often the case in industrial robots, you'll have parallel cameras, in cars you'll often mount your cameras parallel and realistically our eyes are actually a set of cameras and we can infer depth and our eyes are fairly parallel however just due to the way that we image with the satellite images and the fact that the satellites have to turn in order to take these images this is not the case for us in our case our sight lines actually interact and we have a slightly more complicated case the slightly more general case where we actually have both positive and negative disparities, and we have to be able to deal with both. So we found we had to go do some neural network surgery, we had to go really dig into the details and change these models quite significantly just so they can actually handle this more general case. Then, after a little bit of training, we started getting to the point where our models learned, and we were very happy about this, of course. And this is the same location as you saw in the first image. Again, in red, it's the LiDAR, which we're trying to achieve the same quality of. The black was the quality of our first attempts using the classical computer vision methods. And the green that you see now, I hope you can see that, is the outcome of our deep learning methods. And as you can see, we're pretty much right there with the LiDAR in terms of determining the height of these trees. And yeah, just a last word on in-production. As the Germans say, ein Mal ist kein Mal. For our large jobs, we often have hundreds to thousands of stereo pairs. That can mean one, you know, in the single digit to up to the lowest double digits, terabytes of data per job that we have to process at 50,000 by 50,000 pixels per image. This can easily be doing inference on hundreds of thousands of image patches at a time. So we're using array and prefect to orchestrate large-scale inference for this, and, yeah, a couple of conclusions. I hope I've convinced you that height plays an important role in the risk that vegetation poses to infrastructure. Stereo satellite imagery is an efficient and cost-effective way of measuring vegetation height at scale, which is very important. We don't care about... We're not measuring single trees here. We're measuring thousands of kilometres of power lines, as an example. Traditional computer vision techniques, while interesting, are inadequate, especially in the case of vegetation. And deeper learning-based techniques can overcome these limitations and actually accurately reconstruct vegetation height for us, making it possible for us to solve these problems. So yeah, thank you very much for coming to my talk. My name is Ferenc Henk. I work for Lavio. And yeah, we are currently actually hiring. So go have a look at our jobs page if you want to come join us.
Speaker 2 [24:39]
Thank you. That was very interesting. I really enjoyed it and learned a new German phrase. I know this kind of language I really like. Yeah, it seems the audience enjoyed it because we've had quite an active question session. So I'll kind of group them into kind of general topics. So there's a question about the impact of seasons on the height estimation accuracy. if you considered or calculated the movements of top of a tree because it's kind of compared to the TV tower again whether you require subtle images from different seasons yeah
Speaker 1 [25:20]
Maybe let me start with the seasons one. Yeah, we found the seasons one is very important. We only really image, we're quite strict on only imaging these areas in what we call the leaf on season, so when the trees actually have leaves. So we'll generally, depending slightly where in the world this is, only take images from April to about September. In some more southern places, you have a little bit more leeway. In some more northern places, you have a little bit less. but definitely it's very difficult to actually reconstruct trees in the winter so we basically don't even try we only take images in the summer what was the other question?
Speaker 2 [25:58]
There are questions along these lines of kind of how difficult is it to differentiate between terrain and vegetation, lighter points.
Speaker 1 [26:09]
So for LiDAR, so we actually, so for LiDAR to determine, to differentiate between ground points and vegetation points is relatively simple, actually. You can basically look, like LiDAR has a couple of cool things, like you can see the return, like how long it took for points to return. So within a specific grid, you can have, you'll have a bunch of points, you'll have a bunch of returns. and you can see the you do some infrared geometry and based on how long things took to return you can see was it a ground point or was this a point in vegetation and we actually still use lidar we use lidar for the terrain models because it's still the best way to get terrain models
Speaker 2 [26:53]
And there's, again, a couple of questions relating to data sets. So if the data source is accessible for everyone, if it is feasible to generate some synthetic data sets for training, another, if it's possible to use image generation AI to generate learning data.
Speaker 1 [27:11]
Yeah, so the last one, generative AI, I'm not sure whether that is really within the realm of possibility just yet.
Speaker 2 [27:12]
Yeah.
Speaker 1 [27:20]
However, we're really interested in synthetic data. We haven't really used synthetic data, but we're busy looking into that. And that's a pretty great way of doing it because there you actually have very accurate locations of the objects you're looking at. However, we just want to make sure that we're, because we're really interested in vegetation, that we are accurately, you know, making accurate 3D models of the vegetation and that the way the light acts with it is accurate. We don't think we can use purely synthetic data because, you know, there are some, it's really, really hard to model, you know, atmospheric effects and, you know, real-world trees and things like this, but we're pretty interested in that as kind of like a pre-training task. There was another question. What was the first part of the question?
Speaker 2 [28:02]
So is the data set publicly available?
Speaker 1 [28:02]
So is. The data set that we have is not publicly available. However, there are actually some open source data sets available. One of them is called the US3D data set, if you want to have a look at that. And there's another one called the WHU stereo data set. It doesn't work on the same satellite sensors that we're working on, but they are publicly available and you can train on them.
Speaker 2 [28:30]
And then I guess we have a minute left, but we have so many questions. So just as a note for everybody, you will be around if people want to find you.
Speaker 1 [28:38]
Yeah, I'm going to be around here for the next two days, so just stop me and talk to me. I'd be happy to talk to you about, like, machine learning or geospatial data or anything, really. And otherwise, you know, add me on LinkedIn and send me a message there.
Speaker 2 [28:52]
So I'll maybe try one or two more questions. So if you're training, it's an interesting one, if you're training data is from certain geographies, would the model basically overfit to those specific vegetation types? Yeah.
Speaker 1 [29:02]
Yeah, exactly. So we do take care in getting vegetation data from all over the world. We do serve as customers in Australia and the US and the EU. And so we do try and sample our data from a large range of geographies. Because we exactly, you know, if we only train in Australia, you know, you can think that if you then work in Germany, you might not get, you'll probably get okay results. But, you know, I'd be very wary of just doing that. So we really do try and get data samples from all over the world. Thank you.
Speaker 2 [29:33]
Thank you, everybody, and again, sorry if I didn't get to your question. I think there's 10-plus more left that I haven't gotten to, but thank you for your interest, and see you for the next talk.