What we learned from scraping 1 billion webpages every month

As Prisync, we crawl a large portion of the web every day for 6 years. First we approach the problem with a naive aspect, but we learned our lesson via experience. Developers create workarounds and hacks all over the time. But doing so has –most probably, unexpected– consequences. Some of the glitches we experiences so far:

There are ;

  • websites not responding properly
  • websites responding different output to identical requests
  • websites not responding at all
  • websites not obeying HTTP at all
  • websites with broken firewall rules
  • websites served on archaic webservers, which even are not aware of current state of transfer protocol
  • websites taking advantage of vulnerabilities (a.k.a. "clever hacks")

In this talk, I share examples of those "hacks" and I propose some methods to keep the web healthy.

This session took place in track PyConDE and was classified suitable for some domain / none python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:02]

Thank you. Thank you, Eran. Hello, everyone. My name is Samet, and welcome to my talk. This is my first ever PyCon talk, so I'm super excited, really super excited. And English is not my main language, so please don't push too hard. Sorry. My name is Samer Attak and I am the CTO of a company which is called Pricing. We do large-scale web scraping. We serve like several companies from all around the world. We have German customers as well. Metro, Grossmarket, MediaMarkt, Hafele. How do you pronounce Hafele? all right um so please drop me a line if you want to or follow me on twitter so uh research data shows that 37 percent of all customers start doing an online price research before buying anything by anything i mean anything like electronics laptops uh tablets smartphones, smart watches, home appliances, shoes like refrigerators, televisions, whatever you think about, people start doing an online price research because we are, I mean, sort of money is valuable. So we are looking for cheaper prices. So if we do that, the companies who are selling those products are willing to know if there are lower prices available in the market so they are looking for their competitors prices so we scrape the web and we visit more than 2 billion websites each month and scrape web collect data process that data to create information like actionable information and then we deliver them sorry we deliver them to the companies this is the actual number that i checked out yesterday night. And we have more than 250,000 e-commerce websites, different e-commerce websites in our database. We collect data from them and we report that data to more than 500 companies all around the world. So that is actually what I do. So in this presentation, I'm going to talk about our observations on the web while scraping the web. So, what we observe is actually the web is in danger. So, the web is a shared resource and like all of the shared resources, it should stand on several principles. There are two main principles that the web stands on currently. The first one is openness. The internet and full resources of it should be easily accessible by anyone who are sharing it. By anyone, I mean all individuals, all organizations, and all the companies. So, we are sharing the web. We are living on the web. We are living on a network of several items. The items contains machines, devices, software, and the individuals, pre-human beings. So, since we are sharing those resources, It should be accessible by anyone who are accessing it. The first principle is being open. The second principle is being dumb, actually. Dumb is sort of an ugly word, but the sources should be as dumb as possible with little or no management of its use patterns. Whatever you do with the Internet, the Internet should be dumb to your actions. I mean, if you want to read several Wikipedia articles, then do it. If you want to watch porn, then do it. Do whatever you want. So, it should be a dump of your use patterns. So, these are the two principles, being open and being dump. So, this brings us to the concept of dump pipes. Dump pipe concept is an analogy to city water supply system. so the city's water supply system contains several pipes and the pipes are sort of dumb they are not aware of who you are they are not aware of what you are doing with your water so you can drink that water you can like wash your dishes wash your car whatever you do you can just like waste that water it doesn't aware of it is not aware of what you are doing with the water you are getting. And it doesn't care who you are. It doesn't look for, it doesn't check your passport. It doesn't check the country that you are from. I mean, it's fully unaware of who you are and what you are doing, actually. It just provides you a steady current of water whenever you want. So that is the dump pipe analogy. Internet should use that analogy. Internet should be as close as possible to that concept, actually, but it's not. It is a slightly more complex system of delivering something. Internet is closer to smart pipes concept. So it's not carrying just water. It's carrying data. Data is sort of more complex than carrying water. Carrying data is more complex than carrying water. So, it is sort of more complex. Carrying, sorry, the pipes are slightly more sophisticated than water pipes because we have several types of network devices, several types of consumers who are consuming. Types of the consumers are also different than the consumers of water. So, internet is closer to smart pipes than dump pipes. sorry oops alright I'm gonna do it so internet when looking to when looking into the network people use use layered analogies so you know sometimes they focus on 7 layers of OSI or sometimes they focus on like four layers of network. So I propose a layered pipe concept again. The internet should be taught as in these four layers. First is the physical layer, the cables and or the wireless connections that connect you to the internet. And the second layer is the logical layer, which actually contains the core itself. All the protocols that you are working on TCP IP itself or the HTTP SSH or FTP SFTP whatever protocol that you are working on lays in the logical layer and over that layer there comes the application layer application layer itself contains the applications if you are using web browsers that is an application if you are using an email client a Spotify desktop application or your mobile applications they are all lay in this layer and the content layer represents the content type that you are consuming so if you are consuming text over wikipedia that that is represented in this layer or you if you are consuming video music whatever you are consuming the type of the content is represented in this layer so there is always something broken in each layer so whichever layer that we focus on we see that something is broken so let's focus each layer and see what are some problems of them first let's start with the top most layer the content layer most probably your traffic which you paid for which you already paid for the traffic the internet connection is shaped shaped by your internet service provider depending on according to the content type that you are consuming if you are consuming too much video then probably you'll be slowed down by your internet service provider isp because I mean, it's a cost for them, so it's not very profitable to provide you unlimited video content to you. So, in DumpPipe's analogy, it should be unaware of the content that you are consuming, but they know what you are consuming, so they shape your traffic. Probably, they slow down your traffic by the content type. So, sometimes some stakeholders of the internet, your favorite search engine, favors faster websites or favors the websites who implemented the technologies which is invented by that search engine. Google prefers the websites who are using accelerated mobile pages or progressive web apps. So if you implement AMP or PWA, you are favored and you are getting higher ranks in the search results. So it conflicts with the dumbness principle of the web. So when we come to the application layer, we have other problems. um your again your favorite search engine prefers if you use favors it favors you if you use their browser if you use chrome instead of internet explorer you get like faster faster results or sometimes they serve you different results uh if you use their browser and sometimes you throttle by the by the uh application you are using let's say you you use spotify again your internet service provider may be aware of that and then they can like shape your traffic according to that when we come to the logical layer i think me and everyone else in this room and in everyone else in the conference is uh like somehow related to this layer because we are hacking all the layers all the components in this layer uh we are discriminated by the originating ip addresses Let's say you fired up a VPN on like AWS or DigitalOcean, like $5 droplet. And then when you visit as a real user with a real browser, when you are trying to visit a website, probably you'll get blocked because most of the IP ranges of popular cloud providers like AWS, DigitalOcean, Linode, And they are all blocked by large websites because most of the scrapers and bots and harmful attacks are originated from those cloud providers. And like Best Buy, Amazon, eBay, they all block the IP ranges of them. So as a real person, you are discriminated by the IP you are using. And sometimes you are discriminated by the country that you are from. This happens mostly in the island-ish countries, let's say Australia. Since shipping something from Australia to outside is sort of expensive, they don't do that. It's very expensive. And the websites who are based in Australia does not send any item. And even more, they don't accept any orders coming outside from Australia. And even more, they don't serve you the content to outside of Australia because it's not profitable. They will not accept any orders outside from Australia. So, let's say you are an Australian and you are not here and you want to buy something to your wife, a birthday gift. And you want to visit the Australian website and you are blocked. You see a 403 error. So it sort of conflicts with the openness principle. There are broken web server configurations. Sometimes people, most of the time small websites, small web shops, leave configurations as it is. And most of the default configurations are too strict to be used on the production. So let's say you go to Amazon and search something quickly and click a link in the results and click another link. So you click, you visited three links within a minute. And I think the default configuration of light TPD thinks you are a bot because you are like doing quick actions. so you get banned like five minutes or so. So actually you are a real person but the default configuration thinks you are a bot so you are blocked. And sometimes you get different responses to identical GET requests. HTTP is a stateless protocol so a GET request should return the very same results but But sometimes you get different response to very same requests. And there is a huge fight with the bots. And there are two types of bots, good bots and bad bots. 80% of the bots are good bots, actually. So there's a minority of bots who are not good. But people take binary decisions. So they block all the bots. There's a way managing the bots, actually, robots.txt. And people are not aware of that. so people block all the bots. If you block all the bots, actually, you are in trouble because Google, Yandex, DuckDuckGo, everyone indexes you. If you block all of them, then you are in problem. Sometimes data center firewalls falsely mark you as a bad bot, and then even the website is not aware of the bot but you are all blocked uh currency handling is a problem with the dominance problem dominance principle again uh let's say you visited the u.s company and it shows you a product with the price of 200 and you say i'm a european union citizen and i want to see the prices in euros and it shows you 200 euros actually 200 euros is sort of more expensive more expensive than $200 so it is again conflicts with the dumbness of the neutrality principle of the web and there are more hacks people heavily depend on javascript instead of doing proper http usage they are like they depend on heavily javascript they discriminate you by your browser's capabilities let's say the newest version of chrome supports something and your browser does not uh so they they serve that shiny content to the newest version of the chrome but not you so you are discriminated by your browser's version uh instead of proper http redirects people do like mystery directs by javascript and there are more hacks actually people are not obeying http protocol that we contracted on so http is a contract that lays among all of us so to solve quick problems people apply quick hacks and then we are getting more and more problem actually problems actually and when we come to the physical layer again there are problems your isp is a company that should provide a cable or a wireless connection between you and the internet. But sometimes that ISPs get partnerships from different services. And then let's say you use Verizon to connect the internet and Verizon, let's say, get a partnership with Hulu. And then if you use Amazon Prime Video, they're aware of it. So they are trying to slow it down to make you use Hulu. So it again conflicts with the dominance principle. Device neutrality is a If you visit like flight ticket websites and if you are using MacBooks they think you are slightly richer so they may show you higher prices. Again device neutrality lays on the physical layer. So these are several problems that we observe actually. We have one link, we visit that link and we see like those problems at all. So, what to do? Actually, since we have two principles, we should obey those two principles. Let's try to be as open as possible. I mean, you are providing some resources, you are adding some resources to the internet and make that resource fully accessible by anyone, any individual, any organization, any company, by any company. So, make it accessible for everyone who are sharing the web. Since the web is a network of items, and we are not the first species living on a network of items. Animals live in forests. Forest is a concept representing a more sophisticated network of items and resources. So, we are trying to achieve, live on those shared resources. So, let's be as open as possible. And the second one, be cooperative. Again, to share those resources, we should be as cooperative as possible. And lastly, be slightly more dumb. So, be dumber, please. Instead of getting clever hacks, instead of getting clever workarounds, let's try to get slightly dumber and obey the protocols that we contracted on. so be open be cooperative and be dumb actually that is the end of my talk so thank you for listening to me thank you if you want to make further reading on that net neutrality and the semantic web are two keywords uh you can like uh catch them on wikipedia if you have any questions any questions on the call

Speaker 2 [19:49]

Thank you for this talk. Thank you. Basically, I want to ask you about what's the reason for some content creators to be open for bots? I mean, when Google parses you or Yandex or somebody else, this is a win-win situation because they parse you and they show your content to other users. But when, for example, other companies try to parse your content and use it for their on their own purposes benefits yeah thank you I don't see the reason why I should be open for for that kind of companies

Speaker 1 [20:44]

Yeah, that's a very correct approach. You may want to block some bots accessing to your content, and there is a proper way of doing that. So I suggest you to block them in a proper way instead of fighting blindly. So use the first robots.txt and this all of that bots say that you are not allowed to enter here. If they don't obey robots.txt, then there comes the other ways. so you may want to block someone to access your content because it's your content in e-commerce domain there is a huge potential on the competition so people wants their prices to be seen everywhere else so in our domain people wants to share their prices so it's a sort of an advertisement it's sort of an announcement so if they have cheaper prices they want them they want everyone else to see that In your case, you may want to block them, of course, and then there's a proper way of doing that. But people do not apply that proper way. Instead, they're fighting blindly. I suggest that. Okay, thank you. Thank you. Okay, come in.

Speaker 3 [22:09]

Thanks. Thanks for the good talk. I've been on the receiving end of these bots on our legacy interfaces. And we've had the problem that we can only fight back badly. So I see the problem that even if 1% of the bots programmers do not follow these nice instructions, they will bring down your system. So the only chance we have actually to block them is to fight back blindly and randomly. So it would be a nice idea if we wouldn't have to do it. But as a comparison, for example, we got 38,000 requests per hour on an API that can do four requests per second. And we have 20 real requests. So the only thing we have to do to fight this is to make educated guesses on the people who try to hide very hard and then fight back badly. so I see no chance to change it or do you have suggestions how to do that?

Speaker 1 [23:05]

to do that yeah i understand that if you if if you are like if they are getting your systems down then there is nothing to do i mean you need to be alive so i understand that so there's always false positives like uh since you are blocking all of them all of the requests maybe you can do like whitelisting instead of like making it if you are talking about the apis maybe you can do white listing I don't know I'm not aware of the whole system that you are suggesting but

Speaker 3 [23:36]

Yeah, what we did is we analyzed the kind of requests and then made educated guesses if this could be a real person or not. So if you're programming bots, ask in English and we'll block you more.

Speaker 1 [23:47]

Yeah, yeah. I mean, there are like bot lists and they are like keeping the list of good bots and bad bots. So maybe you can do whitelisting the good bots and then restricting all the others. I understand that. I mean, if you are getting down, if you're like becoming down, I understand that. And I don't have any other suggestions. I mean, just do the whitelisting thing and kill them all. Kill the rest. Okay, any more questions? um hi i hope everybody forgives me for asking two questions one of them will be like really short um do you scrape any web pages where data is rendered on something like a canvas element where you cannot you know just parse structured data and more proper question what is like the worst page to scrape ever how would you describe it oh yeah uh we don't scrape canvas anymore we used to do that we used to scrape images at first and then like making ocr and then try to get the prices but nowadays since mobile traffic is increasing people are reducing their content size so instead of serving the prices in images they are like trying to make it more mobile accessible so we don't do any canvas scraping or image scraping anymore we used to do that and the worst is uh there's a framework i don't remember the name maybe open card there's a framework that you can use and in older versions of open card there is something broken there's a i guess but it's very largely deployed so there are more than like 100 e-commerce websites are using them so whenever we visit them they render completely wrong prices the content is wrong actually but they are wrong if you are visiting them not as a boat but as a real person you see the wrong prices let's say you are looking for a like five bucks uh phone case and they're like they they may show you the price of an iphone so i mean that was the uh problem because it is inconsistent and hard to track hard to find um that is the like hardest problem that i as far I remember and there is like a few lines of code just written for to catch that type of websites in our code base so if if the version is open card something something to something there's a like there's a code block in the in our code base. Thanks. Okay we have time for more questions. Looking at the other side what are the actions that you are taking on your side to get around blocking and And defense is on the website.

Speaker 3 [26:56]

the websites that you're scraping.

Speaker 1 [26:58]

you're scraping or you try to scrape i mean if you understand uh that you the website does not want us uh we don't do any workarounds i mean in in in any indication uh if they indicate that they don't want us we don't try to find any work workarounds but if we think that uh we are blocked by accident we are blocked by the data center we are blocked by like physical region then we try to find ways we use proxy servers we use like visiting them as as real browsers as possible I mean mimicking the real browsers we have several like workarounds but other than that if they somehow indicate that they don't want us then okay we are not visiting them so we are a good board in terms of cloud Okay, we have time for one last question. There's one over here, I guess. We have a website, and we also have a lot of bots trying to scrape our data, so we put in the robots file that if they want our data, they should just email us, and then for a small price that covers covers the extra cost of our cloud computing they can get an api key and that's perfect but we see that actually almost nobody reads this robots.txt file do you have like a better way to communicate this to bots or whatever um i mean is there a particular footprint of that bots do you know which bots are accessing your data people and so we have a footprint i mean we are not hiding so we have a footprint uh we if we visit you uh you will be aware of us i mean we are not changing and we have a particular user agent so we are not like hiding actually uh so if some websites are providing us an api that is the perfect solution because that's a contract so it doesn't change if we scrape the web when we scrape the web the two percent of whole websites changes in two months so in two months we need to uh change our two percent of our scraper engine code base because the websites are changing so if the website provides us an api that's the perfect solution because that's a contract and we will be aware if that contract changes so if we can find an api we use that api Other than that, we don't, like, we obey robots.txt completely. If you don't allow us, we are not visiting you. I'm not aware of any communication channels about the bad bots. There are, like, forums where people are trying to scrape the websites. So, please drop me a line or I will find you. If we scrape you, we can use your API. At least I can do that. Thank you, I guess. Thank you for coming to my talk. It's nice to be here. Thank you very much.

Samet Atdag

About — in the speaker's own words

I'm Samet Atdag, a seasoned developer, the co-founder and CTO of Prisync. Prisync is a startup focused on information retrieval and data processing in e-commerce domain. I develop systems for crawling a large portion of the web.

I'm the organizer of Python Istanbul user group. With more than 8000 members, Python Istanbul is one of the largest user groups of Europe.

Social card for talk: What we learned from scraping 1 billion webpages every month