Thanks for your comment.
I would not say that web scraping is simple. Everything depends on the scale. Probably for very small sites you can use curl or wget. But I am talking about cases when you need to scrape millions of pages. And in these cases, things are way more complex, as you have to solve concurrency and resource management problems. Also, finding a good strategy of crawling a million pages is a challenge (also consider cases when URLS are dynamically generated)!
To answer your comments:
- Server-Side Rendered page
easy to scrape with curl or any HTTP Client (+ HTML Parser)
Will not scale. Also, you will have to avoid visiting pages twice and filtering out duplicates
- Javascript Rendered page
A. You can use Headless / Browser Automation, but, it will be slow (+ HTML Parser)
In most of the cases, you will be able to find how a web page (e.g. a product page) is fetching data from API, so in most of the cases, you don’t need selenium.
B. Do “little bit” Reverse Engineering on their Web API (FASTER)
Unfortunately, this does not work. Most of the web sites do not have API. And those who have, would not provide a full and up to date data. Even more, some of the APIs are just horrible and can’t be used for data extraction.
Important Point :
- Make sure your scrapper support Proxy Usage
It does
- If your site target has anti-scraper / crawler / bot (Like your bot follow pagination, 1->2->3 and so on) and it block your IP, you can use IP Rotation Service like geosurf.com and luminati.io
You’re right. But please take into that proxies management is a complex stand-alone task. There are some quite advanced systems which allow overcoming bans with proxies, and I was developing one of them in the past.
Also nowadays in some cases, it’s just not enough just to perform a request through another proxy, as the most advanced system would also perform 3-4 levels of request fingerprint analysis. With this regards, I would suggest looking at Crawlera
- In some Country / Site, Web Scraping are prohibited
Well.. is that correct to assume the internet is also prohibited in these countries? Please take into account that no search engine can work without web scraping. And I don’t see the web without search these days. (But it’s just an opinion).


















