# Web scraping tools

**URL:** <https://forum.elixirforum.com/t/web-scraping-tools/4823>\
**Category:** Questions\
**Tags:** web-scraping\
**Created:** [April 27, 2017, 8:40pm UTC](https://forum.elixirforum.com/t/web-scraping-tools/4823 "2017-04-27T20:40:10Z")\
**Posts on this page:** 1\
**Showing post:** 28

<div class="post-metadata">

**Author:** ![adammokan](https://forum.elixirforum.com/user_avatar/forum.elixirforum.com/adammokan/32/26407_2.png) [@adammokan](https://forum.elixirforum.com/u/adammokan)\
**Post date:** [September 12, 2019, 2:09am UTC](https://forum.elixirforum.com/t/web-scraping-tools/4823/28 "2019-09-12T02:09:04Z")

</div>

> Can you clarify on that? I’ve heard sites like Amazon deliberately give you wrong prices if they detect a bot, is that true for them and others? Or what kind of trash data?

Yes. I don’t want to speak for specific sites - but I can tell you many ‘popular’ sites will start delivering results that are either not ordered correctly (_if you think in terms of SERP rank data as an example, where order matters to buyers_), E-commerce sites will start throwing bogus prices, and so-on. The only real way to protect/test this is to do manual A/B comparisons. The other thing to consider is we’re fully immersed in a world of personalized content - so absolutely it gets difficult, even if not being thrown bogus data, is that your clients will think your results are wrong because they are viewing personalized content if they manually compare results.

> Additionally, how do you even test your bot against tools like Panopticlick at all? Have your bot GET their root page and click the “Go” button? Is that what you meant, or do they (or others) have a dedicated bot testing toolkit?

I should have been more specific on that. I mostly reused the same boilerplate headless crawl scripts and would include the site-specific nav/logic separately. So this meant all of my ‘pre-crawl’ setup to plug gaps like setting a legit `navigator.platform` and making sure `navigator.webdriver` returns false, etc (_there are quite a few of these you need to cover_).

Anyhow, I would traditionally include that initial setup logic and then have a custom script that would navigate the Panopticlick site, run the full checks, screenshot the results, and study later. Just to make sure I wasn’t missing something obvious. So yes, click the go button - load the full results and screenshot the page.

Aside from that there are many other “how private is my browser” checks out there that test for hardware-level info that may be worth mocking on some targets.

This reminds me of the fun I had with geolocation/gps coords - again, depending on what you are going for. Most common case for that effort was a well-known map site. Some hints for geolocation - and this may be outdated, but it was always important to mock both `navigator.geolocation.getCurrentPosition` as well as `navigator.geolocation.watchPosition` to spoof lat/long. A little bit of noise in the `coords.accuracy` attribute went a long way here 👍

---

_[View the full topic](https://forum.elixirforum.com/t/web-scraping-tools/4823)._
