Could The brand new Google Spider Be Creating Concerns With Web-sites?

Aug 7, 2012 by erikfahlstedt04

Around some time Google declared “Big Daddy,” there was a brand new Googlebot roaming the internet. Considering the fact that then I have heard tales from clients of web sites and servers going down and formerly unindexed content receiving indexed.

I started digging into this and you’d be surprised at what I located out.

1st, let’s look in the timeline of events:

In Late September some astute spider watchers over at Webmasterworld spotted unique Googlebot exercise. In actual fact, it absolutely was within this thread: webmasterworld/forum3/25897-9-10.htm the bot was very first noted on. It involved some posters who thought that maybe this could possibly be frequent customers masquerading because the popular bot.

Early on in addition, it appeared that the new bot wasn’t obeying the Robots.txt file. This can be the protocol which will allow or denies crawling to parts of the web-site.

Speculation grew on what the new crawler was until Matt Cutts talked about a brand new Google check data heart mattcutts/blog/good-magazines/#comment-5293. For those that don’t know, Matt Cutts is usually a senior engineer with Google and one with the handful of Google staff speaking to us “regular folk.” This refer to occurred in November.

There was not much mention of Massive Daddy till early January of this year when Matt once more blogged about this asking for suggestions. mattcutts/blog/bigdaddy/

A lot suggestions was provided about the accuracy in the results. There were also these that requested if the Mozilla Googlebot (known as “Mozilla/5.0 (compatible; Googlebot/2.1; +google/bot.html)” in your visitor logs) and Large Daddy had been connected, but no response was produced.

Now I am likely to begin some of my own speculation:

I do in fact think the two are connected. In truth, I feel this new crawler will eventually replace the previous crawlers just as Large Daddy will change the present data infrastructure. textlinkbrokers/blogs/comments/310_0_1_0_C/

Why is this significant?

Depending on my observations, this crawler might have the ability to do a lot much more compared to old crawler.

For one, it emulates a more modern browser. The previous bot was based on the Lynx text based browser. Even though I’m positive Google added capabilities as time went on, the fundamental Lynx web browser is simply that -basic.

Which explains why Google couldn’t handle items like JavaScript, CSS and Flash.

On the other hand, with the new spider, built about the Mozilla motor, you will find a lot of opportunities.

Just appear at what your Mozilla or Firefox browser can perform alone -render CSS, study and execute JavaScript along with other scripting languages, even emulate other browsers.

But that is not all.

I have talked to some of my clientele and their sites are finding hammered by this new spider. It’s gotten so negative that a few of their servers have gone down as a result of the quantity of traffic from this one spider!

About the additionally side, I’ve customers who went from a few hundred thousand indexed pages to more than 10 million in only a handful of weeks! Literally due to the fact December, 2005 there is been a 3500% boost in indexed pages more than an 8 week period! Just so you know, this really is also the client’s site that went down due to the huge volume of crawling occurring.

But that’s nevertheless not all.

I’ve a different client which makes use of IP recognition to serve content based on a person’s geographic location. Should you live within the US you will get American content material and pricing; in the event you live within the United kingdom you receive Uk content and pricing. As you may envision, the United kingdom, US, Canadian and Australian content material is all quite similar. In fact about the only thing noticeably different would be the pricing facet.

This can be my problem -if the replicate content material will get indexed by Google what will they do? There is an excellent likelihood the website would be penalized or perhaps banned for violation in the webmaster high quality guidelines set forth by Google here: google/webmasters/guidelines.html#quality

This is why we implemented IP recognition -so that Googlebot, which crawls from US IP addresses only sees one model with the site.

Even so, a critique of the server logs displays that this new Googlebot is going to not just the US content but also the content material in the other sections from the web-site. Naturally, I desired to confirm the IP recognition was operating. It is actually. This leads me to speculate then; can this web browser spoof its place and/or use a proxy?

Consider that -the web browser is intelligent enough to complete a few of its personal screening by viewing the website from several IP addresses. If that’s the case then those who cloak websites are going to have issues.

In almost any case, from your limited observations I’ve produced, this new Google -both the data heart and the spider -are likely to modify the way we do issues.

Authorised Michael Vick Jersey Online Store brings all sorts of inexpensive LeSean McCoy Jersey straight away with Super fast Shipping and delivery, Secure Payment & Superb Customer Support.