Sunday, March 21, 2010

Why we need HTTP Compression


HTTP compression isn't something you put in your code. Instead, check with your sys admin to see if it's installed (or can be) on the server that's dishing out your site's pages.
What does it do? HTTP compression accelerates the transmission of pages from server to surfer. It allows for server-side HTML compression so server apps (like Apache or Microsoft Web Server) can compress the source code of your page before sending it out over the wires.

HTML compression works on almost every browser these days, but the code is savvy enough to dish out un-compressed files to unsupported browsers. Compression alone improves the download size of pages by up to 300%. (Exclamation mark.)
Other nuggets in the HTTP compression nougat include persistent connections (between server and client) and pipelining (which allows the server to rattle off files without waiting for a client's "uh-huh, got that one, next please" response). Crackin' good stuff, all of it. The only negative thing we can say about HTTP compression is that it's probably already installed on your server, seeing how it's well-adopted now (for obvious reasons). Still, if you suspect your hosting company is slow-with-the-program, it couldn't hurt to ask.



Link Prefetching

Link prefetching is a feature in some browsers which makes good thrift of a browser's idle time by downloading files that aren't on the current page but might be needed a page or two down the road. Don't worry - link prefetching doesn't slow pages down. The extra downloads don't kick in until after the current page has finished loading and the browser doesn't have any pressing matters to attend to.
If you had a photo gallery, for example, with Previous and Next navigation links by each picture, you could add a prefetch link tags pointing to both of those destinations. That way, while the user is staring at one amazing travel picture, two other pages are downloading in the background. If the user does click either the Previous or Next links, those pages will already have been downloaded, and the content will display instantaneously. More technically speaking, "Guy clicks the link, see, Bada-bing, Bada-boom, it's the next page, already. Aow!"
Link prefetching isn't automatic - it relies on you, the developer, to encode hints as to what files are likely to be needed next. (Browsers aren't prescient, but good webmasters understand a site's traffic flow.) You can encode prefetch hints in either the HTTP header or the page's HTML. There are a number of allowed variations on how to code a prefetch, but a simple HTML  link¢ tag with a relation type of "next" is small, easy, and our favorite. Like these:

Cache In

Network (or Proxy) Caching
We previously discussed how browser-side caches store commonly-used images on the users' hard drives, but it's important to note that similar caches exist all alongside the highways and byways of the internet network. These "Network Caches" make websites appear more responsive because information doesn't have to travel nearly as far to reach the user's computer.
Some webmasters are leery of network caches. They worry that remote caches might serve out-of-date versions of their site - an understandable concern, especially for sites like blogs that update frequently. But even with a constantly-updating site, there are images and other pieces of content which don't change all that often. Said content would download a lot faster from a nearby network cache than it would from your server.
Thankfully, you can get a site "dialed in" pretty nicely with just a basic knowledge of cache-controls. You can force certain elements to get cached days on end while keeping other elements from being stored at all. META tags in your document won't cut it. You'll need to configure some HTTP settings to make caching really work for you.
Every Bit Counts
Alrighty. So you've done the big stuff - dropped bit depths on your every PNG, cranked up the HTTP compression, and taken a (metaphorical) weedwhacker to your old, convoluted table layouts. Yet you're still obsessed with how small, how fast, and how modem-user-friendly you can make your site. Ready to jump into some seriously obsessive-complusive optimization?
You know those TV commercials where they zoom in on a supposedly "clean" kitchen counter, only to reveal wee anthropomorphic germ-creatures at play?
Well, you can similarly clean every extraneous detail from a site's layout, and still have some nasty, nasty cruft living in the source code. What's the point of novel-length meta keyword lists and content tags? C'mon, do you still believe that search engines care about that stuff? Not in this millenium. You'll get better search referrals by thinking carefully about the real content on your pages and building an authoritative site that's linked to widely.
Streamlining the  head¢er section of unneeded meta keyword/author/description content, and likewise junking giant scripts makes a bigger impact, kilobyte-per-kilobyte, than sacrifices made elsewhere on the page. Having a short  head to ¢ to your document ensures the initial chunks of data the user receives contain some "real" content, which gets displayed immediately. That's another notch for "perceived speed" improvements.
Of course, there are plenty of regular  body¢ bytes still worth tossing. Start with HTML comments, redundant white space, and returns. Stripping all these invisible items from your source code yields extra kilobytes of space savings on the average.


URL Abbreviation

Ever spot how links on the Yahoo frontdoor are generally just a few characters long? Go to the site and move your mouse over some of the news links near the top. You'll see they all start with http://yahoo.com/s/ and then list a string of six numbers. Links, generally, run on much longer than that, especially if they include redirect codes or CGI variables. Put enough normal links on a page, (viz. the Yahoo frontdoor, again) and a sprinkling of kilobytes (and seconds of download time) is also added to the code.
So what's Yahoo doing with those funny links, anyway? They're abbreviating their URLs, using the mod_rewrite Apache module, so that a link like "/s/882142" redirects to "mysite.com/content/unregistered/News". Implementing this requires getting your hands dirty with some server configuration. Specifically, you need to get mod_rewrite installed and poke around with the srm.conf file. Dirty work for many of us, but the payoff is worth several solid kilobytes on a link-heavy page.



Monday, March 1, 2010

Keyword Effectiveness Index

The Keyword Effectiveness Index (KEI) compares the Count result (number of times a keyword has appeared in  data) with the number of competing web pages to pinpoint exactly which keywords are most effective for your campaign.
In a nutshell: Look for the keywords near the top. The higher the KEI, the more popular your keywords are, and the less competition they have. Which means you have a better chance of getting to the top.
The article below is a much more detailed look at the KEI and why we have decided to use it.


DETAILED EXPLANATION

The KEI is a measure of how effective a keyword is for your web site. The derivation of the formula for KEI is based on three axioms:

1) The KEI for a keyword should increase if its popularity increases. Popularity is defined as the number of hits coming. This axiom is self-explanatory.

2) The KEI for a keyword should decrease if it becomes more competitive. Competitiveness is defined as the number of sites which a search engine e.g. AltaVista displays when you search for that keyword using exact match search.

Exact match search means that a search engine searches for only those sites, which use the keyword exactly as typed in by the user. It is the equivalent of entering: It Basically used by " "

“beach wedding dresses”


Partial match search means that a search engine also searches for sites which contain the individual words of the keyword but not necessarily occurring together or in the order typed in by the user. It is the equivalent of entering:
beach wedding dresses


Partial match search presents a distorted picture of the competitiveness of a keyword because when you optimize your site for a particular keyword, you are actually competing with sites which have used the keyword exactly as typed in by the user.

So to clarify, competitiveness is defined as the number of sites which a search engine displays when you search for that keyword using exact match search, that is with quotes surrounding the term. Rather than those web sites returned when entering the phrase only partially, that is without quotes.

Note: When you select KEI Analysis, quotes will be added temporarily to each of your search terms for the purposes of the search.


3) If a keyword becomes more popular and more competitive at the same time such that the ratio between its popularity and competitiveness remains the same, its KEI should increase. The rationale behind this axiom requires a more detailed explanation. The best way to do this is to take an example:
Suppose the popularity of a keyword is 4 and AltaVista displays 100 sites for that keyword. Then the ratio between popularity and competitiveness for that keyword is 4/100 = 0.04.
Suppose that both the popularity and the competitiveness of the keyword increases. Assume that the popularity increases to 40 and AltaVista now displays 1000 sites for that keyword. Then the ratio between popularity and competitiveness for that keyword is 40/1000 = 0.04.

Hence, the keyword has the same ratio between popularity and competitiveness as before. However, as is obvious, the keyword would be far more attractive in the second case. If the popularity is only 4, there's hardly any point in spending time trying to optimize your site for it even though you have a bigger chance of ending up in the top 30 since there are only 100 sites which are competing for a top 30 position. Each hit is no doubt important, but from a cost-benefit angle, the keyword is hardly a good choice. However, when the popularity increases to 40, the keyword becomes more attractive even though its competitiveness increases. Although it is now that much more difficult to get a top 30 ranking, spending time in trying to do so is worthwhile from the cost benefit viewpoint.
A good KEI must satisfy all the 3 axioms. Let P denote the popularity of the keyword and C the competitiveness.
The formula that we have chosen is KEI = (P^2/C), i.e. KEI is the square of the popularity of the keyword and divided by its competitiveness. This formula satisfies all the 3 axioms:

i) If P increases, P^2 increases and hence KEI increases. Hence, Axiom 1 is satisfied.
ii) If C increases, KEI decreases and hence, Axiom 2 is satisfied.
iii) If P and C both increase such that P/C is the same as before, KEI increases since KEI can be written as
KEI = (P^2/C) = (P/C * P). Since P/C remains the same, and P increases, KEI must increase. Hence, Axiom 3 is satisfied.
Note that the formula for KEI is not unique. In fact, this is one of the nice things about the KEI. If, instead of using 2, you use any power of P greater than 1, the resultant formula will also satisfy the 3 axioms. For example, (P^1.5/C) and (P^3/C) both satisfy the 3 axioms. The exact power of P that you choose depends on how much emphasis you want to give to the popularity of a keyword viz-a-viz its competitiveness. Higher the power of P in the formula, higher will be the emphasis on popularity. If you are very confident about your search engine positioning skills, choose a higher value for the power of P. If you are not that confident about your search engine positioning skills, choose a lower value for the power of P (but the power should still be more than 1). Thus, the KEI can be adapted to your skill level! Feeling confused as to which power you should choose? Stick to 2. It maintains a nice balance between both popularity and competitiveness.

Sunday, February 28, 2010

Why Ping is necessary

The last six months has seen a massive rise in content theft blogs and spam blogs, and there’s one thing these blogs usually have in common, and that’s the whole “Blog and Ping” thing, but if you don’t know what Blog and Ping is, don’t feel bad, because most people don’t.
But before I start a word of advice: don’t do it. Knowing and understanding your enemy is important in formulating ways of overcoming and defeating them. People who create these blogs are leeches who deserve nothing less than being banned from the search engines they so desperately seek to be included in.

So what is Blog and Ping?
I’ll give a short answer and a long answer because Blog and Ping comes in a few different flavours.
The Short Answer: Blog and Ping is a online marketing term applied to a system that utilizes blogs and pings to deliver content and/ or sites for indexing in search engines with the ultimate aim of profit.
But that doesn’t really explain a lot, because basically that’s what blogs already do.
Background
For those of you too young or who never got into writing more traditional web pages, basically getting your webpage into Google, Yahoo or other search engines has always been somewhat difficult. The time it takes for your static web site to be indexed by Yahoo for example after you submit it to them without taking up the paying option is around 6 weeks, and sometimes longer. Blogs changed the rules of indexing because where as you use to have to wait for the search engines to index you, all of a sudden bloggers could wave a big red flag with the words “I’m over here” written on it and the search engines would come. Pinging a central server such as Weblogs.com meant that a central list was available of blogs that were posting at a given point in time, and the search engine spiders then followed the links on these lists and bingo: your blog gets indexed.
The vultures circle
When your on a good thing, people notice, and the vultures started to notice blogs getting a really good run in search engines, and they pretty quickly worked out why….introducing Blog and Ping.
Flavours of Blog and Ping
Blog and Ping comes in different flavours and variations to the theme. Each Blog and Ping promoter has a different sales pitch that also states that their flavour is the best. The commonality to all of them is that they involve pinging sites such as Weblogs.com as a means to deliver the search engine spiders to a blog. But this is where things get different.
Real blogs, stolen content
One strain of blog and ping promotes blogs as being a great way to make money from Adsense and affiliate programs with scripts that steal content from other blogs and repost it in an attempt to create a legitimate looking site that gets traffic from search engines.
Michelle Timothy’s RSStoBlog.com is a leading pusher of this flavour of Blog and Ping. To quote Ms Timothy:
” Imagine for a minute that you had a tireless assistant that worked hard night and day finding fresh, relevant content for you to post to your blog. Now also imagine that this tireless assistant not only finds this relevant, keyword specific content for you but also posts it to your blog at exactly the times you want it posted … and they do this day in and day out, until you beg them to stop!”
I’ve not provided a link to the site, but by all means cut and paste the URL into your browser, because the people using it are stupid enough to give testimonials as well. Their example WordPress blog also includes a stolen story to the Blog Herald as well.
Spam blogs
I’ve split spam blogs into a separate category because the stolen content blogs can in reality function and look like real blogs if they are done properly, and many people would be unable to tell the difference. Spam blogs on the other hand stand out like a sore thumb, and these are another flavour of Blog and Ping. The theory with these sites is essentially a new form of link farms, in that they are never really created for viewing by the general public, but as a way for the search engines to discover other sites, and to reward those sites for multiple links. A three step process: the search engines spider finds your spam blog ping at Weblogs.com, follows it back to the blog, then discovered links to static web pages and then goes through to index them. The difficult thing of course is without content the search engine spiders won’t visit, so these sites create all sorts of rubbish as content, often with keywords scattered throughout posts, to assist the spider going onto index the money making static web site.
There are a number of sites promoting this, and each one usually has a slight variation on this theme. BloggingEqualizer.com states the following on their method:
The technique consists of 4 steps:
1. Build a search engine “spider trap” by creating a free blog on Blogger.com (they can’t resist freshly updated content).
2. Grab a free account at MyYahoo.com.
3. Post links to the Web pages you want spidered inside your blog.
4. Then, just call the search engine spiders to dinner by sending (or “pinging”), with the click of a button, the blog to your MyYahoo account.
Sites like Instantblogandping.com promotes a version which involves automatically reposting your own static webpages to Blogger with links back to their source.
Is there money to be made in Blog and Ping?
Yes. The same way as there is money to be made in Amway, because Blog and Ping is really just another variation on the old multi level marketing theme without as many circles on a whiteboard; the only people who make money are the promoters and the people at the very top of the pyramid.
First and foremost the promoters are most likely raking it in. Most of these programs/ scripts are on the market for between $100 and $500 USD, and guess what: they wouldn’t be offering them if there weren’t thousands of suckers out there who’d fall for the slick marketing spiel and promises of automated riches and happily put their hands into their pockets to buy them. They say there is a market for everything, and with the internet being so large and with so many people using it this also holds true for Blog and Ping programs.
Some of the earlier users would have made some money out of Blog and Ping as well, in around the middle of last year when spam blogs were still relatively unknown and those using Blog and Ping techniques would have been few and far between. Today, with literally millions of spam and content theft blogs out there it would be nearly impossible for anyone to even earn enough money back to cover the cost of Blog and Ping package they’ve bought if they are looking to make money off their blog.
In terms of SEO strategies the honest to god truth today is that if you’re looking only to get your static site indexed into one of the big search engines Blog and Ping actually does work, but only in the same way that a few text links from a few decent blogs would deliver the same thing, and you’ll pay a lot less for a proper text link from say here at the Blog Herald then you’ll pay for most of these programs.
The end is nigh
Already some in the SEO industry are saying that Blog and Ping is dead due to the massive increase in users, content theft sites and spam blogs. If you’re getting any benefit out of Blog and Ping now, you won’t be for much longer because already some search engines are talking about excluding your sites.
Is there any good in Blog and Ping
Yes, but in the same way that there is good in porn, because the pursuit of money often drives technological change. Services such as reblog have potential to be used for real meta-blogging in the same way you can set up a link blog at Bloglines today. The challenges presented by people using Blog and Ping strategies will force search engines and others to find new ways of filtering the rubbish out and that will be a good thing for the blogosphere.
Note to Google and the Blogger team: sooner rather than later please.

Monday, February 15, 2010

How to recover virus infected site which is banned by Google

The number of sites affected by malware/badware grew from a handful a week to thousands per week.You can do from Webmaster Tools which provides malware reviews.

If you find that your site is affected by malware, either through malware-labeled search results or in the summary for your site in Webmaster Tools just do following step.
  1. View a sample of the dangerous URLs on your site in Webmaster Tools.
  2. Make any necessary changes to your site according to StopBadware.org's Security tips.
  3. New: Request a malware review from Google For evaluate your site.
  4. New: Check the status of your review.
    • If Google feel the site is still harmful, then Google will provide an updated list of remaining dangerous URLs
    • If Google determined the site to be clean, you can expect removal of malware messages in the near future (usually within 24 hours).

    We encourage all webmasters to become familiar with Stopbadware's malware prevention tips. If you have additional questions, please review this documentation or post to the discussion group. We hope you find this new feature in Webmaster Tools useful in discovering and fixing any malware-related problems, and thanks for your diligence for awareness and prevention of malware.

    Friday, February 12, 2010

    What is the Google Dance?

    Approximately once a month, Google update their index by recalculating the Pageranks of each of the web pages that they have crawled. The period during the update is known as the Google dance.
    Because of the nature of Page Rank, the calculations need to be performed about 40 times and, because the index is so large, the calculations take several days to complete. During this period, the search results fluctuate; sometimes minute-by minute. It is because of these fluctuations that the term, Google Dance, was coined. The dance usually takes place sometime during the last third of each month.
    Google has two other servers that can be used for searching. The search results on them also change during the monthly update and they are part of the Google dance.
    For the rest of the month, fluctuations sometimes occur in the search results, but they should not be confused with the actual dance. They are due to Google's fresh crawl and to what is known "Everflux".

    Google has two other searchable servers apart from www.google.com. They are www2.google.com and www3.google.com. Most of the time, the results on all 3 servers are the same, but during the dance, they are different.
    For most of the dance, the rankings that can be seen on www2 and www3 are the new rankings that will transfer to www when the dance is over. Even though the calculations are done about 40 times, the final rankings can be seen from very early on. This is because, during the first few iterations, the calculated figures merge to being close to their final figures. You can see this with the Pagerank Calculator by checking the Data box (top left) and performing some calculations. After the first few iterations the search results on www2 and www3 may still change, but only slightly.
    During the dance, the results from www2 and www3 will sometimes show on the www server, but only briefly. Also, new results on www2 and www3 can disappear for short periods. At the end of the dance, the results on www will match those on www2 and www3.
    This Google Dance Tool allows you to check your rankings on www, www2 and www3 and on all of data centers simultaneously.

    Google currently has 12 data centers, any one of which can provide the Toolbar PageRank of any page. As the dance progresses, these data centers are updated one by one. Before the dance begins, they all return the same, current PageRank value for a given page, but during the dance they are updated, one by one, to the new PageRank value. Checking each of the centers during the dance reveals the new PageRank values as they gradually spread through the centers. If the PageRank isn't going to change, the centers show the same values throughout, of course.
    Querying the data centers
    For this, it is necessary to have the Google Toolbar installed and the PageRank indicator on. Every time a page is received by the browser, the Toolbar requests its PageRank from one of Google's data centers. The information is returned as a one-line text file and stored in the Temporary Internet Files folder.
    The Toolbar's request URL includes the URL of the page that it wants the PageRank for (the target page), and a checksum that matches that URL. Of course, the checksum must match the target page's URL.
    A fat URL for a typical Toolbar request (all in one line):-
    http://216.239.33.102/search
    ?client=navclient-auto
    &ch=5150615727
    &features=Rank:FVN
    &q=info:http%3A%2F%2Fwww%2Eexampledomain%2Ecom%2F

    If you copy and paste that fat URL into your browser, you will get Google's "forbidden" page back. That's because the target page and checksum don't match - it's just an example of the request URL.
    Notice that the target page is in escaped format - some of the characters are represented by hexadecimal codes (e.g. %2F).
    To get the new PageRank for a particular page, you need to make the same request that the Toolbar makes for it. I.e. you need the fat URL that the Toolbar uses. And you need to request the PageRank from all of Google's data centers. The method is a bit long-winded but it works. Here's how to do it:-

  1. Use your browser to browse to the page. This makes sure that the page and the Toolbar's PageRank request are in your Temporary Internet Files folder. You only need to do this once - not every time.




  2. Open the index.dat file from the Temporary Internet Files folder into a text editor, and perform a search in it for the target page. You'll find the entire fat URL, similar to the one above, for the Toolbar's PageRank request. NOTE: Because the target page is escaped in the fat URL, search only for an unescaped part; e.g. "exampledomain".




  3. When you've found the fat URL, copy and paste it into your browser's address box and press Return or click Go. If the page is in Google's directory, the returned line includes the directory path. The last element in the first part of the line is the Toolbar PageRank value for the target page. To see the page's new PageRank spread across the centers during the dance, use the same fat URL, but replace the IP address with each of the data centers. This is also a good way to see the progress of the dance in general.
    Data centers

    216.239.33.100 :: www-ex.google.com
    216.239.35.100 :: www-sj.google.com :: currently offline
    216.239.37.100 :: www-va.google.com
    216.239.39.100 :: www-dc.google.com
    216.239.41.100 :: www-fi.google.com
    216.239.51.100 :: www-ab.google.com
    216.239.53.100 :: www-in.google.com
    216.239.55.100 :: www-zu.google.com
    216.239.57.100 :: www-cw.google.com
    216.239.59.100 :: www-gv.google.com
    66.102.11.100 :: www-kr.google.com
    66.102.7.100 :: www-mc.google.com
    TIP: If you want to check the same pages during future dances, save the fat URLs into a text document so that you don't need to go through the process of finding them in the Temporary Internet Files folder each time.





  4. Google Dance - The Index Update of the Google Search Engine

    The name "Google Dance" has often been used to describe the index update of the Google search engine. Google's index update occurred on average once per month. During an index update there was significant movement in search results and Google showed new backward links for pages. However, in mid-2003 Google started to update it's index continuously. It appears that, still, there has to be an update of the complete index once in a while and during this time new backward links are shown. But, because of the continuous update, the effects on search results seem to be rather insignificant.
    We will keep this site up running because it provides some information beyond the Google Dance. But there will no longer be a monitoring of updated data centers during a "Dance".
    The Technical Background of the Google Dance
    The Google search engine pulls its results from more than 10,000 servers which are simple Linux PCs that are used by Google for reasons of cost. Naturally, an index update cannot be proceeded on all those servers at the same time. One server after the other has to be updated with the new index.
    Many webmasters think that, during the Google Dance, Google is in some way able to control if a server with the new index or a server with an old index responds to a search query. But, since Google's index is inverse, this would be very complicated. As we will show below, there is no such control within the system. In fact, the reason for the Google Dance is Google's way of using the Domain Name System (DNS).
    Google Dance and DNS
    Not only Google's index is spread over more than 10,000 servers, but also these servers are, as of now, placed in 13 different data centers. These data centers are mainly located in the US (i.e. Santa Clara, California and Herndon, Virginia) and in Dublin, Ireland.
    In order to direct traffic to all these data centers, Google could thoeretically record all queries centrally and then send them to the data centers. But this would obviously be inefficient. In fact, each data center has its own IP address (numerical address on the internet) and the way these IP addresses are accessed is managed by the Domain Name System.
    Basically, the DNS works like this: On the Internet, data transfers always take place in-between IP addresses. The information about which domain resolves to which IP address is provided by the name servers of the DNS. When a user enters a domain into his browser, a locally configured name server gets him the IP address for that domain by contacting the name server which is responsible for that domain. (The DNS is structured hierarchically. Illustrating the whole process would go beyond the scope of this paper.) The IP address is then cached by the name server, so that it is not necessary to contact the responsible name server each time a connection is built up to a domain.
    The records for a domain at the responsible name server constitute for how long the record may be cached by a caching name server. This is the Time To Live (TTL) of a domain. As soon as the TTL expires, the caching name server has to fetch the record for a domain again from the responsible name server. Quite often, the TTL is set to one or more days. In contrast, the Time To Live of the domain www.google.com is only five minutes. So, a name server may only cache Google's IP address for five minutes and has then to look up the IP address again.
    Each time, Google's name server is contacted, it sends back the IP address of only one data center. In this way, Google queries are always directed to different data centers by changing DNS records. On the one hand, the DNS records may be based on the load of the single data centers. In this way, Google would conduct a simple form of load balancing by its use of the DNS. On the other hand, the geographical location of a caching name server may influence how often it receives the single data centers' IP addresses. So, the distance for data transmissions can be reduced.
    How data centers, DNS and Google Dance are related, is easily answered. During the Google Dance, the data centers do not receive the new index at the same time. In fact, the new index is transferred to one data center after the other. When a user queries Google during the Google Dance, he may get the results from a data center which still has the old index at one point im time and from a data center which has the new index a few minutes later. From the users perspective, the index update took place within some minutes. But of course, this procedure may reverse, so that Google switches seemingly between the old and the new index.
    Finally, it shall be noted that Google did the DNS load balancing by themselves until September 2003. Since then, they use the services and, hence, the name servers of Akamai Technologies, Inc.
    IP Addresses and Domains of Google's Data Centers
    The progression of a Google Dance could basically be watched by querying the IP addresses of Google's data centers. But queries on the IP addresses are normally redirected to www.google.com. However, Google has domains which resolve to the single data centers' IP addresses. These domains as well as their IP addresses are shown in the following list.
    Domain IP-Adresse
    www-ex.google.com 216.239.33.100
    www-sj.google.com 216.239.35.100
    www-va.google.com 216.239.37.100
    www-dc.google.com 216.239.39.100
    www-ab.google.com 216.239.51.100
    www-in.google.com 216.239.53.100
    www-zu.google.com 216.239.55.100
    www-cw.google.com 216.239.57.100
    www-fi.google.com 216.239.41.100
    www-gv.google.com 216.239.59.100
    www-kr.google.com 66.102.11.100
    www-mc.google.com 66.102.7.100
    www-lm.google.com 66.102.9.100
    Those that keep an eye on Google's index updates often think that the Google Dance is over, when they see the new index at www.google.com or when they don't see the old index at www.google.com for some time. In fact, the update is not finished until all the domains listed above provide results from the new index.
    The index updates at the single data centers seem to happen at one point in time. As soon as one data center shows results from the new index, it won't switch back to the old index. This happens most likely because the index is redundant at each data center and at first, only one part of the servers (eventually half of them) is updated. During this period, only the other half of the servers is active and provides search results. As soon as the update of the first half of servers is finished, they become active and provide search results while the other half receives the new index. Thus, from the user's perspective, the update of one data centers happens at one point in time.
    Finally, it shall be noted that the access to the single data centers is generally controlled by the DNS only, but sometimes queries are redirected. However, this is easy to detect: When for a query at one of the domains listed above, the links to Google's cache do not comply with the IP address that belongs to the domain, then the query is redirected. If this happens, Google inhibits - for whatever reason - the access to one data center.
    The Google Dance Test Domains www2 and www3
    The beginning of a Google Dance can always be watched at the test domains www2.google.com and www3.google.com. Those domains normally have stable DNS records which make the domains resolve to only one (often the same) IP address. Before the Google Dance begins, at least one of the test domains is assigned the IP address of the data center that receives the new index first.
    Building up a completely new index once per month can cause quite some trouble. After all, Google has to spider some billion documents an then to process many TeraBytes of data. Therefore, testing the new index is inevitable. Of course, the folks at Google don't need the test domains themselves. Most certainly, they have many options to check a new index internally, but they do not have a lot of time to conduct the tests.
    So, the reason for having www2 and www3 is rather to show the new index to webmasters which are interested in their upcoming rankings. Many of these webmasters discuss the new index at the Google forums out on the web. These discussions can be observed by Google employees. At that time, the general public cannot see the new index yet, because the DNS records for www.google.com normally do not point to the IP address of the data center that is updated first when the update begins.
    As soon as Google's test community of forums members does not find any severe malfunctions caused by the new index, Google's DNS records are ready to make www.google.com resolve the the data center that is updated first. This is the time when the Google Dance begins. But if severe malfunctions become obvious during this test phase, there is still the possibility to cancel the update at the other data centers. The domain www.google.com would not resolve to the data center which has the flawed index and the general public could not take any notice about it. In this case, the index could be rebuilt or the web could be spidered again.
    So, the search results which are to be seen on www2.google.com and www3.google.com will always appear on www.google.com later on, as long as there is a regular index update. However, there may be minor fluctuations. On the one hand, the index at one data center never absolutely equals the index at another data center. We can easily check this by watching the number of results for the same query at the data center domains listed above, which often differ from each other. On the other hand, it is often assumed that the iterative PageRank calculation is not finished yet, when the Google Dance begins so that preliminary values exert influence on rankings at that point in time.
    The New PageRank Values during the Google Dance
    Most webmasters are interested in ranking changes for their website during the Google Dance. But, besides that, many also want to know about their new PageRank values. Normally, the Google Toolbar fetches the PageRank values from the data center that is specified by its IP address in the actual DNS record for www.google.com. Hence, when the Google Dance begins, the Toolbar usually displays the old PageRank values.
    Google submits PageRank values in simple text files to the Toolbar. In former times, this happened via XML. The switch to text files occured in August 2002. The PageRank files can be requested directly from the domain www.google.com. Basically, the URLs for those files look like follows (without line breaks):
    http://www.google.com/search?client=navclient-auto&
    ch=0123456789&features=Rank&q=info:http://www.domain.com/
    There is only one line of text in the PageRank files. The last cipher in this line is PageRank.
    The parameters incorporated in the above shown URL are inevitable for the display of the PageRank files in a browser. The value "navclient-auto" for the parameter "client" identifies the Toolbar. Via the parameter "q" the URL is submitted. The value "Rank" for the parameter "features" determines that the PageRank files are requested. If it is omitted, Google's servers still transmit XML files. The parameter "ch" transfers a checksum for the URL to Google, whereby this checksum can only change when the Toolbar version is updated by Google.
    The PageRank files that are requested by the Google Toolbar are cached by the Internet Explorer. So, their URLs and the checksums can simply been found out by having a look at the folder Temporary Internet Files. Knowing the checksums of your URLs, you can view the PageRank files in your browser. Since the PageRank files are kept in the browser cache and, thus, are clearly visible, and as long as requests are not automated, watching the PageRank files in a browser should not be a violation of Google's Terms of Service. However, you should be cautious. The Toolbar submits its own User-Agent to Google. It is:
    Mozilla/4.0 (compatible; GoogleToolbar 1.1.60-deleon; OS SE 4.10)
    1.1.60-deleon is a Toolbar version which may of course change. OS is the operating system that you have installed. So, Google is able to identify requests by browsers, if they do not go out via a proxy and if the User-Agent is not modified accordingly.
    Now, let's see how we can get the new PageRank values. Taking a look at IE's cache, you will notice that the PageRank files are not requested from the domain www.google.com but from IP addresses like 216.239.33.102. Additionally, the PageRank files' URLs often contain a parameter "failedip" that is set to values like "216.239.35.102;1111" (Its function is not absolutely clear). However, it is pretty easy to get the new PageRank values. Simply modify the IP addresses in the URL so that the request goes to one of the data centers that already has the new index. The necessary information is given above.

    Sunday, February 7, 2010

    Measuring Results

    Is your site generating more leads, higher quality leads, or more sales? What keywords are working? You can look at your server logs and an analytics program to track traffic trends and what keywords lead to conversion.

    Outside of traffic another good sign that you are on the right track is if you see more websites asking questions or talking about you. If you start picking up high quality unrequested links you might be near a Tipping Point to where your marketing starts to build on itself.

    Search engines follow people, but lag actual market conditions. It may take search engines a while to find all the links poiting at your site and analyze how well your site should rank. Depending on how competitive your marketplace is it may take anywhere from a few weeks to a couple years to establish a strong market position. Rankings can be a moving target as at any point in time


    * you are marketing your business
    * competitors are marketing their businesses and reinvesting profits into building out their SEO strategy
    * search engines may change their relevancy algorithms
     
    rantop.com
    ....Our Business Partners....

    Rainrays Web Directory


    Earn upto Rs. 9,000 pm checking Emails. Join now!