Search Engines and Spying
ave you ever pondered or worried about what search engines know about you? If not, you should. It is no secret that even the most prominent and powerful companies see value in tracking their users. Companies do not have any desire to infringe the user’s right for privacy, but information about their users allows them to improve their services, e.g. by behaviour learning which leads to optimisation and targetted content.
A few clarifications are worth making: If you use Windows and, more particularly, if you use Internet Explorer, your privacy is jeopardised the most. I am not pointing my finger at Microsoft, but spyware is targetted at the most prolific platform and application. It is the most cost-effective spyware development process. Have you heard about Alexa rating? Ranks are most likely based on spyware at some level or another. The watchers know what sites you visit. The watchers might also know what pages you bookmark, how long you ‘sit’ on a SERP, which pages you follow and the list goes on and on.
Google should be able to infer, based on a small statistical sample that is exposed to spying, which page got bookmarked the most after it was reached from one given SERP (e.g. leaving SERP, then hitting “back” and requesting the SERP again). This way, spamfull sites can be distinguished from sites the user actually had interest in and ended up bookmarking. It is a machine learning approach. Results adapt to users’ behaviour. With a cunning Google cookie that precisely identifies users, you can only imagine the tonage of data that is accumulated about individuals. Shall it be needed, let it be taken ‘off the shelf’…
On to Microsoft, they already exploit cookies by swapping them across sites, which of course is questionable if not an illegal practice. They have recently bought a spyware company (Claria who make Bonzi and other related ‘goodies’). Consequently, they downgraded Claria’s threat level, which was a cause for controversy, much like a controversy with Firefox. Users reported that Firfox was labelled harmful by Microsoft’s spyware detection application; Firefox removal was recommended too. All in all, it appears as though MSN have similar plans in mind and aggressive behaviour remains one of their innate features.
Disclaimer: I am not entirely sure about each one of the facts above, but spying is nothing unprecedented. What I argued is based on text (and I do not mean science-fiction or conspiracy theory resources), though many speculations were involved and such communication must be taken with a grain of salt.






Filed under: 

t was roughly a decade ago when a mysterious journalism trend emerged. All of a sudden, in a matter of a few years, major newspapers began to migrate their content to the Internet and shield their on-line popularity. A frantic stampede — that was — to what would become the future of journalism. Editors came to realise that it was only a matter of time before readers would take advantage of technology. This new form of publication had very many conceivable advantages:
ver a thousand people have reached this site when looking for an FTP client for Firefox. As my 

itemaps are probably the most crucial pages in all Web sites. They might not be most helpful to human visitors, but they greatly assist crawlers, thereby attracting, promoting and inviting more traffic from search engines. Site maps can be perceived as ‘crawling maps’, regardless of how shallow or deep they are. These maps not just the spine, but possibly the entire skeleton of complex Web sites where crawling will take a huge number of distinct routes through pages.
Google have recently introduced RSS sitemaps. This means that new site content will be appended to the map whose form is a long aggregated feed, i.e. links with minimal content and without unimportant media and layout detail. This move by Google encouraged many Webmasters to go on the RSS wagon and XML their Web sites. This benefits Google in a variety of ways. First of all, there is a clear pairing between content and dates. If the site delivers timely news, this will become a significant factor. Secondly, as any site is described fully by its map, there is no need for repeated crawling. There are significant savings in terms of bandwidth, which appear to allow search engines to crawl a sparser portion of the Web, as well as reduce the burden on Web servers, whose load involves a great deal of pages being served to crawlers or robots.
Let us think of nature for a moment. Think of rabbits whose reproduction rate is potentially very high. Once you have too many rabbits in the wild, not enough resources exist in nature in order to feeds them all. What’s more, predators that feed on rabbit will prosper and rapidly bring the number of rabbit down, thereby quickly reproducing themselves — the predators. It is a cycle in nature that has been explained by biologists many times in the past.
Also on the same topic, it is rather frustrating that it took several years until Microsoft decided to fix the bugs inherent in Internet Explorer 6 and add valuable yet fundamental features like support for feeds. This move was a result of Firefox, of course, but more sadly, 