Introduction About Site Map

XML
RSS 2 Feed RSS 2 Feed
Navigation

Main Page | Blog Index

Archive for the ‘Internet’ Category

Search Engines and Spying

CCTV

Have you ever pondered or worried about what search engines know about you? If not, you should. It is no secret that even the most prominent and powerful companies see value in tracking their users. Companies do not have any desire to infringe the user’s right for privacy, but information about their users allows them to improve their services, e.g. by behaviour learning which leads to optimisation and targetted content.

A few clarifications are worth making: If you use Windows and, more particularly, if you use Internet Explorer, your privacy is jeopardised the most. I am not pointing my finger at Microsoft, but spyware is targetted at the most prolific platform and application. It is the most cost-effective spyware development process. Have you heard about Alexa rating? Ranks are most likely based on spyware at some level or another. The watchers know what sites you visit. The watchers might also know what pages you bookmark, how long you ‘sit’ on a SERP, which pages you follow and the list goes on and on.

Google should be able to infer, based on a small statistical sample that is exposed to spying, which page got bookmarked the most after it was reached from one given SERP (e.g. leaving SERP, then hitting “back” and requesting the SERP again). This way, spamfull sites can be distinguished from sites the user actually had interest in and ended up bookmarking. It is a machine learning approach. Results adapt to users’ behaviour. With a cunning Google cookie that precisely identifies users, you can only imagine the tonage of data that is accumulated about individuals. Shall it be needed, let it be taken ‘off the shelf’…

Google Cookie
The Google Cookie

On to Microsoft, they already exploit cookies by swapping them across sites, which of course is questionable if not an illegal practice. They have recently bought a spyware company (Claria who make Bonzi and other related ‘goodies’). Consequently, they downgraded Claria’s threat level, which was a cause for controversy, much like a controversy with Firefox. Users reported that Firfox was labelled harmful by Microsoft’s spyware detection application; Firefox removal was recommended too. All in all, it appears as though MSN have similar plans in mind and aggressive behaviour remains one of their innate features.

Disclaimer: I am not entirely sure about each one of the facts above, but spying is nothing unprecedented. What I argued is based on text (and I do not mean science-fiction or conspiracy theory resources), though many speculations were involved and such communication must be taken with a grain of salt.

Evolve or Die

Longhorn

This picture was funnier when Windows Vista was still called Longhorn

It was roughly a decade ago when a mysterious journalism trend emerged. All of a sudden, in a matter of a few years, major newspapers began to migrate their content to the Internet and shield their on-line popularity. A frantic stampede — that was — to what would become the future of journalism. Editors came to realise that it was only a matter of time before readers would take advantage of technology. This new form of publication had very many conceivable advantages:

  • Being able to isolate uninteresting bits from the unmissable (content tailoring)
  • Ability to save article (electronic storage)
  • Reduction in cost (physical properties)
  • The ability to share articles with friends and colleagues (reproducibility)
  • Being able to performs searches (indexing technology)

Whatever gains one can imagine, most likely computers will have them. So, paper was bound to become obsolete sooner or later.

Several years later, particularly in the beginning of this millennium, blogs (Web logs) began to emerge. Suddenly, people had easy access to powerful publication platforms. The growth of blogs in terms of number and amount of content was explosive and their extent soon became excessive. How could a mainstream newspaper keep up with blogs and attract the masses? Journalists became threatened. Later came he feeds frenzy, which unlike blogs, has not reached its anti-climax, yet. Soon enough, operating systems (as we know them) might be put aside in favour of on-line operating systems. So who will be the next victim if not Microsoft? And who will it be that inherits the Earth? Google and Yahoo are among the contenders. For Microsoft, who have made many enemies, the last resort has become software patents, which they are piling up like connonballs for application at the Patents Office.

Slow movement towards on-line data management has been well-comprehended by all fronts. Microsoft attempt to conquer the Net as they are investing a heavy load of resources to reach that goal. From this site alone, MSNBot fetched 13968 pages, trailing by just a decent margin behind Google with 20627 pages in the month of July. Without a doubt, Microsoft have gained plenty of bandwidth and computer power.

It was not long ago that Bill Gates gave his engineers just 100 days to steal Google’s idea of AJAX-enabled maps (satellite and hybrid maps too). In response to Mozilla Firefox, which kept Microsoft on their toes, a decision was make to update rusty Internet Explorer 6 and leap to Internet Explorer 7. Not surprisingly, many ideas were stolen from Firefox, which was considered the main danger at the time.

Later on, Microsoft decided to embed RSS support in the kernel of Windows Vista (formerly Longhorn). Again, this step was taken in order to compensate for the considerable lag behind some recent technologies. Gates et al. were then scheming to buy off popular bloggers — a shrewd idea due to general Microsoft disdain in popular and influential blogspheres. The main complaints of bloggers, along with some major voices in the media, was the lack of inter-operability, being proof that Microsoft are patronising the rest of the IT world with utter platform discrimination. They simply fail to internalise that it is nice to be important, but even more important to be nice (open).

Firefox & FTP

FireFTP

FireFTP – Click image for the full-sized version
Firefox is Mac OSX-themed (Download/install)

Over a thousand people have reached this site when looking for an FTP client for Firefox. As my previous post on the topic was succinct and vague, I decided to elaborate, capture some screenshots, report on my experiences, conduct some comparisons and make recommendations.

I have run FireFTP beta (version 0.88) on two separate computers with different operating systems since late April/early May when it was released to the public. The time and content of my previous post indicates that I got notified via RSS feeds so FireFTP must still be very new and ‘fresh’. Bugs are rare, but they appear recurrent for instance, when recursive, deep downloads/uploads are pushed to the extreme. Nonetheless, in this particular problem domain which is file transfer, nothing is mission-critical by nature. The data never gets corrupted, only the flow of control in the application lacks reliability. Overall, I would happily assign a rating of 7/10 to FireFTP.

As my essays tend to (knowingly) drift away to separate, yet related topics, I shall survey a few alternatives which also make good FTP clients. My experience with FTP clients goes back to age of 15 or 16. Among the applications I have continuously used for the task are:

  • CuteFTP
  • SmartFTP
  • WSFTP
  • CrystalFTP
  • Hummingbird
  • IE
  • …and surely a few more Windows applications, which go back too far for me to recall

I ordered the list above by frequency or duration of use; it is not chronological. Although most of the above were shareware (especially the former), there were many UI ‘nags’ involved. Around the year 2000 I began to drift to Linux, which by nature, offered nag-free applications that were solid workhorses. My favourite FTP client, which is most valuable for the majority of tasks, is GFTP. I run it remotely from RedHat Fedora Code II clusters. GFTP achieves extremely high transfer rates, perhaps by pipelining (the site corrected me by saying that GFTP is multi-threaded. I believe that in the Windows-oriented list above, only IE supports multi-threading). That powerful feature allows me to download and upload 1,000 files in a matter of seconds. This is useful, for example, whenever I want to apply a massive search-and-replace to a comprehensive site section.

GFTP

GFTP screenshot – Click image to see it full-sized

Linux users are encouraged to use GFTP for any traffic where latency becomes a hindering factor, in particular where a large number of files is involved. Unfortunately, one lost advantage is that GFTP cannot be integrated with the browser. FireFTP gets close to it (it gets embedded in a tab), but KDE‘s Konqueror (father of Safari) is another alternative, which is an HTTP, HTTPS, FTP and file management ‘platform’. It competes in terms of speed with GFTP. It is much superior in terms of usability too. Finally, worth adding is the observation that aforementioned applications are separate from the browser at one level or another. That is something that Microsoft have addressed, may it be good or bad. When the filesystem tool (explorer.exe) is hard-coded into the browser (or vice versa), there is not much room to ‘dance’ from one browser to another (and retain good use of resources).

Konqueror as an FTP client

Screenshot of Konqueror as an FTP client
click image to see it full-sized

Dangers of RSS Sitemaps

Sitemaps are probably the most crucial pages in all Web sites. They might not be most helpful to human visitors, but they greatly assist crawlers, thereby attracting, promoting and inviting more traffic from search engines. Site maps can be perceived as ‘crawling maps’, regardless of how shallow or deep they are. These maps not just the spine, but possibly the entire skeleton of complex Web sites where crawling will take a huge number of distinct routes through pages.

Map of EuropeGoogle have recently introduced RSS sitemaps. This means that new site content will be appended to the map whose form is a long aggregated feed, i.e. links with minimal content and without unimportant media and layout detail. This move by Google encouraged many Webmasters to go on the RSS wagon and XML their Web sites. This benefits Google in a variety of ways. First of all, there is a clear pairing between content and dates. If the site delivers timely news, this will become a significant factor. Secondly, as any site is described fully by its map, there is no need for repeated crawling. There are significant savings in terms of bandwidth, which appear to allow search engines to crawl a sparser portion of the Web, as well as reduce the burden on Web servers, whose load involves a great deal of pages being served to crawlers or robots.

There are hidden dangers in moving towards RSS sitemaps. Typical HTML/XHTML/other site maps get neglected as they are laborious to maintain and as time goes by they may appear redundant, much like an older generation of pages that have gone completely out of date. One of the worrying implications is that opponents — in this particular case MSN and Yahoo predominantly — get ‘robbed’ of the true, old-styled site maps, which are conceded altogether or simply neglected in favour of RSS sitemaps. Therefore, they must follow suit and take advantage of the public RSS site maps, just as Google did. They might need to woo Webmasters and get those sitemap submitted to them as well. Is it going to be an easy task? Probably not, especially while Google’s impact is by far superior.

The phenomenon above is the introduction and absorbance of new technologies by force. Google are forcing, not malevolently though, trends of crawling and push for methods to change. Microsoft exhibited some similar behaviour by introducing new elements to HTML without consent from the community. They incorporated these element into Internet Explorer; they considered the Internet to be a source to serve Internet Explorer rather than the reverse, whereby the Web is open and accessible to all. By doing so, Microsoft encouraged Web developers to construct Web pages that work exclusively under Internet Explorer and slowly killed their main opponent: Netscape Navigator. But that is all history as Internet Explorer (version 6) lags behind Opera, Firefox and arguably behind Safari too.

In practical terms, do RSS sitemaps lead to any gains? We have discussed this issue to death at the primary SEO-related newsgroup. Borek expressed his skepticism about Google sitemaps:

So far Google just fetches my sitemaps 4 times a day. One site is PR3 5 months old, second is PR2 several years old, redesigned in June. No signs of crawl on either (and there are not spidered pages on both sites).

The bottom line, from my point-of-view, is that for news-delivering sites, RSS sitemaps provide a good opportunity to conquer valuable SERP‘s very quickly. For most standard sites with a decent amount of bandwidth to spare, RSS sitemaps appear like an overkill, even when HTML to XML implementation are virtually available ‘off-the-shelf’. RSS sitemaps may also be valuable for blogs where the nature of publication is linear rather than hierarchical or lateral.

Other related threads on the topic:

Related News: Google to Patent Ads in Feeds

Internet and Nature

RabbitLet us think of nature for a moment. Think of rabbits whose reproduction rate is potentially very high. Once you have too many rabbits in the wild, not enough resources exist in nature in order to feeds them all. What’s more, predators that feed on rabbit will prosper and rapidly bring the number of rabbit down, thereby quickly reproducing themselves — the predators. It is a cycle in nature that has been explained by biologists many times in the past.

The same principles apply to the Web. When demand for Web sites is high (e.g. the birth in the early nineties), smashing success hits the few trailblazers and other envious business follow suit by joining. Yet, at some stage, the bubble must burst and there will be no more place for new sites to be accommodated. Is Internet reaching an anti-climax? I doubt so. According to one source, we are now observing a second dot-com boom develop. According to insightful predictions at Wired Magazine, we are yet to see the best of the Web:

2015

The Web continues to evolve from a world ruled by mass media and mass audiences to one ruled by messy media and messy participation. How far can this frenzy of creativity go? Encouraged by Web-enabled sales, 175,000 books were published and more than 30,000 music albums were released in the US last year. At the same time, 14 million blogs launched worldwide. All these numbers are escalating. A simple extrapolation suggests that in the near future, everyone alive will (on average) write a song, author a book, make a video, craft a weblog, and code a program. This idea is less outrageous than the notion 150 years ago that someday everyone would write a letter or take a photograph.

By 2015, desktop operating systems will be largely irrelevant. The Web will be the only OS worth coding for. It won’t matter what device you use, as long as it runs on the Web OS. You will reach the same distributed computer whether you log on via phone, PDA, laptop, or HDTV.

I have made merely identical predictions in the past:

No Competition = No Innovation

According to Bloomberg, innovation definitely required motivation:

Microsoft has had the ability to develop a satellite map service for MSN since 1998. Chairman Bill Gates decided in April, the same month Google released its version, to rush the project. He gave his engineers a deadline of 100 days. Stephen Lawler, who runs the Virtual Earth group, and his team figured they would need about a year to get the satellite mapping technology ready.

Horse raceAlso on the same topic, it is rather frustrating that it took several years until Microsoft decided to fix the bugs inherent in Internet Explorer 6 and add valuable yet fundamental features like support for feeds. This move was a result of Firefox, of course, but more sadly, Internet Explorer 7 imitates Firefox, much like Virtual Earth intends to match Google Maps and show that development in the Microsoft compus is not dead, yet. Note the subtle use of the word development, not innovation.

Internet Command Line

Computer shell
Getting data more quickly using a CLI

Some time ago I came across YubNub – YubNub.org which is a command-line interface for the Web. It allows you to operate in and on the Web much more efficiently than by using the traditional GUI‘s. You can, for example, query Google for ‘saturn’ by typing in

g saturn

YubNub provides a very quick way of navigation the Web as chaining of commands is possible too and nearly 100 commands exist. The site was a great find, but I do not believe it can replace a keyboard-navigable portal (see example). Maybe I can embed the YubNub command-line in portals at some stage, which would make it ever more flexible. While on the issue of portals, Google have extended features offered by their Portal (personalised homepage) service. Among the new features:

  • Adding your own bookmarks
  • Selecting from more news feeds
  • Adding your own RSS news feeds

Retrieval statistics: 21 queries taking a total of 0.095 seconds • Please report low bandwidth using the feedback form
Original styles created by Ian Main (all acknowledgements) • PHP scripts and styles later modified by Roy Schestowitz • Help yourself to a GPL'd copy
|— Proudly powered by W o r d P r e s s — based on a heavily-hacked version 1.2.1 (Mingus) installation —|