Introduction About Site Map

XML
RSS 2 Feed RSS 2 Feed
Navigation

Main Page | Blog Index

Archive for the ‘Internet’ Category

Google and Pet Peeves

Dog fine sign
Letting that pet go loose

RECENT observation of Google’s moves has led to the accumulation of several remarks. Hereby, I would like to list a few of them, getting them off my chest for what it’s worth.

  • Google Base. For those who do not know, Google have just launched a new services that awakens the desire to stand up and shout “All your bases are belong (sic) to Google”, which is a phrase that goes far back in time. For those unfamiliar with the phrase, it stems from poor translation of a computer game and its rapid spread is primarily attributed to UseNet. Google Base, as the service was entitled, appears to be eyeing that gigantic, non-profitable Craig’s List, intending to use that new platform for embedment of yet more targetted ads. This argument is nothing beyond speculation nonetheless.
  • Literature domination. The highly controversial book scanning initiative (Google Print), supposedly to be followed by Microsoft rather soon. Where there is potential evil, there must be a Microsoft stampede.
  • Updates with personality. Google are naming their updates, e.g. Bourbon and now Jagger (due to complete at the beginning of next month). That naming convention is reminiscent of that which is ascribed to hurricanes. Do they have a list of names queued up for future assignment?
  • Indexing Obsession. Google have declared their desire to crawl and organise the entire human knowledge. They also said that indexing may take 300 years (a wild speculation by their CEO), but should they not understand content rather than simply index it all? I have recently proposed an alternative, which relies on semantics and factual data.

Now that my rants have been voiced, I feel surprisingly relieved. I like Google and often rave about their search performance. All in all, I hope my criticisms are all constructive rather than unnecessarily excruciating.

CORRECTION (29/10/2005): the current update was dubbed “Jagger” by WebmasterWorld.

Windows and Web-based Software

Windows AJAX
Windows XP in AJAX

An article from CNN, which as expected does not descend to technicalities, explains why Web-based applications put Windows and Microsoft Office under realistic threat.

NEW YORK (AP) — A quiet revolution is transforming life on the Internet: New, agile software now lets people quickly check flight options, see stock prices fluctuate and better manage their online photos and e-mail.

Such tools make computing less of a chore because they sit on distant Web servers and run over standard browsers. Users thus don’t have to worry about installing software or moving data when they switch computers.

And that could bode ill for Microsoft Corp. and its flagship Office suite, which packs together word processing, spreadsheets and other applications.

The threat comes in large part from Ajax, a set of Web development tools that speeds up Web applications by summoning snippets of data as needed instead of pulling entire Web pages over and over.

In Defence of Google Print

Book scanning

ERIC Scmidt, who is the CEO of Google, speaks out in defence of Google Print. Goole Print is the controversial initiative to scan literature, infringing copyrights in the process. The intended service, titled Google Print, is to provide surfers with instant and comprehensive coverage of books from the shelf.

Imagine sitting at your computer and, in less than a second, searching the full text of every book ever written. Imagine an historian being able to instantly find every book that mentions the Battle of Algiers. Imagine a high school student in Bangladesh discovering an out-of-print author held only in a library in Ann Arbor. Imagine one giant electronic card catalog that makes all the world’s books discoverable with just a few keystrokes by anyone, anywhere, anytime.

Collaborative Effort to Crawl the Web

Iuron

ONLY a few days ago, somebody had me aware of the Majestic-12 distributed search engine. The idea behind the engine is persistent use of other people’s computer power and bandwidth. The goal is mind is to crawl and potetially index the Web reasonably well.

This sudden ‘enlightement’, to me at least, provided somewhat of an insight. It affected the matter of practicability in my Open Source Iuron, which is in its early stages and more of a porposal at this stage. As explained before, Iuron does not index pages; it aspires to gain actual knowledge from the Internet instead. This can potentially make PageRank (or equivalents) obsolete, I believe, thereby reducing spam and search engine cheats.

Within a few days, I will be meeting the person who is arguably the father of the Semantic Web. My project will be difficult to lift off the ground without some support. Nonetheless, this now appears to be a hindrance with a simple solution. It is, after all, the kind of project where the vast requirement for bandwidth and computer power can be obtained in more or less the same way as Majestic-12. Since it is Open Source, willingness on the public’s behalf should not be a considerable peril.

On an unrelated topic which is paranoia, I recently noticed a referral reduction from Google. It became conspicuosuly significant in recent days so I thought it was an attempt to silence me. It finally turns out to have been merely a side-effect of a large-scale update at Google’s end. Many Web sites were in fact affected by this and distress became apparent in a few newsgroups. It was even pointed out that msn.com was assigned PageRank 2!

Firefox Fork

Firefox in the dock

I have just become aware of Flock, which is an interesting fork of Firefox 1.5. The much-anticipated version 1.5 has not been formally released yet, which makes this a somewhat controversial scenario. Flock is now being promoted by WordPress.com, where a download of Flock warrant a free WordPress blog (at least for the time being).

With all due respect, I am always slightly apprehensive when it comes to adopting, thus relying on forks. I am also aware of the problem associated with forking one’s own application. Once you lag behind, the long-invested dedication can wind up being disposed of. Flock appears to me like the conspicuous rationale behind Mozilla becoming a foundation and going by the identity of Mozilla.com.

As regards WordPress.com and the Flock relationship, I might give it a try, but I would certainly seek plenty of convincing arguments before I do so. My past experiences with Firefox 1.5 betas (AKA Deer Park) have been fairly disappointing and led to regrets. I have made two such attempts to migrate to a version that was not finalised.

Having said that, a certain other fact is worrying me slightly more. According to ZDNet, the lead developer of Flock said:

“Please note that this is a developer preview and that there are still plenty of bugs, many of which we are aware of.”

To me that sounds as if the application is not quite ready for “prime time” so WordPress.com are possibly getting users on a dangerous wagon.

It was also said, however:

“In architecting our software, build systems and engineering processes, we have given considerable thought to how our code will be able to evolve alongside the Mozilla code, without forking it”

This sounds rather re-assuring and I sure hope these folks will walk as they preach.

Related item: Best Technology Products of 2005

The Demise of PageRank

PageRank versus traffic
The number of sites with PageRank 10 is tiny when compared to the number of sites with PageRank 0. Conversely, traffic is largely centralised in sites with a high PR (more details)

A little tour around Google has led me to a questionably out-of-date article from the Register. This Article is 2 years old, but it seems more true than ever these days because scraping and link-related spam/attacks are constantly on the rise.

Google has made no secret of its goal to “understand” the web, an acknowledgement that its current brute-force text index produces search results with little or no context. The popularity of Teoma demonstrates that even a small index can produce superior results for certain kind of searches. Teoma leans on existing classification systems.

While Google relied on PageRankâ„¢ to provide context, all was well. But PageRank is now widely acknowledged to be broken, so new, smarter tricks are required.

Regarded as heresy when we raised the issue last spring, now some of Google’s warmest admirers, MetaFilter’s Matt Haughey and web designer Jason Kottke have acknowledged the problem.

As Gary Stock noted here last May, Google “didn’t foresee a tightly-bound body of wirers. They presumed that technicians at USC would link to the best papers from MIT, to the best local sites from a land trust or a river study – rather than a clique, a small group of people writing about each other constantly. They obviously bump the rankings system in a way for which it wasn’t prepared.”

The intersting fact is that Google themselves acknowledge the problem and I am sure difficulties have intensified, if anything, in the past 2 years. The specific reference to bloggers proves that very point as a new blog is set up every second these days.

Regarding the point about pages being indexed rather than learned from or understood, that is one of the catalysts that led me to starting Iuron. That site has attracted tremendous levels of interest since the idea had been conceived on the night on October 9th. I set up the Web site and made an official announcement the following day. Yesterday I finished a 1-page formal proposal and I contacted the person who is perceived by some as the father of the Semantic Web. He was once my lecturer.

WikiMirror – Vile Ripoff

Sky scrapers
Content scrapers: where is the original content? Which one is the ripoff?

MUCH that we see on the Internet these days are mirrors, although we are rarely aware of it. Several crooks make good money out of it. Some search engines are crippled by the fact that they have no knowledge as to which sites are known ‘mirror culprits’ and which ones can be trusted. Consequently, search engines like Yahoo tend to return many references to content scrapers, which is a deterrent.

As I go about looking at some SERP‘s I suddenly come across a commercial site — a Wikipedia mirror — that had been registered since January 2005. Its description in Google (judge for yourselves):

Wikimirror.com – Free encyclopedia search a b c d e f g h i j k l m n o p q r s t u v w x y z _ · Google Web Encyclopedia.
 

WTF? Is the Creative Commons Licence an invitation to massive, huge-scale ripoffs? A community of WikiPedians around the world is voluntarily spending time in vain? To get somebody else rich(er)? I have had a look around the site in question. It appears like a complete mirror, which I must stress cannot be edited, unlike its much superior source. I recently discussed the issue of people mirroring the CIA Factbook, which is public content, in a relevant newsgroup. When will this end? And why has Google not banned Wikimirror.com yet? Does the domain name not say something? Helloooooo…?

I spoke to Chris Pirillo about Blogspot spam yesterday. After our discussion he posted an item that makes a nice little read. That item is titled Google: Kill Blogspot Already!!!, which is a venturous and strong title to be used by somebody as prominent as Chris.

Also while on the subject, have a look at Networkmirror.com [rel='nofollow']. In its defence, this one-among-many Slashdot mirrors does not archive content and it serves a defensible purpose — that of mirroring sites before they go down due to the Slashdot Effect. Then mirrors sites at least get exposure while they cannot cope with the demand.

UPDATE: 1-script argues that I may have been a little hasty in posting this item. The site states:

Content Credit

Wikimirror financially supports the Wikimedia Foundation. Displaying this page does not burden Wikipedia hardware resources. This article is from Wikipedia. All text is available under the terms of the GNU Free Documentation License. Contact: info [AT] wikimirror [DOT] com

As a side note, I still think that Wikipedia contributers (me included) ought to be aware of these facts (or the existence of this mirror), which may imply that their contributions become commercial and thus money-making.

Retrieval statistics: 21 queries taking a total of 0.129 seconds • Please report low bandwidth using the feedback form
Original styles created by Ian Main (all acknowledgements) • PHP scripts and styles later modified by Roy Schestowitz • Help yourself to a GPL'd copy
|— Proudly powered by W o r d P r e s s — based on a heavily-hacked version 1.2.1 (Mingus) installation —|