Introduction About Site Map

XML
RSS 2 Feed RSS 2 Feed
Navigation

Main Page | Blog Index

Archive for the ‘Internet’ Category

PageRank and Traffic

Google’s PageRank mechanism refelects on popularity estimates in most engines and directories, which presently adopt similar ranking methods that are based on the notion of citations (links).

The average webmaster would fall victim to an illusion. It always appears as if the majority of other sites have higher ranks. A question then springs to mind: how come most pages being visited have PageRank 4 or higher, for example, if they are only a small minority? The matter of fact is that much of the Web, including one’s own site have much lower ranks. That remote part of the Web is often virtually invisible to the errant surfer.

Let us recognise the fact that Google, Yahoo and CNN, for example, receive far more traffic than other sites. Their traffic is many orders of magnitude higher than that of the average site. To visualise this, I drew two exponential curves. Some figures suggest that the curves I drew for illustrational purposes should be far steeper. Some would suggest that at any level of PageRank the number of sites may be 3 times smaller and the traffic 3 times greater than that of the previous PageRank level.

PageRank versus traffic
The number of sites with PageRank 10 is tiny when compared to the number of sites with PageRank 0. Conversely, traffic is largely centralised in sites with a high PR.

To give an idea of how vast the World Wide Web actually is, the number of sites is slowly approaching 100 million (all registered global domains, as well individual parts of the world). At present, PageRank 0 may fit 50 million sites (some are not identified, not listed, or parked), PageRank 2: 5 million sites, PageRank 4: 500,000 sites and so forth. Of course this rough guess does not lead to the true numbers, which end up at just dozens of sites with PageRank 10.

In conclusion, always remember that the vast majority of sites is on the left-hand-side of the figure above, whereas some of the far more popular pages are on the (comparably) tiny number of pages on the right, e.g. the Google’s search page, the BBC front page or the W3 consortium pages.

Weather RSS Feeds

Man with binoculars

Methods exist for fetching weather forecasts as XML (RSS feeds). Weather Underground is the international giant, but its feeds are not sufficiently extensive. It reflects on current conditions in various places in the world, but provides no forecast. To subscribe to such feeds, go to Weather Underground, search for your town/city and identify the RSS2 button (presently at the top-right of the resulting page).

For more extensive and useful details, make use of the brand new XML’d information from the American National Weather Service:

The folks in the United Kingdom can enjoy the BBC Weather RSS feeds, which are in fact scraped from BBC Weather They are not quite legitimate, but they provide a concise yet detailed 5-day forecast.

Wikipedia Statistics

Wikipedia statistics

Maybe purely by mistake, or maybe intentionally so, Wikipedia Stats pages are available for public viewing.

Interesting figures to notice:

  • Internet Explorer’s share among Wikipedia users/visitors is 70% (it gets around 52% on schestowitz.com)
  • Referrals by search engines:
    • Google: 3082040
    • Yahoo: 1156329
    • MSN: 263755

This implies that certain search engines favour Wikipedia more than others (judge for yourself). It also suggests that Firefox users tend to like Wikipedia more than Internet Explorer users.

Yahoo to Compete with Skype

Yahoo telephone

The Computer Business Review reports that Yahoo will soon join the VoIP market and target the hugely popular Skype. VoIP, which has become synonymous with Skype, is a method/service allowing telephone calls to be made over the Internet, making such calls virtually free. With a standard telephone connected to a standard line, call fares are very low too. In the latter case, international call fares equate to these of local calls.

Yahoo is planning to become a full-service VoIP provider, but declined to give a timeline on new product announcements beyond “coming months,” said spokesperson Terrell Karlsten.

The Rise of the Wiki

Only yesterday, one of the regulars in a search engines discussion group (nntp://alt.internet.search-engines) raised the following question/complaint:

Everytime I search for something on google Wikipedia appears in the top 100 results! Is it taking over the Internet?

Last night, IT guru Joel Spolsky asked for advice on Wiki for his site’s translations. He could no longer resist the productivity of a Wiki, which allows visitors to manage content while vandalistic changes can be quickly rollbacked (reversed). The matter of fact is that Wikis have become extremely ubiquitous and they are conquering the Net, much like blogs. They simplify the composition of Web pages, particularly when composition is handled and managed by a group.

Schestowitz.com has two Wikis — a private one and a public one. At some stage I found that Wikis can conveniently be used for correspondence or for collaboration in academia, e.g. when writing or revising a paper jointly.

Wiki
The Public Wiki section on this domain

There are all sort of reasons for using the Wiki’s simplified mark-up and not WYSIWYG interfaces (as clean as they may be) or plain HTML.

  • Firstly, there is the issue of security such as off-site linking. You need to provide a subset of features. Having said that, this can be prevented by imposing restrictions on whichever composition language/method you choose.
  • Secondly, one has to worry about the issue of consistency. I have a friend who manages the content of a site and I cannot recall how many times he did something invalid that broke the integrity of the entire site, which I then had to fix urgently. “More restrictive” means “less things to go wrong”. After a few weeks I completely gave up on making his pages valid XHTML/CSS. He had too much freedom and too little experience.

Finally, on the issue of software recommendation, I cannot comment on the popular MediaWiki, which is used to power Wikipedia. I have never used MediaWiki myself, but PHPWiki has been fairly stable ever since I set it up 6 months ago and had it updated many times a day. If you are looking for ease-of-use and trivial embedment of media, I suggest you go with more mainstream packages such as the admirable MediaWiki, which I am familiar with as a user, but not as an administrator.

Google OS – What if?

Linux and Google

A venturous article from Softpedia makes speculations with regards to a Google homebred operating system that is based on GNU/Linux. Google’s Open Source affinity is out of the bag now, so looking 5 years ahead, will there be a tighter integration between Google’s on-line services and a Google-controlled Linux distrubution?

Google is non-conformist enough, and it has sufficient money and knowledge to be able to venture in this domain. The recipe? You take a Linux distribution (Google’s appetite for Open Source and Linux is no longer a secret), you mix it with Google’s knowledge on Internet searching, e-mail, security, programming and document indexing, you give it a name that includes Google and OS, and there you have it.

Content Spam Prevention

Browser searchContent spam grows worryingly fast. This new type of spam shows its face in a form which is different from spam that we very well know as uninvited E-mail. Such spam is placed on the Web and later infiltrates search engine results pages (SERP’s). If you recently came across a page full of links and ads that did not provide useful information, you would probably understand what content spam is.

Spam content pages are often generated automatically (robot-made). Such robots would run a search, pick up HTML code from the results page, and finally add some advertisement. This generates a large volume of pages, which warrant presence in many SERP‘s. One very large site that I recently came across got tempted and uses this technique to woo visitors and offer them subscription. If you come across such sites, be sure to report them as its our only chance of banning illicit sites and preventing content spam (and referrer spam, i.e. organic and fake links/referrals to sites) from expanding. Google provide an on-line form for this very particular purpose.

Retrieval statistics: 21 queries taking a total of 0.086 seconds • Please report low bandwidth using the feedback form
Original styles created by Ian Main (all acknowledgements) • PHP scripts and styles later modified by Roy Schestowitz • Help yourself to a GPL'd copy
|— Proudly powered by W o r d P r e s s — based on a heavily-hacked version 1.2.1 (Mingus) installation —|