Introduction About Site Map

XML
RSS 2 Feed RSS 2 Feed
Navigation

Main Page | Blog Index

Archive for the ‘Internet’ Category

WikiMirror – Vile Ripoff

Sky scrapers
Content scrapers: where is the original content? Which one is the ripoff?

MUCH that we see on the Internet these days are mirrors, although we are rarely aware of it. Several crooks make good money out of it. Some search engines are crippled by the fact that they have no knowledge as to which sites are known ‘mirror culprits’ and which ones can be trusted. Consequently, search engines like Yahoo tend to return many references to content scrapers, which is a deterrent.

As I go about looking at some SERP‘s I suddenly come across a commercial site — a Wikipedia mirror — that had been registered since January 2005. Its description in Google (judge for yourselves):

Wikimirror.com – Free encyclopedia search a b c d e f g h i j k l m n o p q r s t u v w x y z _ · Google Web Encyclopedia.
 

WTF? Is the Creative Commons Licence an invitation to massive, huge-scale ripoffs? A community of WikiPedians around the world is voluntarily spending time in vain? To get somebody else rich(er)? I have had a look around the site in question. It appears like a complete mirror, which I must stress cannot be edited, unlike its much superior source. I recently discussed the issue of people mirroring the CIA Factbook, which is public content, in a relevant newsgroup. When will this end? And why has Google not banned Wikimirror.com yet? Does the domain name not say something? Helloooooo…?

I spoke to Chris Pirillo about Blogspot spam yesterday. After our discussion he posted an item that makes a nice little read. That item is titled Google: Kill Blogspot Already!!!, which is a venturous and strong title to be used by somebody as prominent as Chris.

Also while on the subject, have a look at Networkmirror.com [rel='nofollow']. In its defence, this one-among-many Slashdot mirrors does not archive content and it serves a defensible purpose — that of mirroring sites before they go down due to the Slashdot Effect. Then mirrors sites at least get exposure while they cannot cope with the demand.

UPDATE: 1-script argues that I may have been a little hasty in posting this item. The site states:

Content Credit

Wikimirror financially supports the Wikimedia Foundation. Displaying this page does not burden Wikipedia hardware resources. This article is from Wikipedia. All text is available under the terms of the GNU Free Documentation License. Contact: info [AT] wikimirror [DOT] com

As a side note, I still think that Wikipedia contributers (me included) ought to be aware of these facts (or the existence of this mirror), which may imply that their contributions become commercial and thus money-making.

Permissive Web Hosts

Computer shell
The command-line is the front line

NOT many Linux hosts permit SSH access to one’s own Webspace. There are concerns that the extra freedom and flexibility of the *NIX shell might lead to servers going down. If servers are shared, the problem is somewhat magnified and the threat is just not worth it.

There are, however, workarounds if you use cPanel as I revealed hack. I have also recently been told about PHP Console by its author. That shell can be installed on your Webspace and allow you to do what you otherwise could not.

It is neither the time nor the place to post promotional links, but I am very satisfied with my host that helped me combat the recent zombie attack on 2 of my sites. The attacks have now ended fully and I removed all my defences, namely filtering through re-directions. If you experienced some locational oddities recently, this must have been the cause.

Non-Evolving Search Engines and Operating Systems

3 Monkeys

Refuse to explore and stay a monkey forever

Google algorithms are very complex at present. The very core of these remains PageRank — a mechanism that often gets misused and leads to disasters (referrer spam among other link spamming techniques). For my recent zombie attacks I blame:

  • Google – for unintentionally leading to ‘link greed’, not understanding or anticipating streetsmarts and penetration of second- and third-world countries into the Internet
  • ISP‘s – for apathetically harbouring traffic that is pure spam or targetted attacks
  • Microsoft – for creating an operating system that is so easy for crooks to capture

Google’s algorithms have become like a horse that has lots of decorations upon it, but is nothing more than a brute-force horse underneath. You can take a pig, put it in a dress and take it out for dinner. But it’s still a pig in dress, not a girlfriend.

Search must Evolve. A flawed or limited principle at the heart of something is bound to fail no matter how many bits you hang atop to patch it up and improve it. This is why traditional page indexing is not a good method for approaching the problem of information extraction and discovery. Microsoft’s operating system, for instance, suffers from the very same problem where a flawed and too complex an operating system was build from ‘code spaghetti’. It was recently heard through the grapevine that Longhorn was thrown away and reverted merely to ground zero to be based on the XP-related Server 2003 code. This comes to show that weekly updates were merely patching a mordid mess. Microsoft recognise their inability to complete with Linux performance (uptime, flexibility — the a reason for Monad). Linux just took a right approach — a right paradigm if you like — all along and was therefore able to sweep along all the best programmers in the world.

Returning to Google, by relying heavily on PageRank and making ad-hoc improvements, no real innovation will be made. This is why I intimated Iuron a few days ago. It ought to turn a large pool of indexed page into actual knowledge and provide definitely answers rather than a linear scatter of related pages.

Also comes to mind are AltaVista and other antiquated search engines with very fundamental and not-so-cunning methods for scanning pages. These were very quickly relaced by backlink-based engines, i.e. link counting in 1998. No progress has been made in nearly 8 years, however. One may begin to conceive a Google killer rather than a Windows killer. The required resources, however, in particular data centres, make the (financial) entry barrier too high to initiate a substantial enough threat. Proprietaries, however, are no concrete barrier, in contrary to the case with operating systems. So, I remain optimistic and I might soon meet some professors whose expertise is the semantic Web.

With reference to the famous 3-monkey image on top (I also have one on top of my monitor), those who refuse to evolve (Ballmer) show ‘zoo symptoms’ already. I vividly recall the day when Scoble quoted Microsoft CEO, Steve Ballmer, saying that RSS has no future. This was roughly 7 months ago. Ballmer also said that Google would vanish in 5 years and promised that MSN search was bound to ‘kill’ Google. I say: live in the past, be the past.

Moving on to a different topic, the latest article with the theme of aging at Microsoft came out on the day when I first composed this item: Pity poor Microsoft’s midlife crisis. Another recent article I have just been informed of is At 30, Microsoft Grapples With Growing Up.

Recommended reading:

Q: What about all the people in the corporate environments who are forced to use MS products and aren’t allowed the option/choice to use Mac/Linux/UNIX?

A: Kick your boss’s ass, or, choose to work for a company who have decisions that you liked.

Windows Attacks the Web

Dear Windows users,

Please get your act together and always patch up your operating system, if not migrate to a less vulnerable operating system. The Macs, for example, are not so hard to use.

Your current machines are occasionally getting infected and then used as zombies in the midst of our network. Subsequently, under the control of evil hands, they commence a collateral onslaught on Web sites. Such sites, if not computers in your network, can be powered by Mac O/S, Linux or other UNIX variants. You must become responsible as you reside in a networked environment and can affect it tremendously without you being aware of it.

I am currently suffering from an international army of infected Windows machines. This is no laughing matter as I can confirm all zombies are Windows-driven and there are hundreds of them all over the world.

Windows out-of-the-box (i.e. unpatched) is somewhat of a weapon. It can easily drain bandwidth of other users and inject spam content. It can also lead to downtime of others in a global village that is the Internet. We are now hearing about a Dutch network of people who exploited vulnerable Windows machines world-wide. Be more cautious than ever before or else get disconnected by your ISP. Here at the University we already charge people if they continously get disconnected due to viruses in their Windows machines. Viruses (virii as the common slang) are causing a great deal of distress to network administrators, which is the reason for monetary penalty.

If the attacks on my site do not reach a halt or some solution is found, it might have to be isolated (READ: brought down), which is unacceptable. If not isolated, my site may take other sites down along with it in the future. I am not alone in the recent batch of attacks according to Tao of Mac.

Related items:

E-mail Count

Stuffed mailboxes

I am never too aure as to how E-mail should be counted (if it can be counted at all). There is a certain amount of traffic that can be quantified when it comes to E-mail, but not all E-mails should be equally treated. Points to ponder:

  • Should spam be counted?
  • Or spam that sank in a BoxTrapper and gets reviewed on occasions?
  • Should mailing lists be counted?
  • What about newsletters that anyone can sign up for?
  • E-mails with multiple recipients?
  • Uninvited E-mails that are not spam?
  • Automated messages?
  • E-mail one-liners?
  • E-mails sent to deprecated accounts that are rarely checked, if ever?

I have reached the conclusion that if E-mail bulk is ever to be counted, there must be some weighting applied to the variety of E-mail ‘types’. I have never counted my mail, but I am somewhat fed up with people who brag about the quantity of mail (or spam) they get. I have recently read somewhere that there services that one can sign up for in order to increase the amount of incoming spam traffic. Amazing, is it not? More amazing is the fact that people may actually sign up and welcome the increase in junk. It’s a market niche, right?

Knowledge Engine

Iuron

EARLY this morning I bought iuron.com. The host was admirably rapid in its automated response and setup. Subsequently, after just minutes since the purchase I was able to access the domain (DNS was up-to-date already) and initialise things, e.g. content, passwords, customised error pages, siteinfo.xml, robots.txt, favicon.ico, etc.

In the intermin I created a site/project logo using The GIMP. It took me roughly 10 minutes to produce the compact GIF version, of which I keep a non-lossy version (bitmap) as often required.

A detailed but not yet comprehensive proposal is bound to go public in the brand-new domain soon. The specifications are not public and are password-protected for the time being. Let the title and excerpt give you a rough indication of where the project is headed. I am yet to seek funding opportunities so many more details are (hopefully) soon to follow.

Feed Readers Everywhere

Man and his dogMany on-line services incorporate RSS feeds these days. Briefly observe, for instance, the vendor-specific feeds management for this Web log.

It turns out that Google have joined that game with what they call Google Reader. GReader is bound to deliver quite a lot given what we know about Google engineers and their love-affair with Web applications (more latterly videos).

For feeds, I am still using RSSOwl (JRE, hence inter-operable), of which I am a tester. I am also a tester of the Web-based, AJAX-rich Feedlounge (screenshot below), but admittedly I never contribute immensely as a tester. I use RSSOwl (screenshot below) almost exclusively. I am rather loyal to RSSOwl though I occasionally take Feedlounge for a short spin. Feedlounge is somewhat irresponsive, is still in its testing phases, facing an uncertain future and boasts commercial aspirations.

Feedlounge

Feedlounge in action

RSSOwl screenshot

RSSOwl back in April (click to enlarge)

Retrieval statistics: 21 queries taking a total of 0.093 seconds • Please report low bandwidth using the feedback form
Original styles created by Ian Main (all acknowledgements) • PHP scripts and styles later modified by Roy Schestowitz • Help yourself to a GPL'd copy
|— Proudly powered by W o r d P r e s s — based on a heavily-hacked version 1.2.1 (Mingus) installation —|